EDBT 2026 Demo / reviewers in the wild / expert
Jun Wei 0006
dblp:68/6546-6
· DBLP profile ↗
16ranked-venue papers
9as first author
13since 2021 · last 2027
0000-0002-1263-5260ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 9 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | MedDATP: Adapting CLIP for few-shot medical image classification via domain adapter and task prompts
Zixun Zhang, Yuncheng Jiang 0002, Jun Wei 0006, Huazhu Fu, Shuguang Cui, Tao Luo 0014, Zhen Li 0026 |
Expert Syst. Appl. | 3 |
| 2024 | WeakPCSOD: Overcoming the Bias of Box Annotations for Weakly Supervised Point Cloud Salient Object DetectionabstractPoint cloud salient object detection (PCSOD) is a newly proposed task in 3D dense segmentation. However, the acquisition of accurate 3D dense annotations comes at a high cost, severely limiting the progress of PCSOD. To address this issue, we propose the first weakly supervised PCSOD (named WeakPCSOD) model, which relies solely on cheap 3D bounding box annotations. In WeakPCSOD, we extract noise-free supervision from coarse 3D bounding boxes while mitigating shape biases inherent in box annotations. To achieve this, we introduce a novel mask-to-box (M2B) transformation and a color consistency (CC) loss. The M2B transformation, from a shape perspective, disentangles predictions from labels, enabling the extraction of noiseless supervision from labels while preserving object shapes independently of the box bias. From an appearance perspective, we further introduce the CC loss to provide dense supervision, which mitigates the non-unique predictions stemming from weak supervision and substantially reduces prediction variability. Furthermore, we employ a self-training (ST) strategy to enhance performance by utilizing high-confidence pseudo labels. Notably, the M2B transformation, CC loss, and ST strategy are seamlessly integrated into any model and incur no computational costs for inference. Extensive experiments demonstrate the effectiveness of our WeakPCSOD model, even comparable to fully supervised models utilizing dense annotations. Jun Wei 0006, Shaohua Kevin Zhou, Shuguang Cui, Zhen Li 0026 |
AAAI | 1 |
| 2024 | MixPolyp: Integrating Mask, Box and Scribble Supervision for Enhanced Polyp SegmentationabstractLimited by the expensive labeling, polyp segmentation models are plagued by data shortages. To tackle this, we propose the mixed supervised polyp segmentation paradigm (MixPolyp). Unlike traditional models relying on a single type of annotation, MixPolyp combines diverse annotation types (mask, box, and scribble) within a single model, thereby expanding the range of available data and reducing labeling costs. To achieve this, MixPolyp introduces three novel supervision losses to handle various annotations: Subspace Projection loss $\left({{{\mathcal{L}}_{{\mathcal{S}}{\mathcal{P}}}}}\right)$, Binary Minimum Entropy loss $\left({{{\mathcal{L}}_{{\mathcal{B}}{\mathcal{M}}{\mathcal{E}}}}}\right)$, and Linear Regularization loss $\left({{{\mathcal{L}}_{{\mathcal{L}}{\mathcal{R}}}}}\right)$. For box annotations, ${{\mathcal{L}}_{{\mathcal{S}}{\mathcal{P}}}}$ eliminates shape inconsistencies between the prediction and the supervision. For scribble annotations, ${{\mathcal{L}}_{{\mathcal{B}}{\mathcal{M}}{\mathcal{E}}}}$ provides supervision for unlabeled pixels through minimum entropy constraint, thereby alleviating supervision sparsity. Furthermore, ${{\mathcal{L}}_{{\mathcal{L}}{\mathcal{R}}}}$ provides dense supervision by enforcing consistency among the predictions, thus reducing the non-uniqueness. These losses are independent of the model structure, making them generally applicable. They are used only during training, adding no computational cost during inference. Extensive experiments on five datasets demonstrate MixPolyp’s effectiveness. Yiwen Hu 0001, Jun Wei 0006, Yuncheng Jiang 0002, Shuguang Cui, Zhen Li 0026 |
BIBM | 2 |
| 2024 | Let Video Teaches You More: Video-to-Image Knowledge Distillation using Detection TRansformer for Medical Video Lesion DetectionabstractAI-assisted lesion detection models play a crucial role in the early screening of cancer. However, previous image-based models ignore the inter-frame contextual information present in videos. On the other hand, video-based models capture the inter-frame context but are computationally expensive. To mitigate this contradiction, we delve into Video-to-Image knowledge distillation leveraging DEtection TRansformer (V2I-DETR) for the task of medical video lesion detection. V2I-DETR adopts a teacher-student network paradigm. The teacher network aims at extracting temporal contexts from multiple frames and transferring them to the student network, and the student network is an image-based model dedicated to fast prediction in inference. By distilling multi-frame contexts into a single frame, the proposed V2I-DETR combines the advantages of utilizing temporal contexts from video-based models and the inference speed of image-based models. Through extensive experiments, V2I-DETR outperforms previous state-of-the-art methods by a large margin while achieving the real-time inference speed (30 FPS) as the image-based model. Yuncheng Jiang 0002, Zixun Zhang, Jun Wei 0006, Chun-Mei Feng 0001, Guanbin Li, Shuguang Cui, Zhen Li 0026 |
BIBM | 3 |
| 2024 | Towards a Benchmark for Colorectal Cancer Segmentation in Endorectal Ultrasound Videos: Dataset and Model Development
Yuncheng Jiang 0002, Yiwen Hu 0001, Zixun Zhang, Jun Wei 0006, Chun-Mei Feng 0001, Xuemei Tang, Yong Liu 0026, Shuguang Cui, Zhen Li 0026 |
MICCAI (8) | 4 |
| 2023 | ScribblePolyp: Scribble-Supervised Polyp Segmentation through Dual Consistency AlignmentabstractAutomatic polyp segmentation models play a pivotal role in the clinical diagnosis of gastrointestinal diseases. In previous studies, most methods relied on fully supervised approaches, necessitating pixel-level annotations for model training. However, the creation of pixel-level annotations is both expensive and time-consuming, impeding the development of model generalization. In response to this challenge, we introduce ScribblePolyp, a novel scribble-supervised polyp segmentation framework. Unlike fully-supervised models, ScribblePolyp only requires the annotation of two lines (scribble labels) for each image, significantly reducing the labeling cost. Despite the coarse nature of scribble labels, which leave a substantial portion of pixels unlabeled, we propose a two-branch consistency alignment approach to provide supervision for these unlabeled pixels. The first branch employs transformation consistency alignment to narrow the gap between predictions under different transformations of the same input image. The second branch leverages affinity propagation to refine predictions into a soft version, extending additional supervision to unlabeled pixels. In summary, ScribblePolyp is an efficient model that does not rely on teacher models or moving average pseudo labels during training. Extensive experiments on the SUN-SEG dataset underscore the effectiveness of ScribblePolyp, achieving a Dice score of 0.8155, with the potential for a 1.8% improvement in the Dice score through a straightforward self-training strategy. Zixun Zhang, Yuncheng Jiang 0002, Jun Wei 0006, Hannah Cui, Zhen Li 0026 |
BIBM | 3 |
| 2023 | WeakPolyp: You only Look Bounding Box for Polyp Segmentation
Jun Wei 0006, Yiwen Hu 0001, Shuguang Cui, Shaohua Kevin Zhou, Zhen Li 0026 |
MICCAI (3) | 1 |
| 2022 | Weakly Supervised Object Localization Through Inter-class Feature Similarity and Intra-class Appearance Consistency
Jun Wei 0006, Sheng Wang 0001, Shaohua Kevin Zhou, Shuguang Cui, Zhen Li 0026 |
ECCV (30) | 1 |
| 2022 | BoxPolyp: Boost Generalized Polyp Segmentation Using Extra Coarse Bounding Box Annotations
Jun Wei 0006, Yiwen Hu 0001, Guanbin Li, Shuguang Cui, Shaohua Kevin Zhou, Zhen Li 0026 |
MICCAI (3) | 1 |
| 2022 | Prior knowledge facilitates low homologous protein secondary structure prediction with DSM distillationabstractMOTIVATION: Protein secondary structure prediction (PSSP) is one of the fundamental and challenging problems in the field of computational biology. Accurate PSSP relies on sufficient homologous protein sequences to build the multiple sequence alignment (MSA). Unfortunately, many proteins lack homologous sequences, which results in the low quality of MSA and poor performance. In this article, we propose the novel dynamic scoring matrix (DSM)-Distil to tackle this issue, which takes advantage of the pretrained BERT and exploits the knowledge distillation on the newly designed DSM features. Specifically, we propose the DSM to replace the widely used profile and PSSM (position-specific scoring matrix) features. DSM could automatically dig for the suitable feature for each residue, based on the original profile. Namely, DSM-Distil not only could adapt to the low homologous proteins but also is compatible with high homologous ones. Thanks to the dynamic property, DSM could adapt to the input data much better and achieve higher performance. Moreover, to compensate for low-quality MSA, we propose to generate the pseudo-DSM from a pretrained BERT model and aggregate it with the original DSM by adaptive residue-wise fusion, which helps to build richer and more complete input features. In addition, we propose to supervise the learning of low-quality DSM features using high-quality ones. To achieve this, a novel teacher-student model is designed to distill the knowledge from proteins with high homologous sequences to that of low ones. Combining all the proposed methods, our model achieves the new state-of-the-art performance for low homologous proteins. RESULTS: Compared with the previous state-of-the-art method 'Bagging', DSM-Distil achieves an improvement about 5% and 7.3% improvement for proteins with MSA count ≤30 and extremely low homologous cases, respectively. We also compare DSM-Distil with Alphafold2 which is a state-of-the-art framework for protein structure prediction. DSM-Distil outperforms Alphafold2 by 4.1% on extremely low-quality MSA on 8-state secondary structure prediction. Moreover, we release a large-scale up-to-date test dataset BC40 for low-quality MSA structure prediction evaluation. AVAILABILITY AND IMPLEMENTATION: BC40 dataset: https://drive.google.com/drive/folders/15vwRoOjAkhhwfjDk6-YoKGf4JzZXIMC. HardCase dataset: https://drive.google.com/drive/folders/1BvduOr2b7cObUHy6GuEWk-aUkKJgzTUv. Code: https://github.com/qinwang-ai/DSM-Distil. Qin Wang 0011, Jun Wei 0006, Mingzhi Lin, Ruobing Ren, Sheng Wang 0001, Shuguang Cui, Zhen Li 0026 |
Bioinform. | 2 |
| 2021 | Shallow Feature Matters for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to localize objects by only utilizing image-level labels. Class activation maps (CAMs) are the commonly used features to achieve WSOL. However, previous CAM-based methods did not take full advantage of the shallow features, despite their importance for WSOL. Because shallow features are easily buried in background noise through conventional fusion. In this paper, we propose a simple but effective Shallow feature-aware Pseudo supervised Object Localization (SPOL) model for accurate WSOL, which makes the utmost of low-level features embedded in shallow layers. In practice, our SPOL model first generates the CAMs through a novel element-wise multiplication of shallow and deep feature maps, which filters the background noise and generates sharper boundaries robustly. Besides, we further propose a general class-agnostic segmentation model to achieve the accurate object mask, by only using the initial CAMs as the pseudo label without any extra annotation. Eventually, a bounding box extractor is applied to the object mask to locate the target. Experiments verify that our SPOL outperforms the state-of-the-art on both CUB- 200 and ImageNet-1K benchmarks, achieving 93.44% and 67.15% (i.e., 3.93% and 2.13% improvement) Top-5 localization accuracy, respectively. Jun Wei 0006, Qin Wang 0011, Zhen Li 0026, Sheng Wang 0001, Shaohua Kevin Zhou, Shuguang Cui |
CVPR | 1 |
| 2021 | Adaptive Residue-wise Profile Fusion for Low Homologous Protein Secondary Structure Prediction Using External KnowledgeabstractProtein secondary structure prediction (PSSP) is essential for protein function analysis. However, for low homologous proteins, the PSSP suffers from insufficient input features. In this paper, we explicitly import external self-supervised knowledge for low homologous PSSP under the guidance of residue-wise (amino acid wise) profile fusion. In practice, we firstly demonstrate the superiority of profile over Position-Specific Scoring Matrix (PSSM) for low homologous PSSP. Based on this observation, we introduce the novel self-supervised BERT features as the pseudo profile, which implicitly involves the residue distribution in all native discovered sequences as the complementary features. Furthermore, a novel residue-wise attention is specially designed to adaptively fuse different features (i.e., original low-quality profile, BERT based pseudo profile), which not only takes full advantage of each feature but also avoids noise disturbance. Besides, the feature consistency loss is proposed to accelerate the model learning from multiple semantic levels. Extensive experiments confirm that our method outperforms state-of-the-arts (i.e., 4.7% for extremely low homologous cases on BC40 dataset). Qin Wang 0011, Jun Wei 0006, Zhen Li 0026, Sheng Wang 0001, Shuguang Cui |
IJCAI | 2 |
| 2021 | Shallow Attention Network for Polyp Segmentation
Jun Wei 0006, Yiwen Hu 0001, Ruimao Zhang, Zhen Li 0026, Shaohua Kevin Zhou, Shuguang Cui |
MICCAI (1) | 1 |
| 2020 | F³Net: Fusion, Feedback and Focus for Salient Object DetectionabstractMost of existing salient object detection models have achieved great progress by aggregating multi-level features extracted from convolutional neural networks. However, because of the different receptive fields of different convolutional layers, there exists big differences between features generated by these layers. Common feature fusion strategies (addition or concatenation) ignore these differences and may cause suboptimal solutions. In this paper, we propose the F3Net to solve above problem, which mainly consists of cross feature module (CFM) and cascaded feedback decoder (CFD) trained by minimizing a new pixel position aware loss (PPA). Specifically, CFM aims to selectively aggregate multi-level features. Different from addition and concatenation, CFM adaptively selects complementary components from input features before fusion, which can effectively avoid introducing too much redundant information that may destroy the original features. Besides, CFD adopts a multi-stage feedback mechanism, where features closed to supervision will be introduced to the output of previous layers to supplement them and eliminate the differences between features. These refined features will go through multiple similar iterations before generating the final saliency maps. Furthermore, different from binary cross entropy, the proposed PPA loss doesn't treat pixels equally, which can synthesize the local structure information of a pixel to guide the network to focus more on local details. Hard pixels from boundaries or error-prone parts will be given more attention to emphasize their importance. F3Net is able to segment salient object regions accurately and provide clear local details. Comprehensive experiments on five benchmark datasets demonstrate that F3Net outperforms state-of-the-art approaches on six evaluation metrics. Code will be released at https://github.com/weijun88/F3Net. Jun Wei 0006, Shuhui Wang, Qingming Huang |
AAAI | 1 |
| 2020 | Label Decoupling Framework for Salient Object DetectionabstractTo get more accurate saliency maps, recent methods mainly focus on aggregating multi-level features from fully convolutional network (FCN) and introducing edge information as auxiliary supervision. Though remarkable progress has been achieved, we observe that the closer the pixel is to the edge, the more difficult it is to be predicted, because edge pixels have a very imbalance distribution. To address this problem, we propose a label decoupling framework (LDF) which consists of a label decoupling (LD) procedure and a feature interaction network (FIN). LD explicitly decomposes the original saliency map into body map and detail map, where body map concentrates on center areas of objects and detail map focuses on regions around edges. Detail map works better because it involves much more pixels than traditional edge supervision. Different from saliency map, body map discards edge pixels and only pays attention to center areas. This successfully avoids the distraction from edge pixels during training. Therefore, we employ two branches in FIN to deal with body map and detail map respectively. Feature interaction (FI) is designed to fuse the two complementary branches to predict the saliency map, which is then used to refine the two branches again. This iterative refinement is helpful for learning better representations and more precise saliency maps. Comprehensive experiments on six benchmark datasets demonstrate that LDF outperforms state-of-the-art approaches on different evaluation metrics. Jun Wei 0006, Shuhui Wang, Zhe Wu 0006, Chi Su, Qingming Huang, Qi Tian 0001 |
CVPR | 1 |
| 2019 | Active Perception Network for Salient Object DetectionabstractTo get better saliency maps for salient object detection, recent methods fuse features from different levels of convolutional neural networks and have achieved remarkable progress. However, the differences between different feature levels bring difficulties to the fusion process, thus it may lead to unsatisfactory saliency predictions. To address this issue, we propose Active Perception Network (APN) to enhance inter-feature consistency for salient object detection. First, Mutual Projection Module (MPM) is developed to fuse different features, which uses high-level features as guided information to extract complementary components from low-level features, and can suppress background noises and improve semantic consistency. Self Projection Module (SPM) is designed to further refine the fused features, which can be considered as the extended version of residual connection. Features that pass through SPM can produce more accurate saliency maps. Finally, we propose Head Projection Module (HPM) to aggregate global information, which brings strong semantic consistency to the whole network. Comprehensive experiments on five benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches on different evaluation metrics. Jun Wei 0006, Shuhui Wang, Liang Li 0003, Qingming Huang |
MMAsia | 1 |