Qing En

dblp:189/4347 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0003-0173-7437ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing wheat pest detection: an edge-enhanced deformable attention network approach
Dongxue Liu, Yingchun Yuan, Qing En, Wei Ma 0008, Chunshan Wang, Zhenxue He, Fangfang Liang
Vis. Comput.3
2025 Attention-based unsupervised prompt learning for SAM in leaf disease segmentation
Luda Tian, Yingchun Yuan, Qing En, Wei Ma 0008, Fangfang Liang
Knowl. Based Syst.3
2025 Dynamic text prompt joint multimodal features for accurate plant disease image captioning
Fangfang Liang, Zhenxue He, Qing En
Vis. Comput.5
2024 Cross-model Mutual Learning for Exemplar-based Medical Image Segmentation
abstract
Medical image segmentation typically demands extensive dense annotations for model training, which is both time-consuming and skill-intensive. To mitigate this burden, exemplar-based medical image segmentation methods have been introduced to achieve effective training with only one annotated image. In this paper, we introduce a novel Cross-model Mutual learning framework for Exemplar-based Medical image Segmentation (CMEMS), which leverages two models to mutually excavate implicit information from unlabeled data at multiple granularities. CMEMS can eliminate confirmation bias and enable collaborative training to learn complementary information by enforcing consistency at different granularities across models. Concretely, cross-model image perturbation based mutual learning is devised by using weakly perturbed images to generate high-confidence pseudo-labels, supervising predictions of strongly perturbed images across models. This approach enables joint pursuit of prediction consistency at the image granularity. Moreover, cross-model multi-level feature perturbation based mutual learning is designed by letting pseudo-labels supervise predictions from perturbed multi-level features with different resolutions, which can broaden the perturbation space and enhance the robustness of our framework. CMEMS is jointly trained using exemplar data, synthetic data, and unlabeled data in an end-to-end manner. Experimental results on two medical image datasets indicate that the proposed CMEMS outperforms the state-of-the-art segmentation methods with extremely limited supervision.
Qing En, Yuhong Guo
AISTATS1
2024 Adaptive Parametric Prototype Learning for Cross-Domain Few-Shot Classification
abstract
Cross-domain few-shot classification induces a much more challenging problem than its in-domain counterpart due to the existence of domain shifts between the training and test tasks. In this paper, we develop a novel Adaptive Parametric Prototype Learning (APPL) method under the meta-learning convention for cross-domain few-shot classification. Different from existing prototypical few-shot methods that use the averages of support instances to calculate the class prototypes, we propose to learn class prototypes from the concatenated features of the support set in a parametric fashion and meta-learn the model by enforcing prototype-based regularization on the query set. In addition, we fine-tune the model in the target domain in a transductive manner using a weighted-moving-average self-training approach on the query instances. We conduct experiments on multiple cross-domain few-shot benchmark datasets. The empirical results demonstrate that APPL yields superior performance to many state-of-the-art cross-domain few-shot learning methods.
Marzi Heidari, Abdullah Alchihabi, Qing En, Yuhong Guo
AISTATS3
2024 Annotation by Clicks: A Point-Supervised Contrastive Variance Method for Medical Semantic Segmentation
Qing En, Yuhong Guo
BMVC1
2024 AKGNet: Attribute Knowledge Guided Unsupervised Lung-Infected Area Segmentation
Qing En, Yuhong Guo
ECML/PKDD (3)1
2024 Multi-prototype Co-saliency Model for Plant Disease Detection
Fangfang Liang, Qing En
PRCV (9)4
2024 Enhancing zero-shot object detection with external knowledge-guided robust contrast learning
Lijuan Duan, Qing En, Zhaoying Liu, Bian Ma
Pattern Recognit. Lett.3
2023 Exemplar-FreeSOLO: Enhancing Unsupervised Instance Segmentation with Exemplars
abstract
Instance segmentation seeks to identify and segment each object from images, which often relies on a large number of dense annotations for model training. To alleviate this burden, unsupervised instance segmentation methods have been developed to train class-agnostic instance segmentation models without any annotation. In this paper, we propose a novel unsupervised instance segmentation approach, Exemplar-FreeSOLO, to enhance unsupervised instance segmentation by exploiting a limited number of unannotated and unsegmented exemplars. The proposed framework offers a new perspective on directly perceiving top-down information without annotations. Specifically, Exemplar-FreeSOLO introduces a novel exemplar-knowledge abstraction module to acquire beneficial top-down guidance knowledge for instances using unsupervised exemplar object extraction. Moreover, a new exemplar embedding contrastive module is designed to enhance the discriminative capability of the segmentation model by exploiting the contrastive exemplar-based guidance knowledge in the embedding space. To evaluate the proposed Exemplar-FreeSOLO, we conduct comprehensive experiments and perform in-depth analyses on three image instance segmentation datasets. The experimental results demonstrate that the proposed approach is effective and outperforms the state-of-the-art methods.
Taoseef Ishtiak, Qing En, Yuhong Guo
CVPR2
2022 Exemplar Learning for Medical Image Segmentation
Qing En, Yuhong Guo
BMVC1
2022 Remember the Difference: Cross-Domain Few-Shot Semantic Segmentation via Meta-Memory Transfer
abstract
Few-shot semantic segmentation intends to predict pixel-level categories using only a few labeled samples. Existing few-shot methods focus primarily on the categories sampled from the same distribution. Nevertheless, this assumption cannot always be ensured. The actual domain shift problem significantly reduces the performance of few-shot learning. To remedy this problem, we propose an interesting and challenging cross-domain few-shot semantic segmentation task, where the training and test tasks perform on different domains. Specifically, we first propose a meta-memory bank to improve the generalization of the segmentation network by bridging the domain gap between source and target domains. The meta-memory stores the intra-domain style information from source domain instances and transfers it to target samples. Subsequently, we adopt a new contrastive learning strategy to explore the knowledge of different categories during the training stage. The negative and positive pairs are obtained from the proposed memory-based style augmentation. Comprehensive experiments demon-strate that our proposed method achieves promising results on cross-domain few-shot semantic segmentation tasks on COCO-20i, PASCAL-Si, FSS-1000, and SUIM datasets.
Wenjian Wang 0002, Lijuan Duan, Yuxi Wang 0001, Qing En, Junsong Fan, Zhaoxiang Zhang 0001
CVPR4
2021 TMD-FS: Improving Few-Shot Object Detection with Transformer Multi-modal Directing
Lijuan Duan, Wenjian Wang 0002, Qing En
PRCV (4)4
2021 Context-sensitive zero-shot semantic segmentation model based on meta-learning
Wenjian Wang 0002, Lijuan Duan, Qing En, Baochang Zhang 0001
Neurocomputing3
2021 Joint Multisource Saliency and Exemplar Mechanism for Weakly Supervised Video Object Segmentation
abstract
Weakly supervised video object segmentation (WSVOS) is a vital yet challenging task in which the aim is to segment pixel-level masks with only category labels. Existing methods still have certain limitations, e.g., difficulty in comprehending appropriate spatiotemporal knowledge and an inability to explore common semantic information with category labels. To overcome these challenges, we formulate a novel framework by integrating multisource saliency and incorporating an exemplar mechanism for WSVOS. Specifically, we propose a multisource saliency module to comprehend spatiotemporal knowledge by integrating spatial and temporal saliency as bottom-up cues, which can effectively eliminate disruptions due to confusing regions and identify attractive regions. Moreover, to our knowledge, we make the first attempt to incorporate an exemplar mechanism into WSVOS by proposing an adaptive exemplar module to process top-down cues, which can provide reliable guidance for co-occurring objects in intraclass videos and identify attentive regions. Our framework, which comprises the two aforementioned modules, offers a new perspective on directly constructing the correspondence between bottom-up cues and top-down cues when ground-truth information for the reference frames is lacking. Comprehensive experiments demonstrate that the proposed framework achieves state-of-the-art performance.
Qing En, Lijuan Duan, Zhaoxiang Zhang 0001
IEEE Trans. Image Process.1
2019 Human-Like Delicate Region Erasing Strategy for Weakly Supervised Detection
Qing En, Lijuan Duan, Zhaoxiang Zhang 0001, Xiang Bai
AAAI1
2019 Deep feature representation based on privileged knowledge transfer
Lijuan Duan, Qing En, Yuanhua Qiao, Laiyun Qing
Pattern Recognit. Lett.2
2016 Human action recognition based on discriminative supervoxels
abstract
Due to the diversity of body movements and uncertainty of recording occasion, human action recognition is still a challenging task, especially in real world. This paper provides a new method of representing the video with mid-level vision representation which is extracted from the discriminative supervoxels. In the proposed method, the discriminative supervoxels we extracted through a learning phase frequently occur within class and are distinguishing enough between classes. They contain the meaningful parts of the video, including specific background of an action and the moving human body. The video is first oversegmented to obtain supervoxels, which are described by the dense trajectories and Bag-Of-Words framework. Afterwards, the discriminative supervoxels are extracted by an iterative procedure through training and selecting. Finally the videos are represented with discriminative supervoxels. Experimental results on KTH, YouTube and UT-Interaction datasets demonstrate comparable performance with state-of-the-art models.
Lijuan Duan, Qing En, Juncheng Chen
IJCNN4