Mengya Han

dblp:264/2816 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0003-3499-3832ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 OACI: Object-aware contextual integration for image captioning
Shuhan Xu, Mengya Han, Wei Yu 0004, Zheng He 0001, Xin Zhou 0003, Yong Luo 0002
Knowl. Based Syst.2
2025 ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
abstract
Xuxu Liu, Siyuan Liang, Mengya Han, Yong Luo, Aishan Liu, Xiantao Cai, Zheng He, Dacheng Tao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xuxu Liu, Siyuan Liang 0004, Mengya Han, Yong Luo 0002, Aishan Liu, Xiantao Cai, Zheng He 0001, Dacheng Tao
ACL (1)3
2025 Open-Vocabulary Fine-Grained Hand Action Detection
abstract
In this work, we address the new challenge of open-vocabulary fine-grained hand action detection, which aims to recognize hand actions from both known and novel categories using textual descriptions. Traditional hand action detection methods are limited to closed-set detection, making it difficult for them to generalize to new, unseen hand action categories. While current open-vocabulary detection (OVD) methods are effective at detecting novel objects, they face challenges with fine-grained action recognition, particularly when data is limited and heterogeneous. This often leads to poor generalization and performance bias between base and novel categories. To address these issues, we propose a novel approach, Open-FGHA (Open-vocabulary Fine-Grained Hand Action), which learns to distinguish fine-grained features across multiple modalities from limited heterogeneous data. It then identifies optimal matching relationships among these features, enabling accurate open-vocabulary fine-grained hand action detection. Specifically, we introduce three key components: Hierarchical Heterogeneous Low-Rank Adaptation, Bidirectional Selection and Fusion Mechanism, and Cross-Modality Query Generator. These components work in unison to enhance the alignment and fusion of multimodal fine-grained features. Extensive experiments demonstrate that Open-FGHA outperforms existing OVD methods, showing its strong potential for open-vocabulary hand action detection. The source code is available at OV-FGHAD.
Ting Zhe, Mengya Han, Xiaoshuai Hao, Yong Luo 0002, Zheng He 0001, Xiantao Cai, Jing Zhang 0037
IJCAI2
2025 DM-PCL: Text-Driven Dual-Modal Prototype Consistency Learning for Weakly-Supervised Few-Shot Part Segmentation
Mengya Han, Yong Luo 0002, Han Hu 0003, Zengmao Wang, Lefei Zhang, Bo Du 0001, Ling-Yu Duan, Dacheng Tao
Int. J. Comput. Vis.1
2025 PartSeg: Few-shot part segmentation via part-aware prompt learning
Mengya Han, Heliang Zheng, Yong Luo 0002, Han Hu 0003, Jing Zhang 0037, Bo Du 0001
Pattern Recognit.1
2024 Textual Enhanced Adaptive Meta-Fusion for Few-Shot Visual Recognition
abstract
Few-shot learning (FSL) is a challenging task that aims to train a classifier to recognize novel categories, where only a few annotated examples are available in each category. Recently, many FSL approaches have been proposed based on the meta-learning paradigm, which attempts to learn transferable knowledge from similar tasks by designing a meta-learner. However, most of these approaches only exploit the information from visual modality and do not utilize ones from additional modalities (e.g., textual description). Since the labeled examples in FSL are limited, increasing the information on the examples is a probable solution to improve the classification performance. This motivates us to propose a novel meta-learning method, termed textual enhanced adaptive meta-fusion FSL (TAMF-FSL), which leverages both the visual information from the visual image and semantic information from language supervision. Specifically, TAMF-FSL exploits the semantic information of textual description to improve the visual-based models. We first employ a text encoder to learn the semantic features of each visual category, and then design a modality alignment module and meta-fusion module to align and fuse the visual and semantic features for final prediction. Extensive experiments show that the proposed method outperforms many recent or competitive FSL counterparts on two popular datasets.
Mengya Han, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Kehua Su, Bo Du 0001
IEEE Trans. Multim.1
2024 Not All Instances Contribute Equally: Instance-Adaptive Class Representation Learning for Few-Shot Visual Recognition
abstract
Few-shot visual recognition refers to recognize novel visual concepts from a few labeled instances. Many few-shot visual recognition methods adopt the metric-based meta-learning paradigm by comparing the query representation with class representations to predict the category of query instance. However, the current metric-based methods generally treat all instances equally and consequently often obtain biased class representation, considering not all instances are equally significant when summarizing the instance-level representations for the class-level representation. For example, some instances may contain unrepresentative information, such as too much background and information of unrelated concepts, which skew the results. To address the above issues, we propose a novel metric-based meta-learning framework termed instance-adaptive class representation learning network (ICRL-Net) for few-shot visual recognition. Specifically, we develop an adaptive instance revaluing network (AIRN) with the capability to address the biased representation issue when generating the class representation, by learning and assigning adaptive weights for different instances according to their relative significance in the support set of corresponding class. In addition, we design an improved bilinear instance representation and incorporate two novel structural losses, i.e., intraclass instance clustering loss and interclass representation distinguishing loss, to further regulate the instance revaluation process and refine the class representation. We conduct extensive experiments on four commonly adopted few-shot benchmarks: miniImageNet, tieredImageNet, CIFAR-FS, and FC100 datasets. The experimental results compared with the state-of-the-art approaches demonstrate the superiority of our ICRL-Net.
Mengya Han, Yibing Zhan, Yong Luo 0002, Bo Du 0001, Han Hu 0003, Yonggang Wen 0001, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2023 Two-Stream Prototype Learning Network for Few-Shot Face Recognition Under Occlusions
abstract
Few-shot face recognition under occlusion (FSFRO) aims to recognize novel subjects given only a few, probably occluded face images, and it is challenging and common in real-world scenarios. Unknown occlusions may deteriorate the class prototypes, while an occluded image in the support set may be critical for recognition if the query image is occluded. This motivates us to propose a novel Two-stream Prototype Learning Network (TSPLN) for FSFR under occlusions by simultaneously considering the quality of support images and their relevance to the query i mage. Specifically, we design a two-stream architecture, which mainly consists of a support-centered stream and query-centered stream, to learn the optimal class prototypes. The former stream is to reduce the negative impact of occluded images on the prototype. This is achieved by exploring the similarities between different images in the support set. In the query-centered stream, we exploit the relevance between the query and support set based on feature alignment (FA). We conduct extensive experiments on two popular datasets: CASIA-WebFace and RMFRD. The experimental results show that our proposed method achieves the state-of-the-art performance for occluded face recognition in the few-shot setting.
Mengya Han, Yong Luo 0002, Han Hu 0003, Yonggang Wen 0001
IEEE Trans. Multim.2
2022 Leveraging GAN Priors for Few-Shot Part Segmentation
abstract
Few-shot part segmentation aims to separate different parts of an object given only a few annotated samples. Due to the challenge of limited data, existing works mainly focus on learning classifiers over pre-trained features, failing to learn task-specific features for part segmentation. In this paper, we propose to learn task-specific features in a "pre-training"-"fine-tuning" paradigm. We conduct prompt designing to reduce the gap between the pre-train task (i.e., image generation) and the downstream task (i.e., part segmentation), so that the GAN priors for generation can be leveraged for segmentation. This is achieved by projecting part segmentation maps into the RGB space and conducting interpolation between RGB segmentation maps and original images. Specifically, we design a fine-tuning strategy to progressively tune an image generator into a segmentation generator, where the supervision of the generator varying from images to segmentation maps by interpolation. Moreover, we propose a two-stream architecture, i.e., a segmentation stream to generate task-specific features, and an image stream to provide spatial constraints. The image stream can be regarded as a self-supervised auto-encoder, and this enables our model to benefit from large-scale support images. Overall, this work is an attempt to explore the internal relevance between generation tasks and perception tasks by prompt designing. Extensive experiments show that our model can achieve state-of-the-art performance on several part segmentation datasets.
Mengya Han, Heliang Zheng, Yong Luo 0002, Han Hu 0003, Bo Du 0001
ACM Multimedia1
2022 Knowledge Graph enhanced Multimodal Learning for Few-shot Visual Recognition
abstract
Few-shot learning (FSL) aims to learn a classifier for novel classes with only a few labeled samples per category available. The mainstream FSL approaches fall in the meta-learning paradigm, where a meta-learner is used to learn transferable knowledge and generalize to new tasks. However, these approaches usually only leverage information from a single modality (e.g., visual image) and fail to explore the information from other modalities (e.g., the knowledge graph). Since the labeled samples are scarce in FSL, increasing the information for each example is a possible solution to improve the performance. This motivates us to develop a new meta-learning framework for few-shot visual recognition termed Knowledge Graph enhanced FSL (KGFSL), which combines the information from multiple modalities: 1) the visual information in images and 2) the rich semantics and structural information in a knowledge graph (KG). Specifically, KGFSL exploits the word embedding of the category and its relationship to other categories to improve the visual-based models. A graph convolutional network (GCN) is first introduced to learn the semantic embeddings for each node (a visual category) in KG. The visual and semantic embeddings are then aligned and combined for final prediction. Finally, the whole framework is trained in an end-to-end manner. We conduct extensive experiments on two widely-used FSL benchmarks: miniImageNet and tieredImageNet. Experimental results demonstrate the effectiveness of the multimodal information for few-shot learning, and our proposed method can significantly outperform the state-of-the-art approaches.
Mengya Han, Yibing Zhan, Baosheng Yu, Yong Luo 0002, Bo Du 0001, Dacheng Tao
MMSP1
2020 Multi-scale feature network for few-shot learning
Mengya Han, Ronggui Wang, Juan Yang 0001, Lixia Xue, Min Hu 0010
Multim. Tools Appl.1