Mengting Hu 0002

dblp:199/5022-2 · DBLP profile ↗
← Back
27ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0003-1536-5400ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 10 first-author · 17 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 AI-Generated Image Homology Detection
Hang Gao 0003, Rui Ba, Kaiye Yu, Han Xing, Mengting Hu 0002
KSEM (3)5
2026 CIDC: Cluster Identification-Guided Dual Correction for Robust Short Text Clustering
abstract
The rapid growth of online short texts has made specialized analysis essential, as these texts are sparse and information-limited. Short text clustering (STC) is critical for automatically grouping unlabeled texts into meaningful clusters, supporting applications such as sentiment analysis, spam filtering, and social media personalization. In the context of massive online content, deep clustering seeks to uncover semantic categories by measuring distances in the representation space. Consequently, aligning clustering pseudo-labels with the true category distribution is crucial for effective self-supervised training, particularly under class imbalance and distribution skew commonly observed in web data. To address this challenge, we propose the Cluster Identification-Guided Dual Correction (CIDC) framework, which generates reliable pseudo-labels to guide deep clustering. Specifically, given cluster partitions and model-estimated class distributions, we perform Cluster Category Identification (CCI) at each training epoch to determine the most probable category for each cluster. This identification provides the foundation for the Pseudo-Label Correction (PLC) and Prototype-Based Correction (PBC) modules, which jointly enhance pseudo-label reliability and representation learning. In the PLC module, samples whose model-estimated class distributions conflict with the assigned cluster category are corrected, thereby improving semantic alignment within clusters. In the PBC module, representative and reliable prototypes are selected according to cluster categories and model predictions to guide training, further strengthening representation discriminability. Extensive experiments demonstrate that CIDC consistently outperforms existing methods in terms of clustering accuracy and mutual information, particularly in unsupervised settings characterized by class imbalance and noisy data.
Yuhua Zhao 0001, Zhixin Han, Peiyu Xu, Hang Gao 0003, Mengting Hu 0002, Tiegang Gao
WWW6
2026 Correlation-guided mixture of experts prompt learning for long-tailed multi-label text classification
Ge Lan, Mengting Hu 0002
Knowl. Based Syst.2
2025 Enhancing Fake News Detection by Incorporating Evidence Credibility
abstract
The evidence-aware fake news detection aims to determine the veracity of claims under the guidance of external evidences. However, existing methods often neglect the credibility of evidences, making them vulnerable to misinformation in real-world scenarios where the evidence credibility is not always guaranteed. In this paper, we incorporate evidence credibility into fake news detection and propose a novel framework named ECFEND, which explicitly models the varying credibility of different evidences. Moreover, we present a new benchmark, SnopesCG, designed to simulate more realistic and challenging scenarios. Each claim in the benchmark is associated with noisy evidences retrieved from web pages as well as generated interference ones. Experimental results demonstrate the superiority of ECFEND over state-of-the-art methods, particularly on SnopesCG. We have open-sourced the code at: https://github.com/nffxdhd88/ECFEND.
Yike Wu 0002, Mengying Liu, Mengting Hu 0002
IJCNN6
2025 HCDS: Hierarchical Clustering for Cold-Start Few-Shot Data Selection
abstract
Deep learning models usually require large labeled datasets to generalize well, but this is computationally and financially costly. Cold-start few-shot data selection enables fast model generalization by selecting a few diverse, representative samples from an unlabeled data pool. To achieve this goal, previous work usually divides the training data into several clusters and performs sampling from these clusters. Yet, such a way tends to have two issues. First, imbalanced data distribution in the training data pool still exists in the selected subset, causing models' performance biases and suboptimal generalization ability. Second, these methods improve sample diversity in each cluster by considering either the feature dissimilarity among instances, or model uncertainty for individual instance. They ignore the entire representativeness of samples within a cluster. To tackle these challenges, we propose a novel framework HCDS : Hierarchical Clustering for Cold-Start Few-Shot Data Selection. Specifically, we first perform class-level clustering, using pseudo-labels for class supervision and applying contrastive clustering to derive class-rich features. We then refine these features within the class-level clusters into semantically meaningful features and perform representation-level clustering. Finally, we sample data from the representation-level clusters based on global similarity to ensure representativeness. Experimental results on six public datasets, including both balanced and imbalanced ones, show that HCDS achieves state-of-the-art performance, particularly with limited and imbalanced data.
Yuhua Zhao 0001, Zhixin Han, Xunzhi Wang, Bitong Luo, Hang Gao 0003, Minlie Huang, Mengting Hu 0002
SIGIR7
2025 OmniNER2025: Diverse and Comprehensive Fine-Grained NER Dataset and Benchmark for Chinese
abstract
As Named Entity Recognition (NER) tasks have evolved, artificial intelligence has been widely applied in this field. However, most benchmarks are limited to English, making it challenging to replicate successful experiences in other languages. To expand NER to informal and diverse Chinese text scenarios, we have proposed a new large-scale Chinese NER dataset, OmniNER2025. This dataset, obtained from user posts on a popular Chinese social media platform Xiaohongshu, contains 195,568 samples and 89 categories, all manually annotated. To our knowledge, it is currently the largest Chinese open-source NER dataset in terms of sample size, category diversity, and domain coverage. This dataset is more challenging than existing Chinese NER datasets and better reflects real-world applications. The large sample size and diverse entity types provide valuable research resources. Additionally, we introduced the ERRTA tool for error analysis and teacher model guidance, significantly reducing model errors and improving performance. In the future, we will refine the ERRTA framework and explore optimization strategies to enhance the practical value of NER models. By releasing the OmniNER2025 dataset and introducing the ERRTA tool, we have advanced fine-grained NER research and improved model performance, promoting its application and development in real-world scenarios.
Shuaipeng Liu, Mengting Hu 0002, Wen Dai, Xiaowei Zhao 0003, Xiujuan Xu
SIGIR4
2024 ToMBench: Benchmarking Theory of Mind in Large Language Models
abstract
Zhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengting Hu, Yunghwei Lai, Zexuan Xiong, Minlie Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhuang Chen 0002, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengting Hu 0002, Yunghwei Lai, Zexuan Xiong, Minlie Huang
ACL (1)8
2024 BvSP: Broad-view Soft Prompting for Few-Shot Aspect Sentiment Quad Prediction
abstract
Yinhao Bai, Yalan Xie, Xiaoyi Liu, Yuhua Zhao, Zhixin Han, Mengting Hu, Hang Gao, Renhong Cheng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yinhao Bai, Yalan Xie, Yuhua Zhao 0001, Zhixin Han, Mengting Hu 0002, Hang Gao 0003, Renhong Cheng
ACL (1)6
2024 Towards Robust Evidence-Aware Fake News Detection via Improving Semantic Perception
abstract
Evidence-aware fake news detection aims to determine the veracity of a given news (i.e., claim) with external evidences. We find that existing methods lack sufficient semantic perception and are easily blinded by textual expressions. For example, they still make the same prediction after we flip the semantics of a claim, which makes them vulnerable to malicious attacks. In this paper, we propose a model-agnostic training framework to improve the semantic perception of evidence-aware fake news detection. Specifically, we first introduce two kinds of data augmentation to complement the original training set with synthetic data. The semantic-flipped augmentation synthesizes claims with similar textual expressions but opposite semantics, while the semantic-invariant augmentation synthesizes claims with the same semantics but different writing styles. Moreover, we design a novel module to learn better claim representation which is more sensitive to the semantics, and further incorporate it into a multi-objective optimization paradigm. In the experiments, we also extend the original test set of benchmark datasets with the synthetic data to better evaluate the model perception of semantics. Experimental results demonstrate that our approach significantly outperforms the state-of-the-art methods on the extended test set, while achieving competitive performance on the original one. Our source code are released at https://github.com/Xyang1998/RobustFND.
Yike Wu 0002, Mengting Hu 0002, Mengying Liu
LREC/COLING3
2024 Towards Robust Information Extraction via Binomial Distribution Guided Counterpart Sequence
abstract
Information extraction (IE) aims to extract meaningful structured tuples from unstructured text. Existing studies usually utilize a pre-trained generative language model that rephrases the original sentence into a target sequence, which can be easily decoded as tuples. However, traditional evaluation metrics treat a slight error within the tuple as an entire prediction failure, which is unable to perceive the correctness extent of a tuple. For this reason, we first propose a novel IE evaluation metric called Matching Score to evaluate the correctness of the predicted tuples in more detail. Moreover, previous works have ignored the effects of semantic uncertainty when focusing on the generation of the target sequence. We argue that leveraging the built-in semantic uncertainty of language models is beneficial for improving its robustness. In this work, we propose Binomial distribution guided counterpart sequence (BCS) method, which is a model-agnostic approach. Specifically, we propose to quantify the built-in semantic uncertainty of the language model by bridging all local uncertainties with the whole sequence. Subsequently, with the semantic uncertainty and Matching Score, we formulate a unique binomial distribution for each local decoding step. By sampling from this distribution, a counterpart sequence is obtained, which can be regarded as a semantic complement to the target sequence. Finally, we employ the Kullback-Leibler divergence to align the semantics of the target sequence and its counterpart. Extensive experiments on 14 public datasets over 5 information extraction tasks demonstrate the effectiveness of our approach on various methods. Our code and dataset are available at https://github.com/byinhao/BCS.
Yinhao Bai, Yuhua Zhao 0001, Zhixin Han, Hang Gao 0003, Chao Xue 0003, Mengting Hu 0002
KDD6
2024 LinkNER: Linking Local Named Entity Recognition Models to Large Language Models using Uncertainty
abstract
Named Entity Recognition (NER) serves as a fundamental task in natural language understanding, bearing direct implications for web content analysis, search engines, and information retrieval systems. Fine-tuned NER models exhibit satisfactory performance on standard NER benchmarks. However, due to limited fine-tuning data and lack of knowledge, it performs poorly on unseen entity recognition. As a result, the usability and reliability of NER models in web-related applications are compromised. Instead, Large Language Models (LLMs) like GPT-4 possess extensive external knowledge, but research indicates that they lack specialty for NER tasks. Furthermore, non-public and large-scale weights make tuning LLMs difficult. To address these challenges, we propose a framework that combines small fine-tuned models with LLMs (LinkNER) and an uncertainty-based linking strategy called RDC that enables fine-tuned models to complement black-box LLMs, achieving better performance. We experiment with both standard NER test sets and noisy social media datasets. LinkNER enhances NER task performance, notably surpassing SOTA models in robustness tests. We also quantitatively analyze the influence of key components like uncertainty estimation methods, LLMs, and in-context learning on diverse NER tasks, offering specific web-related recommendations.
Zhen Zhang 0048, Yuhua Zhao 0001, Hang Gao 0003, Mengting Hu 0002
WWW4
2024 Modeling Category Semantic and Sentiment Knowledge for Aspect-Level Sentiment Analysis
abstract
To classify the sentiment polarity of the aspect entity in a sentence, most existing research evaluates the semantic knowledge among a certain aspect of a sentence and corresponding context as significant clues for the task. However, available accompanying information has not been completely exploited, especially the coarse-grained category-level knowledge in contexts. Such knowledge can help to alleviate polysemy and ambivalence problems. In this paper, we propose a multi-task learning framework Co-interactive Attention Network(CoAN) to jointly learn and handle multiple granularity features at both target and category levels. In order to leverage the fine-grained and coarse-grained knowledge in contexts and get multi-granularity sentiment related sentence representations, we introduce two co-interactive attention layers to conduct accompanying semantic interactions at the word-level and the feature-level. The experimental results on three restaurant review datasets prove that CoAN is superior to the baselines by 1.41% in accuracy and 2.81% in F1-score. Furthermore, ablation studies and attention visualizations show that the multi-task framework and novel co-interactive mechanisms can distinguish and fuse multi-granularity knowledge, which benefits the two subtasks in aspect based sentiment analysis.
Yuan Wang 0021, Peng Huo, Lingyan Tang, Mengting Hu 0002, Qi Yu 0005, Jucheng Yang 0001
IEEE Trans. Affect. Comput.5
2024 Learning Driver-Irrelevant Features for Generalizable Driver Behavior Recognition
abstract
Traffic accidents caused by driver distractions have seriously endangered public safety, with driver distractions typically stemming from behaviors beyond safe driving. Recently, vision-based driver behavior recognition has attracted much attention, achieving great success with deep learning-based schemes. However, the generalization ability of these models in real-world scenarios remains unsatisfactory. In this paper, we conduct an in-depth investigation into the underlying causes of this unsatisfactory generalization and conclude that the behavior features extracted by convolutional neural networks are intertwined with driver identity features. Based on this discovery, we propose a feature decomposition (FD) framework to disentangle these two types of features. The separated behavior features, referred to as driver-irrelevant behavior features, are subsequently leveraged for behavior recognition. Moreover, we introduce a co-training strategy to optimize the FD framework. This strategy enables behavior features and identity features to provide mutual auxiliary signals and encourages each other to drop the information that do not belong to them, so that the learned behavior features can be driver-irrelevant. Rigorous experiments are conducted on two widely-studied datasets, yielding results that demonstrate the superior performance of our framework and its improved generalization capabilities. Importantly, our framework’s fast inference capabilities make it highly suitable for real-world scenarios. Codes are released at https://github.com/gaohangcodes/ LearningDriverIrrelevantFeatures4DBR.
Hang Gao 0003, Mengting Hu 0002, Yi Liu 0002
IEEE Trans. Intell. Transp. Syst.2
2023 rT5: A Retrieval-Augmented Pre-trained Model for Ancient Chinese Entity Description Generation
Mengting Hu 0002, Xiaoqun Zhao, Xiaosu Sun, Zhengdan Li, Yike Wu 0002
NLPCC (1)1
2023 Contrastive knowledge integrated graph neural networks for Chinese medical text classification
Ge Lan, Mengting Hu 0002, Ye Li 0016
Eng. Appl. Artif. Intell.2
2023 Fine-Grained Domain Adaptation for Aspect Category Level Sentiment Analysis
abstract
Aspect category level sentiment analysis aims to identify the sentiment polarities towards the aspect categories discussed in a sentence. It usually suffers from a lack of labeled data. A popular solution is to transfer knowledge from a labeled source domain to an unlabeled target domain by unsupervised domain adaptation. However, most domain adaptation methods in sentiment analysis are coarse-grained, considering the source or target domain as a whole during the adaptation. We argue that these single-source single-target methods are inefficient since they ignore the difference between different aspect categories. In this article, we propose a fine-grained domain adaptation method to address the aspect category level sentiment analysis task by considering the adaptation between subdomains. Specifically, the source/target domain is divided into multiple subdomains according to the hierarchical structure of the aspect categories. We then design a multi-source multi-target transfer network to achieve fine-grained transfer. Extensive experimental results demonstrate the effectiveness of our fine-grained domain adaptation method on aspect category level sentiment analysis.
Mengting Hu 0002, Hang Gao 0003, Yike Wu 0002, Zhong Su, Shiwan Zhao
IEEE Trans. Affect. Comput.1
2023 Hybrid Regularizations for Multi-Aspect Category Sentiment Analysis
abstract
Aspect level sentiment classification aims to identify the sentiment polarity towards a particular aspect in a sentence. Previous attention-based methods generate an aspect-specific representation for each aspect and employ it to classify the sentiment polarity. However, normalized attention scores scatter over every word in the sentence, resulting in two issues. First, the attention may inherently introduce noise and downgrade the performance. Second, the opinion words may be “diluted” by other words, while the opinion feature should dominate for sentiment analysis. The issues become more severe in multi-aspect sentences. In this paper, we address the above two issues via hybrid regularizations, i.e.,aspect-levelandtask-level regularizations. Concretely, the aspect-level regularizations constrain the attention weights to alleviate noise. Among them, orthogonal regularization is designed for multi-aspect sentences and sparse regularization is for single-aspect sentences. To extract sentiment-dominant features, task-level regularization is proposed by introducing an orthogonal auxiliary task, i.e., aspect category detection. This regularization can allocate task-oriented context information for specific downstream tasks. Extensive experimental results on three public datasets demonstrate the effectiveness of the proposed approach in both single-task and multi-task scenarios.
Mengting Hu 0002, Shiwan Zhao, Zhong Su
IEEE Trans. Affect. Comput.1
2022 Classical Sequence Match Is a Competitive Few-Shot One-Class Learner
abstract
Nowadays, transformer-based models gradually become the default choice for artificial intelligence pioneers. The models also show superiority even in the few-shot scenarios. In this paper, we revisit the classical methods and propose a new few-shot alternative. Specifically, we investigate the few-shot one-class problem, which actually takes a known sample as a reference to detect whether an unknown instance belongs to the same class. This problem can be studied from the perspective of sequence match. It is shown that with meta-learning, the classical sequence match method, i.e. Compare-Aggregate, significantly outperforms transformer ones. The classical approach requires much less training cost. Furthermore, we perform an empirical comparison between two kinds of sequence match approaches under simple fine-tuning and meta-learning. Meta-learning causes the transformer models’ features to have high-correlation dimensions. The reason is closely related to the number of layers and heads of transformer models. Experimental codes and data are available at https://github.com/hmt2014/FewOne.
Mengting Hu 0002, Hang Gao 0003, Yinhao Bai
COLING1
2022 Improving Aspect Sentiment Quad Prediction via Template-Order Data Augmentation
abstract
Recently, aspect sentiment quad prediction (ASQP) has become a popular task in the field of aspect-level sentiment analysis.Previous work utilizes a predefined template to paraphrase the original sentence into a structure target sequence, which can be easily decoded as quadruplets of the form (aspect category, aspect term, opinion term, sentiment polarity).The template involves the four elements in a fixed order.However, we observe that this solution contradicts with the order-free property of the ASQP task, since there is no need to fix the template order as long as the quadruplet is extracted correctly.Inspired by the observation, we study the effects of template orders and find that some orders help the generative model achieve better performance.It is hypothesized that different orders provide various views of the quadruplet.Therefore, we propose a simple but effective method to identify the most proper orders, and further combine multiple proper templates as data augmentation to improve the ASQP task.Specifically, we use the pre-trained language model to select the orders with minimal entropy.By fine-tuning the pre-trained language model with these template orders, our approach improves the performance of quad prediction, and outperforms state-ofthe-art methods significantly in low-resource settings 1 .
Mengting Hu 0002, Yike Wu 0002, Hang Gao 0003, Yinhao Bai, Shiwan Zhao
EMNLP1
2022 Automated search space and search strategy selection for AutoML
Chao Xue 0003, Mengting Hu 0002, Xueqi Huang, Chun-Guang Li
Pattern Recognit.2
2021 Multi-Label Few-Shot Learning for Aspect Category Detection
abstract
Mengting Hu, Shiwan Zhao, Honglei Guo, Chao Xue, Hang Gao, Tiegang Gao, Renhong Cheng, Zhong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Mengting Hu 0002, Shiwan Zhao, Chao Xue 0003, Hang Gao 0003, Tiegang Gao, Renhong Cheng, Zhong Su
ACL/IJCNLP (1)1
2021 Knowledge Graph Integrated Graph Neural Networks for Chinese Medical Text Classification
abstract
Text classification is a well-developed task in natural language processing. In the medical area, this task is still difficult since the discriminative ability requires domain-specific knowledge, such as high specialization, terminology, and structured relationships. In this work, we leverage a large Chinese medical knowledge graph (KG) to derive the above knowledge and propose graph neural networks (GNN) to make the best use of such knowledge. Particularly, for a medical text, we build two graphs: text-graph, which is based on the occurrences of contextualized words; text-specific knowledge graph, retrieved from KG in light of the common terms between the text and KG. The above two graphs are bridged by the common terms and then merged into a joint one. We propose GNN to learn this graph. As such, our model builds interactions between adjacent nodes, and meanwhile, the medical knowledge can be propagated from KG to text. To enhance the node representations and improve knowledge interaction, we introduce general prior knowledge to the text-graph and domain-specific prior knowledge to the text-specific knowledge graph. We conduct extensive experiments on three medical datasets. The experimental results show that our model outperforms strong baseline methods significantly.
Ge Lan, Ye Li 0016, Mengting Hu 0002
BIBM3
2021 Efficient Mind-Map Generation via Sequence-to-Graph and Reinforced Graph Refinement
abstract
A mind-map is a diagram that represents the central concept and key ideas in a hierarchical way.Converting plain text into a mindmap will reveal its key semantic structure and be easier to understand.Given a document, the existing automatic mind-map generation method extracts the relationships of every sentence pair to generate the directed semantic graph for this document.The computation complexity increases exponentially with the length of the document.Moreover, it is difficult to capture the overall semantics.To deal with the above challenges, we propose an efficient mind-map generation network that converts a document into a graph via sequenceto-graph.To guarantee a meaningful mindmap, we design a graph refinement module to adjust the relation graph in a reinforcement learning manner.Extensive experimental results demonstrate that the proposed approach is more effective and efficient than the existing methods.The inference time is reduced by thousands of times compared with the existing methods.The case studies verify that the generated mind-maps better reveal the underlying semantic structures of the document.
Mengting Hu 0002, Shiwan Zhao, Hang Gao 0003, Zhong Su
EMNLP (1)1
2019 Learning to Detect Opinion Snippet for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) is to predict the sentiment polarity towards a particular aspect in a sentence.Recently, this task has been widely addressed by the neural attention mechanism, which computes attention weights to softly select words for generating aspect-specific sentence representations.The attention is expected to concentrate on opinion words for accurate sentiment prediction.However, attention is prone to be distracted by noisy or misleading words, or opinion words from other aspects.In this paper, we propose an alternative hard-selection approach, which determines the start and end positions of the opinion snippet, and selects the words between these two positions for sentiment prediction.Specifically, we learn deep associations between the sentence and aspect, and the long-term dependencies within the sentence by leveraging the pre-trained BERT model.We further detect the opinion snippet by selfcritical reinforcement learning.Especially, experimental results demonstrate the effectiveness of our method and prove that our hardselection approach outperforms soft-selection approaches when handling multi-aspect sentences.
Mengting Hu 0002, Shiwan Zhao, Renhong Cheng, Zhong Su
CoNLL1
2019 Domain-Invariant Feature Distillation for Cross-Domain Sentiment Classification
abstract
Mengting Hu, Yike Wu, Shiwan Zhao, Honglei Guo, Renhong Cheng, Zhong Su. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mengting Hu 0002, Yike Wu 0002, Shiwan Zhao, Renhong Cheng, Zhong Su
EMNLP/IJCNLP (1)1
2019 CAN: Constrained Attention Networks for Multi-Aspect Sentiment Analysis
abstract
Mengting Hu, Shiwan Zhao, Li Zhang, Keke Cai, Zhong Su, Renhong Cheng, Xiaowei Shen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mengting Hu 0002, Shiwan Zhao, Li Zhang 0007, Keke Cai, Zhong Su, Renhong Cheng
EMNLP/IJCNLP (1)1
2019 Robust detection of median filtering based on combined features of difference image
Hang Gao 0003, Mengting Hu 0002, Tiegang Gao, Renhong Cheng
Signal Process. Image Commun.2