Kai He 0001

dblp:12/5913-1 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0003-2639-1532ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 5 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Recovering Coherent Affective Patterns: Addressing Modality Missing in Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) seeks to decode human emotions by integrating heterogeneous modalities. However, real-world scenarios often involve missing or misaligned data due to sensor failures or transmission errors, leading to disrupted temporal dynamics and degraded cross-modal correlations. To address these challenges, we propose RECAP (REcovery of Coherent Affective Patterns), a robust two-stage framework to restore temporal and structural emotional integrity under modality incompleteness. The first stage employs a causality-aware adversarial generator for multi-granularity temporal reconstruction, complemented by a contrastive mutual information factorization module that disentangles shared and modality-specific semantics. The second stage introduces a mutual information-guided attention fusion mechanism with a ranking-based objective, enabling adaptive integration of complementary signals for refined prediction. Extensive experiments on MOSI, MOSEI, and SIMS under various missing-modality conditions demonstrate that RECAP consistently outperforms state-of-the-art methods. Notably, it improves ACC-7 on MOSI by 2.71 percentage points and F1 on SIMS by 6.38 percentage points. These results verify the performance of RECAP in terms of capturing fine-grained emotional cues and robustness.
Huiting Huang, Tieliang Gong, Kai He 0001, Wen Wen 0013, Weizhan Zhang, Mengling Feng
AAAI3
2026 MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt Optimization
abstract
Large language models (LLMs) typically operate in a question-answering paradigm, where the quality of the input prompt critically affects the response. Automated Prompt Optimization (APO) aims to overcome the cognitive biases of manually crafted prompts and explore a broader prompt design space. However, existing APO methods often suffer from rigid template structures and inefficient exploration in the prompt space. To this end, we propose a Multi-Agent Adaptive Reasoning with Socratic guidance framework (MARS) for APO. MARS consists of five complementary agents and formulates the optimization process as a Partially Observable Markov Decision Process (POMDP), enabling adaptive prompt refinement through explicit state modeling and interactive feedback. Specifically, a Planner agent generates flexible optimization trajectories, a Teacher-Critic-Student triad engages in Socratic-style dialogue to iteratively optimize the prompt based on pseudo-gradient signals in the text space, and a Target agent executes the prompt in downstream tasks to provide performance feedback. MARS integrates reasoning, feedback, and state transition into a unified hidden-state evolution process, improving both the effectiveness and interpretability of optimization. Extensive experiments on multiple datasets demonstrate that MARS outperforms existing APO methods in terms of optimization performance, search efficiency, and interpretability.
Jian Zhang 0087, Zhangqi Wang, Kangda Cheng, Kai He 0001, Qika Lin, Jun Liu 0002, Erik Cambria
AAAI5
2026 MedForge: Interpretable Medical Deepfake Detection via Forgery-aware Reasoning
abstract
Text-guided image editors can now manipulate authentic medical scans with high fidelity, enabling lesion implantation/removal that threatens clinical trust and safety. Existing defenses are inadequate for healthcare. Medical detectors are largely black-box, while MLLM-based explainers are typically post-hoc, lack medical expertise, and may hallucinate evidence on ambiguous cases. We present MedForge, a data-and-method solution for pre-hoc, evidence-grounded medical forgery detection. We introduce MedForge-90K, a large-scale benchmark of realistic lesion edits across 19 pathologies with expert-guided reasoning supervision via doctor inspection guidelines and gold edit locations. Building on it, MedForge-Reasoner performs localize-then-analyze reasoning, predicting suspicious regions before producing a verdict, and is further aligned with Forgery-aware GSPO to strengthen grounding and reduce hallucinations. Experiments demonstrate state-of-the-art detection accuracy and trustworthy, expert-aligned explanations.
Kai He 0001, Qingyuan Lei, Bin Pu, Jian Zhang 0087, Yuling Xu, Mengling Feng
ACL (1)2
2026 Multi-granularity semantic extraction and multi-task fusion for Chinese medical entity normalization
Kai He 0001, Rui Mao 0010, Mengling Feng
Expert Syst. Appl.2
2026 MTLQ-ViT: Multi-granularity Tail-enhanced Logarithmic Quantization for Vision Transformers
Yan Kang 0003, Shouhao Xu, Qika Lin, Kai He 0001, Zhuangzhuang Chen, Bin Pu
Pattern Recognit.4
2026 Global Entity Relationship Enhancement Network for Multimodal Sarcasm Detection
abstract
Sarcasm functions as a distinct mode of communication, intended to convey a meaning contrary to its literal interpretation. With the rapid proliferation of social networks, various manifestations of multimodal sarcasm have become widespread. Consequently, there is a growing emphasis on discerning sarcasm conveyed through multimodal data. Existing research often provides only a superficial interpretation of images, lacking a thorough exploration of the contextual nuances embedded within, particularly in understanding the scene depicted in the image. In this paper, we aim to delve deeper into the information embedded within images. Specifically, we begin by extracting entity relationships from images to capture the contextual information they convey. Additionally, we utilize image text recognition to extract textual information from the images. After conducting a comprehensive analysis of the image information, we establish consistency modeling between the image content and text using external knowledge. Finally, we employ a graph neural network to process the constructed cross-modal graph and make predictions regarding sarcasm. Extensive experiments validate the state-of-the-art performance of our model on publicly available multimodal Twitter datasets.
Xiaobao Wang, Meng Ge, Lingshan Li, Di Jin 0001, Kai He 0001, Erik Cambria
IEEE Trans. Affect. Comput.5
2026 External Retrievals or Internal Priors? From RAG to Epitome-Augmented Generation by Fuzzy Selection
abstract
Retrieval-Augmented Generation (RAG) offers a promising solution to the limitations of static knowledge and hallucinations in Large Language Models (LLMs). While prior research has introduced numerous enhancements to RAG systems, a significant challenge remains under-explored: the potential conflict between external retrievals and LLMs' internal priors, which can undermine the quality of generated outputs. To tackle this issue, we present theEpitome-AugmentedGeneration (EAG) framework, which strategically aligns queries, external retrievals, and internal priors to produce high-quality LLM generations by selecting fuzzy inputs. EAG employs two novel lightweight modules, Criticism and Distillation, allowing traditional RAGs to be upgraded to EAGs without the need for specialized training data. Extensive experiments on five datasets across general and medical domains, including both open-ended and closed-ended tasks, validate the effectiveness of EAG. Our framework achieves substantial F1 score improvements: 7.03%, 23.35%, and 21.58% over baseline RAGs in medical QA tasks, 11.80% in law domain, 7.95% in finance domain, and 4.13% and 5.16% in general domain. Beyond performance gains, our study delves into the interplay between LLMs' internal priors and external retrievals, uncovering key principles that govern generation quality and providing valuable insights for future retrieval-augmented frameworks.
Kai He 0001, Jiaxing Xu, Qika Lin, Zeyu Gao 0001, Jialun Wu, Mengling Feng
IEEE Trans. Fuzzy Syst.1
2026 Multi-Atlas Brain Network Classification Through Consistency Distillation and Complementary Information Fusion
abstract
Brain network analysis plays a crucial role in identifying distinctive patterns associated with neurological disorders. Functional magnetic resonance imaging (fMRI) enables the construction of brain networks by analyzing correlations in blood-oxygen-level-dependent (BOLD) signals across different brain regions, known as regions of interest (ROIs). These networks are typically constructed using atlases that parcellate the brain based on various hypotheses of functional and anatomical divisions. However, there is no standard atlas for brain network classification, leading to limitations in detecting abnormalities in disorders. Recent methods leveraging multiple atlases fail to ensure consistency across atlases and lack effective ROI-level information exchange, limiting their efficacy. To address these challenges, we propose the Atlas-Integrated Distillation and Fusion network (AIDFusion), a novel framework designed to enhance brain network classification using fMRI data. AIDFusion introduces a disentangle Transformer to filter out inconsistent atlas-specific information and distill meaningful cross-atlas connections. Additionally, it enforces subject- and population-level consistency constraints to improve cross-atlas coherence. To further enhance feature integration, AIDFusion incorporates an inter-atlas message-passing mechanism that facilitates the fusion of complementary information across brain regions. We evaluate AIDFusion on four resting-state fMRI datasets encompassing different neurological disorders. Experimental results demonstrate its superior classification performance and computational efficiency compared to state-of-the-art methods. Furthermore, a case study highlights AIDFusion's ability to extract interpretable patterns that align with established neuroscience findings, reinforcing its potential as a robust tool for multi-atlas brain network analysis.
Jiaxing Xu, Mengcheng Lan, Xia Dong, Kai He 0001, Wayne Zhang 0001, Qingtian Bian, Yiping Ke
IEEE J. Biomed. Health Informatics4
2026 BrainPrompt+: Multi-Level Brain Prompt Learning for Knowledge-Guided Neurological Disorder Identification
abstract
Accurate identification of neurological disorders such as Alzheimer's disease (AD), Parkinson's disease (PD), and Autism Spectrum Disorder (ASD) is challenging due to subtle early-stage symptoms and heterogeneous brain dynamics. Resting-state functional MRI (rs-fMRI) enables the construction of functional brain networks, where Graph Neural Networks (GNNs) have shown promise for disease classification. However, existing GNN-based methods face three key limitations: correlation-based graph construction introduces noise and negative edges; domain knowledge about brain regions is ignored; and demographic or clinical metadata are fused through simplistic encodings. To overcome these limitations, we propose BrainPrompt+, a knowledge-guided framework that integrates Large Language Models (LLMs) with multi-level natural language prompts. Five types of prompts are introduced: spectral (frequency-domain BOLD features), spatial (inter-ROI connectivity), ROI (anatomical and functional knowledge), disease (progression stages), and subject (demographic context). These prompts are encoded by a frozen LLM and incorporated into a GNN pipeline, unifying imaging, clinical, and external knowledge in a semantically enriched and interpretable manner. Experiments on three rs-fMRI datasets show that BrainPrompt+ consistently outperforms state-of-the-art baselines, achieving accuracy gains of up to 8.93%. Biomarker analysis further demonstrates that the highlighted ROIs align with established neuroscience findings, confirming the interpretability of the model. BrainPrompt+ thus establishes a flexible and generalizable paradigm for knowledge-guided brain network analysis. The source code is available at https://github.com/AngusMonroe/BrainPromptPlus.
Jiaxing Xu, Kai He 0001, Wei Li 0231, Mengcheng Lan, Yue Xun, Qika Lin, Peifan Ran, Yiping Ke, Mengling Feng
IEEE Trans. Medical Imaging2
2025 Crab: A Novel Configurable Role-Playing LLM with Assessing Benchmark
abstract
Kai He, Yucheng Huang, Wenqing Wang, Delong Ran, Dongming Sheng, Junxuan Huang, Qika Lin, Jiaxing Xu, Wenqiang Liu, Mengling Feng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kai He 0001, Delong Ran, Dongming Sheng, Junxuan Huang, Qika Lin, Jiaxing Xu, Mengling Feng
ACL (1)1
2025 Self-supervised Quantized Representation for Seamlessly Integrating Knowledge Graphs with Large Language Models
abstract
Qika Lin, Tianzhe Zhao, Kai He, Zhen Peng, Fangzhi Xu, Ling Huang, Jingying Ma, Mengling Feng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Qika Lin, Tianzhe Zhao, Kai He 0001, Zhen Peng 0005, Fangzhi Xu, Ling Huang 0003, Jingying Ma, Mengling Feng
ACL (1)3
2025 Neuro-Symbolic AI in Healthcare
abstract
Medical AI has achieved strong predictive performance, yet most systems remain limited by shallow reasoning, poor transparency, and weak generalisation in safety-critical settings. Neurosymbolic AI offers a path beyond these constraints by combining neural models' ability to learn from complex clinical data with the explicit structure, logic, and domain knowledge of symbolic methods. This article examines how neurosymbolic approaches can address core challenges in healthcare AI through five key areas: hybrid reasoning that unifies learning and logic; symbol grounding that links internal representations to clinically meaningful concepts; clinical interpretability that exposes reasoning steps; human-integrated decision-making that keeps clinicians in control; and knowledge-driven diagnosis that incorporates guidelines, ontologies, and causal understanding. Together, these elements outline how neurosymbolic AI can support systems that are not only accurate but also transparent, clinically aligned, and robust in complex or data-sparse scenarios. Advancing this paradigm will require collaboration across AI research, clinical practice, and knowledge engineering, as well as governance mechanisms that ensure fairness and accountability. Neurosymbolic AI thus represents a promising direction for building trustworthy, knowledge-rich intelligence in healthcare.
Jialun Wu, Xin Mei, Kai He 0001, Jiaxing Xu, Qika Lin, Zeyu Gao 0001, Rui Mao 0010
BIBM3
2025 DivScore: Zero-Shot Detection of LLM-Generated Text in Specialized Domains
abstract
Detecting LLM-generated text in specialized and high-stakes domains like medicine and law is crucial for combating misinformation and ensuring authenticity.However, current zeroshot detectors, while effective on general text, often fail when applied to specialized content due to domain shift.We provide a theoretical analysis showing this failure is fundamentally linked to the KL divergence between human, detector, and source text distributions.To address this, we propose DivScore, a zero-shot detection framework using normalized entropybased scoring and domain knowledge distillation to robustly identify LLM-generated text in specialized domains.We also release a domainspecific benchmark for LLM-generated text detection in the medical and legal domains.Experiments on our benchmark show that Di-vScore consistently outperforms state-of-theart detectors, with 14.4% higher AUROC and 64.0% higher recall (0.1% false positive rate threshold).In adversarial settings, DivScore demonstrates superior robustness to other baselines, achieving on average 22.8% advantage in AUROC and 29.5% in recall.Code and data are publicly available 1 .
Kai He 0001, Yunxiao Zhu, Mengling Feng
EMNLP2
2025 BrainPrompt: Multi-level Brain Prompt Enhancement for Neurological Condition Identification
Jiaxing Xu, Kai He 0001, Wei Li 0231, Mengcheng Lan, Xia Dong, Yiping Ke, Mengling Feng
MICCAI (12)2
2025 GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images
abstract
While recent multimodal large language models (MLLMs) have advanced automated ECG interpretation, they still face two key limitations: (1) insufficient multimodal synergy between ECG time series and ECG images, and (2) limited explainability in linking diagnoses to granular waveform evidence. We introduce GEM, the first MLLM unifying ECG time series, 12-lead ECG images and text for grounded and clinician-aligned ECG interpretation. GEM enables feature-grounded analysis, evidence-driven reasoning, and a clinician-like diagnostic process through three core innovations: a dual-encoder framework extracting complementary time series and image features, cross-modal alignment for effective multimodal understanding, and knowledge-guided instruction data generation for generating high-granularity grounding data (ECG-Grounding) linking diagnoses to measurable parameters ($e.g.$, QRS/PR Intervals). Additionally, we propose the Grounded ECG Understanding task, a clinically motivated benchmark designed to comprehensively assess the MLLM's capability in grounded ECG understanding. Experimental results on both existing and our proposed benchmarks show GEM significantly improves predictive performance (CSN $7.4\%$ $\uparrow$), explainability ($22.7\%$ $\uparrow$), and grounding ($25.3\%$ $\uparrow$), making it a promising approach for real-world clinical applications. Codes, model, and data are available at https://github.com/lanxiang1017/GEM.
Xiang Lan 0004, Feng Wu 0001, Kai He 0001, Qinghao Zhao, Shenda Hong, Mengling Feng
NeurIPS3
2025 Cross-Modal Knowledge Diffusion-Based Generation for Difference-Aware Medical VQA
abstract
Multimodal medical applications have garnered considerable attention due to their potential to offer comprehensive and robust support for medical assistance. Specifically, within this domain, difference-aware medical Visual Question Answering (VQA) has emerged as a topic of increasing interest that enables the recognition of changes in physical conditions over time when compared to previous states and provides customized suggestions accordingly. However, it is challenging because samples usually exhibit characteristics of complexity, diversity, and inherent noise. Besides, there is a need for multimodal knowledge understanding of the medical domain. The difference-aware setting requiring image comparison further intensifies these situations. To this end, we propose a cross-Modal knowlEdge diffusioN-baseD gEneration netwoRk (MENDER), where the diffusion mechanism with multi-step denoising and knowledge injection from global to local level are employed to tackle the aforementioned challenges, respectively. The diffusion process is to gradually generate answers with the sequence input of questions, random noises for the answer masks and virtual vision prompts of images. The strategy of answer nosing and knowledge cascading is specifically tailored for this task and is implemented during forward and reverse diffusion processes. Moreover, the visual and structure knowledge injection are proposed to learn virtual vision prompts to guide the diffusion process, where the former is realized using a pre-trained medical image-text network and the latter is modeled with spatial and semantic graph structures processed by the heterogeneous graph Transformer models. Experiment results demonstrate the effectiveness of MENDER for difference-aware medical VQA. Furthermore, it also exhibits notable performance in the low-resource setting and conventional medical VQA tasks.
Qika Lin, Kai He 0001, Yifan Zhu 0001, Fangzhi Xu, Erik Cambria, Mengling Feng
IEEE Trans. Image Process.2
2024 Contrasformer: A Brain Network Contrastive Transformer for Neurodegenerative Condition Identification
abstract
Understanding neurological disorder is a fundamental problem in neuroscience, which often requires the analysis of brain networks derived from functional magnetic resonance imaging (fMRI) data. Despite the prevalence of Graph Neural Networks (GNNs) and Graph Transformers in various domains, applying them to brain networks faces challenges. Specifically, the datasets are severely impacted by the noises caused by distribution shifts across sub- populations and the neglect of node identities, both obstruct the identification of disease-specific patterns. To tackle these challenges, we propose Contrasformer, a novel contrastive brain network Transformer. It generates a prior-knowledge-enhanced contrast graph to address the distribution shifts across sub-populations by a two-stream attention mechanism. A cross attention with identity embedding highlights the identity of nodes, and three auxiliary losses ensure group consistency. Evaluated on 4 functional brain network datasets over 4 different diseases, Contrasformer outperforms the state-of-the-art methods for brain networks by achieving up to 10.8% improvement in accuracy, which demonstrates its efficacy in neurological disorder identification. Case studies illustrate its interpretability, especially in the context of neuroscience. This paper provides a solution for analyzing brain networks, offering valuable insights into neurological disorders. Our code is available at https://github.com/AngusMonroe/Contrasformer.
Jiaxing Xu, Kai He 0001, Mengcheng Lan, Qingtian Bian, Wei Li 0231, Tieying Li, Yiping Ke, Miao Qiao
CIKM2
2024 PROMISE: A pre-trained knowledge-infused multimodal representation learning framework for medication recommendation
Jialun Wu, Xinyao Yu 0004, Kai He 0001, Zeyu Gao 0001, Tieliang Gong
Inf. Process. Manag.3
2024 Integrating K+ Entities Into Coreference Resolution on Biomedical Texts
abstract
Biomedical Coreference Resolution focuses on identifying the coreferences in biomedical texts, which normally consists of two parts: (i) mention detection to identify textual representation of biological entities and (ii) finding their coreference links. Recently, a popular approach to enhance the task is to embed knowledge base into deep neural networks. However, the way in which these methods integrate knowledge leads to the shortcoming that such knowledge may play a larger role in mention detection than coreference resolution. Specifically, they tend to integrate knowledge prior to mention detection, as part of the embeddings. Besides, they primarily focus on mention-dependent knowledge (KBase), i.e., knowledge entities directly related to mentions, while ignores the correlated knowledge (K+) between mentions in the mention-pair. For mentions with significant differences in word form, this may limit their ability to extract potential correlations between those mentions. Thus, this paper develops a novel model to integrate both KBase and K+ entities and achieves the state-of-the-art performance on BioNLP and CRAFT-CR datasets. Empirical studies on mention detection with different length reveals the effectiveness of the KBase entities. The evaluation on cross-sentence and match/mismatch coreference further demonstrate the superiority of the K+ entities in extracting background potential correlation between mentions.
Yufei Li 0002, Xiaoyong Ma, Penghzhen Cheng, Kai He 0001, Tieliang Gong, Chen Li 0011
IEEE ACM Trans. Comput. Biol. Bioinform.5
2024 Template-Free Prompting for Few-Shot Named Entity Recognition via Semantic-Enhanced Contrastive Learning
abstract
Prompt tuning has achieved great success in various sentence-level classification tasks by using elaborated label word mappings and prompt templates. However, for solving token-level classification tasks, e.g., named entity recognition (NER), previous research, which utilizes N-gram traversal for prompting all spans with all possible entity types, is time-consuming. To this end, we propose a novel prompt-based contrastive learning method for few-shot NER without template construction and label word mappings. First, we leverage external knowledge to initialize semantic anchors for each entity type. These anchors are simply appended with input sentence embeddings as template-free prompts (TFPs). Then, the prompts and sentence embeddings are in-context optimized with our proposed semantic-enhanced contrastive loss. Our proposed loss function enables contrastive learning in few-shot scenarios without requiring a significant number of negative samples. Moreover, it effectively addresses the issue of conventional contrastive learning, where negative instances with similar semantics are erroneously pushed apart in natural language processing (NLP)-related tasks. We examine our method in label extension (LE), domain-adaption (DA), and low-resource generalization evaluation tasks with six public datasets and different settings, achieving state-of-the-art (SOTA) results in most cases.
Kai He 0001, Rui Mao 0010, Tieliang Gong, Chen Li 0011, Erik Cambria
IEEE Trans. Neural Networks Learn. Syst.1
2023 Virtual prompt pre-training for prototype-based few-shot relation extraction
Kai He 0001, Rui Mao 0010, Tieliang Gong, Chen Li 0011, Erik Cambria
Expert Syst. Appl.1
2023 Meta-Based Self-Training and Re-Weighting for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) means to identify fine-grained aspects, opinions, and sentiment polarities. Recent ABSA research focuses on utilizing multi-task learning (MTL) to achieve less computational costs and better performance. However, there are certain limits in MTL-based ABSA. For example, unbalanced labels and sub-task learning difficulties may result in the biases that some labels and sub-tasks are overfitting, while the others are underfitting. To address these issues, inspired by neuro-symbolic learning systems, we propose a meta-based self-training method with a meta-weighter (MSM). We believe that a generalizable model can be achieved by appropriate symbolic representation selection (in-domain knowledge) and effective learning control (regulation) in a neural system. Thus, MSM trains a teacher model to generate in-domain knowledge (e.g., unlabeled data selection and pseudo-label generation), where the generated pseudo-labels are used by a student model for supervised learning. Then, the meta-weighter of MSM is jointly trained with the student model to provide each instance with sub-task-specific weights to coordinate their convergence rates, balancing class labels, and alleviating noise impacts introduced from self-training. The following experiments indicate that MSM can utilize 50% labeled data to achieve comparable results to state-of-arts models in ABSA and outperform them with all labeled data.
Kai He 0001, Rui Mao 0010, Tieliang Gong, Chen Li 0011, Erik Cambria
IEEE Trans. Affect. Comput.1
2023 The Biases of Pre-Trained Language Models: An Empirical Study on Prompt-Based Sentiment Analysis and Emotion Detection
abstract
Thanks to the breakthrough of large-scale pre-trained language model (PLM) technology, prompt-based classification tasks, e.g., sentiment analysis and emotion detection, have raised increasing attention. Such tasks are formalized as masked language prediction tasks which are in line with the pre-training objects of most language models. Thus, one can use a PLM to infer the masked words in a downstream task, then obtaining label predictions with manually defined label-word mapping templates. Prompt-based affective computing takes the advantages of both neural network modeling and explainable symbolic representations. However, there still remain many unclear issues related to the mechanisms of PLMs and prompt-based classification. We conduct a systematic empirical study on prompt-based sentiment analysis and emotion detection to study the biases of PLMs towards affective computing. We find that PLMs are biased in sentiment analysis and emotion detection tasks with respect to the number of label classes, emotional label-word selections, prompt templates and positions, and the word forms of emotion lexicons.
Rui Mao 0010, Qian Liu 0012, Kai He 0001, Wei Li 0076, Erik Cambria
IEEE Trans. Affect. Comput.3
2022 Knowledge Enhanced Coreference Resolution via Gated Attention
abstract
Coreference resolution aims at linking all mentions that refer to the same entity, which are widely adopted in many biomedical and bioinformatics tasks, such as biomedical knowledge graph construction and metabolic pathway integration. Many recent studies focus on improving neural model structures. However, we argue that a practical method that integrates commonsense knowledge can further improve coreference resolution performance, because commonsense delivers extra prior knowledge for reasoning and can enhance related representations, rather than naive mention-context occurrence modeling. In this work, we propose an effective method to integrate external commonsense knowledge into a neural coreference resolution model. Specially, a gated attention mechanism is employed in our method to leverage commonsense according to different contexts. By using ConceptNet as the knowledge base in three span-ranking backbone models, the models can yield significant performance gains on used datasets. We also achieve improvements in tasks of long-term mention detection and cross-sentence coreferences after incorporating knowledge.
Kai He 0001, Yufei Li 0002, Tieliang Gong, Chen Li 0011, Jialun Wu
BIBM1
2022 Uncertainty-guided Mutual Consistency Training for Semi-supervised Biomedical Relation Extraction
abstract
Biomedical relation extraction seeks to automatically extract biomedical relations from biomedical text, which plays an important role in biomedical studies. However, constructing high-quality biomedical annotation data is not only time-consuming but also requires a high level of knowledge in the biomedical field. To alleviate this problem, Semi-supervised Biomedical Relation Extraction aims to extract relation facts from the limited labeled data and the more readily available unlabeled samples. Existing works can be roughly categorized as self-training methods and self-ensembling methods. The former aims to generate pseudo labels, which may lead to the gradual drift problem. The latter aims to encourage the output of one model to be consistent with the other model, where the acquisition of the model is tedious. To alleviate these issues, we propose a novel Uncertainty-Guided Mutual Consistency Training framework(UG-MCT) for semi-supervised Biomedical relation extraction. Specifically, our framework consists of two models with the same structure, which differ only when updating their weights, and then an intersecting pseudo-label mechanism is designed to convert the prediction discrepancies of the two models into mutual consistency training loss, thus promoting the consistency of model predictions. In addition, we utilize uncertainty as guided information to assist the model in focusing on the confident pseudo labels and mitigate the noise of inaccurate pseudo labeling during training. Thus, our model is very simple and efficient while mitigating the noise introduced by pseudo-labels. UG-MCT is evaluated on multiple datasets in different settings and the experimental results demonstrate that our method is highly effective in semi-supervised biomedical relation extraction compared to the state-of-the-art.
Chang Jia, Kai He 0001, Jialun Wu, Tieliang Gong, Chen Li 0011
BIBM4
2022 COPNER: Contrastive Learning with Prompt Guiding for Few-shot Named Entity Recognition
abstract
Distance metric learning has become a popular solution for few-shot Named Entity Recognition (NER). The typical setup aims to learn a similarity metric for measuring the semantic similarity between test samples and referents, where each referent represents an entity class. The effect of this setup may, however, be compromised for two reasons. First, there is typically a limited optimization exerted on the representations of entity tokens after initing by pre-trained language models. Second, the referents may be far from representing corresponding entity classes due to the label scarcity in the few-shot setting. To address these challenges, we propose a novel approach named COntrastive learning with Prompt guiding for few-shot NER (COPNER). We introduce a novel prompt composed of class-specific words to COPNER to serve as 1) supervision signals for conducting contrastive learning to optimize token representations; 2) metric referents for distance-metric inference on test samples. Experimental results demonstrate that COPNER outperforms state-of-the-art models with a significant margin in most cases. Moreover, COPNER shows great potential in the zero-shot setting.
Kai He 0001, Xianli Zhang, Tieliang Gong, Rui Mao 0010, Chen Li 0011
COLING2
2022 JCBIE: a joint continual learning neural network for biomedical information extraction
abstract
Extracting knowledge from heterogeneous data sources is fundamental for the construction of structured biomedical knowledge graphs (BKGs), where entities and relations are represented as nodes and edges in the graphs, respectively. Previous biomedical knowledge extraction methods simply considered limited entity types and relations by using a task-specific training set, which is insufficient for large-scale BKGs development and downstream task applications in different scenarios. To alleviate this issue, we propose a joint continual learning biomedical information extraction (JCBIE) network to extract entities and relations from different biomedical information datasets. By empirically studying different joint learning and continual learning strategies, the proposed JCBIE can learn and expand different types of entities and relations from different datasets. JCBIE uses two separated encoders in joint-feature extraction, hence can effectively avoid the feature confusion problem comparing with using one hard-parameter sharing encoder. Specifically, it allows us to adopt entity augmented inputs to establish the interaction between named entity recognition and relation extraction. Finally, a novel evaluation mechanism is proposed for measuring cross-corpus generalization errors, which was ignored by traditional evaluation methods. Our empirical studies show that JCBIE achieves promising performance when continual learning strategy is adopted with multiple corpora.
Kai He 0001, Rui Mao 0010, Tieliang Gong, Erik Cambria, Chen Li 0011
BMC Bioinform.1
2021 BERT-Based Meta-Learning Approach with Looking Back for Sentiment Analysis of Literary Book Reviews
Hui Bao, Kai He 0001, Xuemeng Yin, Xuanyu Li, Xinrui Bao, Haichuan Zhang 0001, Jialun Wu, Zeyu Gao 0001
NLPCC (2)2
2021 Knowledge enhanced LSTM for coreference resolution on biomedical texts
abstract
MOTIVATION: Bio-entity Coreference Resolution focuses on identifying the coreferential links in biomedical texts, which is crucial to complete bio-events' attributes and interconnect events into bio-networks. Previously, as one of the most powerful tools, deep neural network-based general domain systems are applied to the biomedical domain with domain-specific information integration. However, such methods may raise much noise due to its insufficiency of combining context and complex domain-specific information. RESULTS: In this article, we explore how to leverage the external knowledge base in a fine-grained way to better resolve coreference by introducing a knowledge-enhanced Long Short Term Memory network (LSTM), which is more flexible to encode the knowledge information inside the LSTM. Moreover, we further propose a knowledge attention module to extract informative knowledge effectively based on contexts. The experimental results on the BioNLP and CRAFT datasets achieve state-of-the-art performance, with a gain of 7.5 F1 on BioNLP and 10.6 F1 on CRAFT. Additional experiments also demonstrate superior performance on the cross-sentence coreferences. AVAILABILITY AND IMPLEMENTATION: The source code will be made available at https://github.com/zxy951005/KB-CR upon publication. Data is avaliable at http://2011.bionlp-st.org/ and https://github.com/UCDenver-ccp/CRAFT/releases/tag/v3.1.3. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yufei Li 0002, Xiaoyong Ma, Pengzhen Cheng, Kai He 0001, Chen Li 0011
Bioinform.5