EDBT 2026 Demo / reviewers in the wild / expert
Yang Xiang 0003
dblp:50/2192-3
· DBLP profile ↗
64ranked-venue papers
7as first author
47since 2021 · last 2026
0000-0003-1395-6805ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 4 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 8 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality EvaluationabstractAs Large Language Models (LLMs) are increasingly popularized in the multilingual world, ensuring hallucination-free factuality becomes markedly crucial. However, existing benchmarks for evaluating the reliability of Multimodal Large Language Models (MLLMs) predominantly focus on textual or visual modalities with a primary emphasis on English, which creates a gap in evaluation when processing multilingual input, especially in speech. To bridge this gap, we propose a novel Cross-lingual and Cross-modal Factuality benchmark (CCFQA). Specifically, the CCFQA benchmark contains parallel speech-text factual questions across 8 languages, designed to systematically evaluate MLLMs' cross-lingual and cross-modal factuality capabilities. Our experimental results demonstrate that current MLLMs still face substantial challenges on the CCFQA benchmark. Furthermore, we propose a few-shot transfer learning strategy that effectively transfers the Question Answering (QA) capabilities of LLMs in English to multilingual Spoken Question Answering (SQA) tasks, achieving competitive performance with GPT-4o-mini-Audio using just 5-shot training. We release CCFQA as a foundational research resource to promote the development of MLLMs with more robust and reliable speech understanding capabilities. Yexing Du, Youcheng Pan, Bo Yang 0006, Ming Liu 0004, Yang Xiang 0003 |
AAAI | 8 |
| 2026 | The Visual Prism: Refracting Images into Parallel Multilingual Descriptions with Structured Visual GuidanceabstractParallel corpora, as the foundation of machine translation, remain crucial even in the era of large language models (LLMs) for pre-training and fine-tuning. However, annotating parallel corpora is extremely costly, as it requires annotators to be proficient in multiple languages. To reduce this cost, prior work has explored image-pivoted corpus synthesis, generating multilingual captions for the same image as pseudo-parallel data. Unfortunately, these pseudo corpora suffer from the serious issue of multilingual focus divergence, i.e., the model attending to distinct aspects of the image when generating captions in different languages. To address this problem, we propose a method called PRISMS (Parallel Refracting ImageS into Multilingual descriptions with Structured visual guidance), which leverages semantic graphs as structured visual guidance to unify the focus of multilingual captions. To ensure adherence to this guidance, we introduce two key techniques: supervised fine-tuning using self-generated instructional data, and reinforcement learning with a reward signal based on semantic graph consistency. Experimental results on five languages show that our PRISMS significantly improves the image-pivot parallel corpora synthesis, enabling LLMs to achieve translation performance comparable to that of models trained on manually annotated corpora. Chengpeng Fu, Yichong Huang, Wenshuai Huo, Baohang Li, Yang Xiang 0003, Ting Liu 0001 |
AAAI | 6 |
| 2026 | When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient ReasoningabstractLarge reasoning models (LRMs) have achieved remarkable performance in complex reasoning tasks, driven by their powerful inference-time scaling capability.However, LRMs often suffer from overthinking, which results in substantial computational redundancy and significantly reduces efficiency.Early-exit methods aim to mitigate this issue by terminating reasoning once sufficient evidence has been generated, yet existing approaches mostly rely on handcrafted or empirical indicators that are unreliable and impractical.In this work, we introduce Dynamic Thought Sufficiency in Reasoning (DTSR), a novel framework for efficient reasoning that enables the model to dynamically assess the sufficiency of its chain-of-thought (CoT) and determine the optimal point for early exit.Inspired by human metacognition, DTSR operates in two stages: (1) Reflection Signal Monitoring, which identifies reflection signals as potential cues for early exit, and (2) Thought Sufficiency Check, which evaluates whether the current CoT is sufficient to derive the final answer.Experimental results on the Qwen3 models show that DTSR reduces reasoning length by 28.9%-34.9%with minimal performance loss, effectively mitigating overthinking.We further discuss overconfidence in LRMs and self-evaluation paradigms, providing valuable insights for early-exit reasoning. Yang Xiang 0003, Yixin Ji, Ruotao Xu, Zheming Yang, Juntao Li 0005, Min Zhang 0005 |
ACL (1) | 1 |
| 2025 | Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum LearningabstractYexing Du, Youcheng Pan, Ziyang Ma, Bo Yang, Yifan Yang, Keqi Deng, Xie Chen, Yang Xiang, Ming Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yexing Du, Youcheng Pan, Ziyang Ma 0001, Bo Yang 0006, Yifan Yang 0005, Keqi Deng, Xie Chen 0001, Yang Xiang 0003, Ming Liu 0004, Bing Qin 0001 |
ACL (1) | 8 |
| 2025 | CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention InterventionabstractLarge Vision-Language Models (LVLMs) have demonstrated impressive multimodal abilities but remain prone to multilingual object hallucination, with a higher likelihood of generating responses inconsistent with the visual input when utilizing queries in non-English languages compared to English. Most existing approaches to address these rely on pretraining or fine-tuning, which are resource-intensive. In this paper, inspired by observing the disparities in cross-modal attention patterns across languages, we propose Cross-Lingual Attention Intervention for Mitigating multilingual object hallucination (CLAIM) in LVLMs, a novel near training-free method by aligning attention patterns. CLAIM first identifies language-specific cross-modal attention heads, then estimates language shift vectors from English to the target language, and finally intervenes in the attention outputs during inference to facilitate cross-lingual visual perception capability alignment. Extensive experiments demonstrate that CLAIM achieves an average improvement of 13.56% (up to 30% in Spanish) on the POPE and 21.75% on the hallucination subsets of the MME benchmark across various languages. Further analysis reveals that multilingual attention divergence is most prominent in intermediate layers, highlighting their critical role in multilingual scenarios. Zekai Ye, Libo Qin 0001, Yichong Huang, Baohang Li, Kui Jiang, Yang Xiang 0003, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
ACL (1) | 8 |
| 2025 | Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and ReasoningabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable performance across diverse tasks.Despite great success, recent studies show that LVLMs encounter substantial limitations when engaging with visual graphs.To study the reason behind these limitations, we propose VGCURE, a comprehensive benchmark covering 22 tasks for examining the fundamental graph understanding and reasoning capacities of LVLMs.Extensive evaluations conducted on 14 LVLMs reveal that LVLMs are weak in basic graph understanding and reasoning tasks, particularly those concerning relational or structurally complex information.Based on this observation, we propose a structure-aware fine-tuning framework to enhance LVLMs with structure learning abilities through three self-supervised learning tasks.Experiments validate the effectiveness of our method in improving LVLMs' performance on fundamental and downstream graph learning tasks, as well as enhancing their robustness against complex visual graphs. Yingjie Zhu, Xuefeng Bai 0001, Kehai Chen, Yang Xiang 0003, Jun Yu 0002, Min Zhang 0005 |
ACL (1) | 4 |
| 2025 | Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and EleganceabstractLarge language models (LLMs) have shown remarkable performance in general translation tasks.However, the increasing demand for high-quality translations that are not only adequate but also fluent and elegant.To assess the extent to which current LLMs can meet these demands, we introduce a suitable benchmark (PoetMT) for translating classical Chinese poetry into English.This task requires not only adequacy in translating culturally and historically significant content but also a strict adherence to linguistic fluency and poetic elegance.Our study reveals that existing LLMs fall short of this task.To address these issues, we propose RAT, a Retrieval-Augmented machine Translation method that enhances the translation process by incorporating knowledge related to classical poetry.Additionally, we propose an automatic evaluation metric based on GPT-4, which better assesses translation quality in terms of adequacy, fluency, and elegance, overcoming the limitations of traditional metrics.Our dataset and code will be made available 1 . Andong Chen 0001, Lianzhang Lou, Kehai Chen, Xuefeng Bai 0001, Yang Xiang 0003, Muyun Yang, Tiejun Zhao, Min Zhang 0005 |
EMNLP | 5 |
| 2025 | From Long to Lean: Performance-aware and Adaptive Chain-of-Thought Compression via Multi-round RefinementabstractChain-of-Thought (CoT) reasoning improves performance on complex tasks but introduces significant inference latency due to its verbosity.In this work, we propose Multiround Adaptive Chain-of-Thought Compression (MACC), a framework that leverages the token elasticity phenomenon-where overly small token budgets may paradoxically increase output length-to progressively compress CoTs via multiround refinement.This adaptive strategy allows MACC to dynamically determine the optimal compression depth for each input.Our method achieves an average accuracy improvement of 5.6% over state-of-the-art baselines, while also reducing CoT length by an average of 47 tokens and significantly lowering latency.Furthermore, we show that test-time performance-accuracy and token length-can be reliably predicted using interpretable features like perplexity and compression rate on training set.Evaluated across different models, our method enables efficient model selection and forecasting without repeated fine-tuning, demonstrating that CoT compression is both effective and predictable.Our code will be released in https://github.com/Leon221220/ MACC. Jianzhi Yan, Youcheng Pan, Zike Yuan, Yang Xiang 0003, Buzhou Tang |
EMNLP | 6 |
| 2025 | Beware of Calibration Data for Pruning Large Language ModelsabstractAs large language models (LLMs) are widely applied across various fields, model
compression has become increasingly crucial for reducing costs and improving
inference efficiency. Post-training pruning is a promising method that does not
require resource-intensive iterative training and only needs a small amount of
calibration data to assess the importance of parameters. Recent research has enhanced post-training pruning from different aspects but few of them systematically
explore the effects of calibration data, and it is unclear if there exist better calibration data construction strategies. We fill this blank and surprisingly observe that
calibration data is also crucial to post-training pruning, especially for high sparsity. Through controlled experiments on important influence factors of calibration
data, including the pruning settings, the amount of data, and its similarity with
pre-training data, we observe that a small size of data is adequate, and more similar data to its pre-training stage can yield better performance. As pre-training data
is usually inaccessible for advanced LLMs, we further provide a self-generating
calibration data synthesis strategy to construct feasible calibration data. Experimental results on recent strong open-source LLMs (e.g., DCLM, and LLaMA-3)
show that the proposed strategy can enhance the performance of strong pruning
methods (e.g., Wanda, DSnoT, OWL) by a large margin (up to 2.68%). Yixin Ji, Yang Xiang 0003, Juntao Li 0005, Qingrong Xia, Ping Li 0016, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ICLR | 2 |
| 2025 | ReFEdit: Rehearsal-Free Lifelong Knowledge Editing for Large Language ModelsabstractKnowledge editing has emerged as a promising strategy for updating obsolete or inaccurate knowledge embedded within large language models (LLMs) without costly fine-tuning. The widely adopted locating-then-editing paradigm first locates parameters responsible for knowledge storage and then modifies them to integrate updated knowledge. However, in lifelong knowledge editing scenarios, catastrophic forgetting poses a significant challenge. Existing methods often rely on rehearsal-based techniques, such as storing a feature covariance matrix of previously preserved knowledge to constrain errors, thus raising efficiency and privacy issues. To address this, we introduce ReFEdit, a Rehearsal-Free Lifelong Knowledge Editing framework that enforces an orthogonality restriction on parameter modifications. By aligning the update direction orthogonally to both the latest and initial parameters, ReFEdit minimizes the interference between sequentially edited knowledge while mitigating the impact on previously preserved knowledge, thereby effectively addressing catastrophic forgetting. Extensive evaluations on multiple representative LLMs, including LLaMA3, GPT-J, and GPT2-XL, demonstrate that ReFEdit significantly outperforms most existing rehearsal-based knowledge editing methods while eliminating the need for the rehearsal phase, marking a substantial advancement toward more reliable and flexible lifelong knowledge editing. Our code is available at: https://github.com/Cedric-Mo/ReFEdit Xianjie Mo, Youcheng Pan, Yongshuai Hou, Ping Luo 0001, Yang Xiang 0003 |
ICME | 5 |
| 2025 | A Survey on the Feedback Mechanism of LLM-based AI AgentsabstractLarge language models (LLMs) are increasingly being adopted to develop general-purpose AI agents. However, it remains challenging for these LLM-based AI agents to efficiently learn from feedback and iteratively optimize their strategies. To address this challenge, tremendous efforts have been dedicated to designing diverse feedback mechanisms for LLM-based AI agents. To provide a comprehensive overview of this rapidly evolving field, this paper presents a systematic review of these studies, offering a holistic perspective on the feedback mechanisms in LLM-based AI agents. We begin by discussing the construction of LLM-based AI agents, introducing a generalized framework that encapsulates much of the existing work. Next, we delve into the exploration of feedback mechanisms, categorizing them into four distinct types: internal feedback, external feedback, multi-agent feedback, and human feedback. Additionally, we provide an overview of evaluation protocols and benchmarks specifically tailored for LLM-based AI agents. Finally, we highlight the significant challenges and identify potential directions for future studies. The relevant papers are summarized and will be consistently updated at https://github.com/kevinson7515/Agents-Feedback-Mechanisms. Xuefeng Bai 0001, Kehai Chen, Xinyang Chen 0001, Xiucheng Li, Yang Xiang 0003, Jin Liu 0012, Hong-Dong Li, Yaowei Wang 0001, Liqiang Nie, Min Zhang 0005 |
IJCAI | 6 |
| 2025 | Exploring the Translation Mechanism of Large Language ModelsabstractWhile large language models (LLMs) demonstrate remarkable success in multilingual translation, their internal core translation mechanisms, even at the fundamental word level, remain insufficiently understood.
To address this critical gap, this work introduces a systematic framework for interpreting the mechanism behind LLM translation from the perspective of computational components.
This paper first proposes subspace-intervened path patching for precise, fine-grained causal analysis, enabling the detection of components crucial to translation tasks and subsequently characterizing their behavioral patterns in human-interpretable terms.
Comprehensive experiments reveal that translation is predominantly driven by a sparse subset of components: specialized attention heads serve critical roles in extracting source language, translation indicators, and positional features, which are then integrated and processed by specific multi-layer perceptrons (MLPs) into intermediary English-centric latent representations before ultimately yielding the final translation.
The significance of these findings is underscored by the empirical demonstration that targeted fine-tuning a minimal parameter subset (<5%) enhances translation performance while preserving general capabilities. This result further indicates that these crucial components generalize effectively to sentence-level translation and are instrumental in elucidating more intricate translation tasks. Kehai Chen, Xuefeng Bai 0001, Xiucheng Li, Yang Xiang 0003, Min Zhang 0005 |
NeurIPS | 5 |
| 2025 | TF-Attack: Transferable and fast adversarial attacks on large language models
Kehai Chen, Lemao Liu, Xuefeng Bai 0001, Yang Xiang 0003, Min Zhang 0005 |
Knowl. Based Syst. | 6 |
| 2025 | Hypercomplex Graph Neural Network: Towards Deep Intersection of Multi-Modal Brain NetworksabstractThe multi-modal neuroimage study has provided insights into understanding the heteromodal relationships between brain network organization and behavioral phenotypes. Integrating data from various modalities facilitates the characterization of the interplay among anatomical, functional, and physiological brain alterations or developments. Graph Neural Networks (GNNs) have recently become popular in analyzing and fusing multi-modal, graph-structured brain networks. However, effectively learning complementary representations from other modalities remains a significant challenge due to the sophisticated and heterogeneous inter-modal dependencies. Furthermore, most existing studies often focus on specific modalities (e.g., only fMRI and DTI), which limits their scalability to other types of brain networks. To overcome these limitations, we propose a HyperComplex Graph Neural Network (HC-GNN) that models multi-modal networks as hypercomplex tensor graphs. In our approach, HC-GNN is conceptualized as a dynamic spatial graph, where the attentively learned inter-modal associations are represented as the adjacency matrix. HC-GNN leverages hypercomplex operations for inter-modal intersections through cross-embedding and cross-aggregation, enriching the deep coupling of multi-modal representations. We conduct a statistical analysis on the saliency maps to associate disease biomarkers. Extensive experiments on three datasets demonstrate the superior classification performance of our method and its strong scalability to various types of modalities. Our work presents a powerful paradigm for the study of multi-modal brain networks. Yanwu Yang 0001, Chenfei Ye, Guoqing Cai, Kunru Song, Yang Xiang 0003, Heather Ting Ma |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | ZO-AdaMU Optimizer: Adapting Perturbation by the Momentum and Uncertainty in Zeroth-Order OptimizationabstractLowering the memory requirement in full-parameter training on large models has become a hot research area. MeZO fine-tunes the large language models (LLMs) by just forward passes in a zeroth-order SGD optimizer (ZO-SGD), demonstrating excellent performance with the same GPU memory usage as inference. However, the simulated perturbation stochastic approximation for gradient estimate in MeZO leads to severe oscillations and incurs a substantial time overhead. Moreover, without momentum regularization, MeZO shows severe over-fitting problems. Lastly, the perturbation-irrelevant momentum on ZO-SGD does not improve the convergence rate. This study proposes ZO-AdaMU to resolve the above problems by adapting the simulated perturbation with momentum in its stochastic approximation. Unlike existing adaptive momentum methods, we relocate momentum on simulated perturbation in stochastic gradient approximation. Our convergence analysis and experiments prove this is a better way to improve convergence stability and rate in ZO-SGD. Extensive experiments demonstrate that ZO-AdaMU yields better generalization for LLMs fine-tuning across various NLP tasks than MeZO and its momentum variants. Shuoran Jiang, Qingcai Chen, Youcheng Pan, Yang Xiang 0003, Yukang Lin, Xiangping Wu 0001, Chuanyi Liu, Xiaobao Song |
AAAI | 4 |
| 2024 | Linguistic Rule Induction Improves Adversarial and OOD Robustness in Large Language ModelsabstractEnsuring robustness is especially important when AI is deployed in responsible or safety-critical environments. ChatGPT can perform brilliantly in both adversarial and out-of-distribution (OOD) robustness, while other popular large language models (LLMs), like LLaMA-2, ERNIE and ChatGLM, do not perform satisfactorily in this regard. Therefore, it is valuable to study what efforts play essential roles in ChatGPT, and how to transfer these efforts to other LLMs. This paper experimentally finds that linguistic rule induction is the foundation for identifying the cause-effect relationships in LLMs. For LLMs, accurately processing the cause-effect relationships improves its adversarial and OOD robustness. Furthermore, we explore a low-cost way for aligning LLMs with linguistic rules. Specifically, we constructed a linguistic rule instruction dataset to fine-tune LLMs. To further energize LLMs for reasoning step-by-step with the linguistic rule, we construct the task-relevant LingR-based chain-of-thoughts. Experiments showed that LingR-induced LLaMA-13B achieves comparable or better results with GPT-3.5 and GPT-4 on various adversarial and OOD robustness evaluations. Shuoran Jiang, Qingcai Chen, Yang Xiang 0003, Youcheng Pan, Yukang Lin |
LREC/COLING | 3 |
| 2024 | Mitigating Knowledge Conflicts in Data-to-Text Generation via the Internalization of Fact ExtractionabstractLarge Language Models (LLMs) have made remarkable advancements in Natural Language Generation. Nonetheless, LLMs are prone to encountering knowledge conflicts, scenarios where the generated statement exhibits either twisted facts that contradict the source or unverified facts that lack substantiation from the source. We posit that LLMs should possess the capability to differentiate between facts grounded in evidence and those derived from inference. A promising avenue for mitigating knowledge conflicts involves prompting LLMs to externalize the multi-step reasoning during text generation within the chain-of-thought paradigm. Nevertheless, such externalization of reasoning has only shown benefits for sufficiently large LLMs (e.g., those exceeding 100B parameters). Despite attempts to distill this ability into smaller LLMs, the reliance on large and inefficient teacher models remains a necessity. Thus, we introduce the Internalization of Fact Extraction (IFE), which optimizes instruction-tuning, empowering smaller LLMs to internalize the fact extraction process during text generation within the auxiliary learning paradigm. Consequently, models can directly generate statements where knowledge conflicts are mitigated. Specifically, we design an auxiliary fact extraction task built upon the primary text generation task, requiring models to identify evidential facts and filter out inferential facts after text generation. The auxiliary training dataset for the auxiliary fact extraction task was constructed using word-stemming and string-matching techniques. By leveraging auxiliary learning, the models’ ability to internalize the fact extraction process emerges, thereby enhancing its performance in the primary text generation task. Extensive experiments and analyses are conducted on three benchmark data-to-text generation datasets (i.e., FeTaQA, DialogSum, and SAMSum), covering sub-tasks including table-based text generation and dialogue-based text generation. Experimental results show that our IFE instruction-tuned models outperform the vanilla instruction-tuned baselines in terms of both factuality and quality without introducing any inference latency, demonstrating the superiority of IFE instruction-tuning in mitigating knowledge conflicts. Xianjie Mo, Yang Xiang 0003, Youcheng Pan, Yongshuai Hou, Ping Luo 0001 |
IJCNN | 2 |
| 2024 | Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationabstractLarge language models (LLMs) exhibit complementary strengths in various tasks, motivating the research of LLM ensembling.
However, existing work focuses on training an extra reward model or fusion model to select or combine all candidate answers, posing a great challenge to the generalization on unseen data distributions.
Besides, prior methods use textual responses as communication media, ignoring the valuable information in the internal representations.
In this work, we propose a training-free ensemble framework \textsc{DeePEn}, fusing the informative probability distributions yielded by different LLMs at each decoding step.
Unfortunately, the vocabulary discrepancy between heterogeneous LLMs directly makes averaging the distributions unfeasible due to the token misalignment.
To address this challenge, \textsc{DeePEn} maps the probability distribution of each model from its own probability space to a universal \textit{relative space} based on the relative representation theory, and performs aggregation.
Next, we devise a search-based inverse transformation to transform the aggregated result back to the probability space of one of the ensembling LLMs (main model), in order to determine the next token.
We conduct extensive experiments on ensembles of different number of LLMs, ensembles of LLMs with different architectures, and ensembles between the LLM and the specialist model.
Experimental results show that (i) \textsc{DeePEn} achieves consistent improvements across six benchmarks covering subject examination, reasoning, and knowledge, (ii) a well-performing specialist model can benefit from a less effective LLM through distribution fusion, and (iii) \textsc{DeePEn} has complementary strengths with other ensemble methods such as voting. Yichong Huang, Baohang Li, Yang Xiang 0003, Hui Wang 0030, Ting Liu 0001, Bing Qin 0001 |
NeurIPS | 4 |
| 2024 | TSOANet: Time-Sensitive Orthogonal Attention Network for medical event prediction
Hao Chen 0186, Yang Xiang 0003, Shengye Lu, Buzhou Tang |
Artif. Intell. Medicine | 3 |
| 2024 | MGCoT: Multi-Grained Contextual Transformer for table-based text generation
Xianjie Mo, Yang Xiang 0003, Youcheng Pan, Yongshuai Hou, Ping Luo 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Confounder balancing in adversarial domain adaptation for pre-trained large models fine-tuningabstractThe excellent generalization, contextual learning, and emergence abilities in the pre-trained large models (PLMs) handle specific tasks without direct training data, making them the better foundation models in the adversarial domain adaptation (ADA) methods to transfer knowledge learned from the source domain to target domains. However, existing ADA methods fail to account for the confounder properly, which is the root cause of the source data distribution that differs from the target domains. This study proposes a confounder balancing method in adversarial domain adaptation for PLMs fine-tuning (CadaFT), which includes a PLM as the foundation model for a feature extractor, a domain classifier and a confounder classifier, and they are jointly trained with an adversarial loss. This loss is designed to improve the domain-invariant representation learning by diluting the discrimination in the domain classifier. At the same time, the adversarial loss also balances the confounder distribution among source and unmeasured domains in training. Compared to newest ADA methods, CadaFT can correctly identify confounders in domain-invariant features, thereby eliminating the confounder biases in the extracted features from PLMs. The confounder classifier in CadaFT is designed as a plug-and-play and can be applied in the confounder measurable, unmeasurable, or partially measurable environments. Empirical results on natural language processing and computer vision downstream tasks show that CadaFT outperforms the newest GPT-4, LLaMA2, ViT and ADA methods. Shuoran Jiang, Qingcai Chen, Yang Xiang 0003, Youcheng Pan, Xiangping Wu 0001, Yukang Lin |
Neural Networks | 3 |
| 2024 | BaSFormer: A Balanced Sparsity Regularized Attention Network for TransformerabstractAttention networks often make decisions relying solely on a few pieces of tokens, even if those reliances are not truly indicative of the underlying meaning or intention of the full context. This can lead to over-fitting in transformers and hinder their ability to generalize. Attention regularization and sparsity-based methods have been used to overcome this issue. However, these methods cannot guarantee that all tokens have sufficient receptive fields for global information inference. Thus, the impact of individual biases cannot be effectively reduced. As a result, the generalization of these approaches improved slightly from the training data to new data. To address these limitations, we propose a balanced sparsity (BaS) regularized attention network on top of the transformers, called BaSFormer. BaS regularization introduces the K-regular graph constraint on self-attention connections, which replaces SoftMax with SparseMax in the attention transformation. In BaS-regularized self-attention, SparseMax assigns zero attention scores to low-scoring connections, highlighting influential and meaningful contexts. The K-regular graph constraint ensures that all tokens have an equal-sized receptive field to aggregate information, which facilitates the involvement of global tokens in the feature update of each layer and reduces the impact of individual biases. Given that there is no continuous loss can be used for the K-regular graph regularization, we propose an exponential extremum loss with an augmented Lagrangian function. The experimental results showed that BaSFormer improved the effectiveness of debiasing compared to that of the newest LLMs, such as the GPT-3.5, GPT-4 and LLaMA. In addition, BaSFormer achieves new state-of-the-art (SOTA) results in text generation tasks. Interestingly, this work also shows that BaSFormer can learn hierarchical linguistic dependencies in gradient attributions, which improves interpretability and adversarial robustness. Shuoran Jiang, Qingcai Chen, Yang Xiang 0003, Youcheng Pan, Xiangping Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Learning to Improve Out-of-Distribution Generalization via Self-Adaptive Language MaskingabstractAlthough the pre-trained Transformers learned general linguistic knowledge from large-scale corpus, they still over-fit on the lexical biases when fine-tuning on specific datasets. This problem limits the generalizability of pre-trained models, particularly when learning over out-of-distribution (OOD) data. To address this issue, this paper proposes a self-adaptive language masking (AdaLMask) paradigm to fine-tune the pre-trained Transformers. AdaLMask obviates lexical biases by eliminating the dependence on semantically inessential words. Specifically, AdaLMask learns a Gumbel-Softmax distribution to determine the desired masking positions, and the distribution parameters are optimized via a representation-invariant (RInv) objective to ensure the masked positions are semantically lossless. Four natural language processing tasks are chosen to evaluate the effectiveness of the proposed method on the robustness of lexical biases and OOD generalization. All empirical results demonstrate that the AdaLMask paradigm substantially improves the OOD generalization of pre-trained Transformers. Shuoran Jiang, Youcheng Pan, Qingcai Chen, Yang Xiang 0003, Xiangping Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | BioPRO: Context-Infused Prompt Learning for Biomedical Entity LinkingabstractRecent research tends to address the biomedical entity linking problem in a unified framework solely based on surface form matching between mentions and entities. Specifically, these methods focus on addressing thevarietychallenge of the heterogeneous naming of biomedical concepts. Yet, theambiguitychallenge that the same word under different contexts can be used to refer to distinct concepts is usually ignored. To address this challenge, we propose BioPRO, a two-stage entity linking algorithm to enhance the biomedical entity representations based on context-infused prompt learning. The first stage includes a coarse-grained retrieval from a representation space defined by a bi-encoder that independently embeds the mention and entity's surface forms. Unlike previous one-model-fits-all systems, each candidate is then re-ranked with a fine-grained encoder based on prompt-tuning that sufficiently stimulates knowledge in contextual information of mentions and entities. Furthermore, the trained fine-grained encoder can be utilized to generate deep representations of bio-entities and boost candidate retrieval in the first stage. Extensive experiments show that our model achieves promising performance improvements compared with several state-of-the-art (SOTA) techniques on 4 biomedical corpora. We also observe by cases that the proposed context-infused prompt-tuning strategy is effective in solving both thevarietyandambiguitychallenges in the linking task. Tiantian Zhu 0002, Yang Qin 0001, Ming Feng, Qingcai Chen, Baotian Hu, Yang Xiang 0003 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2024 | Mapping Multi-Modal Brain Connectome for Brain Disorder Diagnosis via Cross-Modal Mutual LearningabstractRecently, the study of multi-modal brain connectome has recorded a tremendous increase and facilitated the diagnosis of brain disorders. In this paradigm, functional and structural networks, e.g., functional and structural connectivity derived from fMRI and DTI, are in some manner interacted but are not necessarily linearly related. Accordingly, there remains a great challenge to leverage complementary information for brain connectome analysis. Recently, Graph Convolutional Networks (GNN) have been widely applied to the fusion of multi-modal brain connectome. However, most existing GNN methods fail to couple inter-modal relationships. In this regard, we propose a Cross-modal Graph Neural Network (Cross-GNN) that captures inter-modal dependencies through dynamic graph learning and mutual learning. Specifically, the inter-modal representations are attentively coupled into a compositional space for reasoning inter-modal dependencies. Additionally, we investigate mutual learning in explicit and implicit ways: (1) Cross-modal representations are obtained by cross-embedding explicitly based on the inter-modal correspondence matrix. (2) We propose a cross-modal distillation method to implicitly regularize latent representations with cross-modal semantic contexts. We carry out statistical analysis on the attentively learned correspondence matrices to evaluate inter-modal relationships for associating disease biomarkers. Our extensive experiments on three datasets demonstrate the superiority of our proposed method for disease diagnosis with promising prediction performance and multi-modal connectome biomarker location. Yanwu Yang 0001, Chenfei Ye, Xutao Guo, Yang Xiang 0003, Heather Ting Ma |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Controllable Contrastive Generation for Multilingual Biomedical Entity LinkingabstractMultilingual biomedical entity linking (MBEL) aims to map language-specific mentions in the biomedical text to standardized concepts in a multilingual knowledge base (KB) such as Unified Medical Language System (UMLS).In this paper, we propose Con2GEN, a prompt-based controllable contrastive generation framework for MBEL, which summarizes multidimensional information of the UMLS concept mentioned in biomedical text into a natural sentence following a predefined template.Instead of tackling the MBEL problem with a discriminative classifier, we formulate it as a sequence-tosequence generation task, which better exploits the shared dependencies between source mentions and target entities.Moreover, Con2GEN matches against UMLS concepts in as many languages and types as possible, hence facilitating cross-information disambiguation.Extensive experiments show that our model achieves promising performance improvements compared with several state-of-the-art techniques on the XL-BEL and the Mantra GSC datasets spanning 12 typologically diverse languages. Tiantian Zhu 0002, Yang Qin 0001, Qingcai Chen, Xin Mu, Changlong Yu, Yang Xiang 0003 |
EMNLP | 6 |
| 2023 | Tensor-based Complex-valued Graph Neural Network for Dynamic Coupling Multimodal brain NetworksabstractThe multi-modal neuroimage study has dramatically facilitated disease diagnosis. Tensor-based methods are commonly used to represent multi-modal data as multi-dimensional arrays and usually implement matrix decomposition. These methods can be seen as a linear algebraic way for the lossy compression of an array. However, involved lossy operations might have a negative impact on performance, and overlook underlying important complementary information between modalities. This study proposes a Tensor-based Complex-valued Graph Neural Network (TC-GNN) to model multimodal neuroimages as complex-valued tensor graphs by investigating underlying complementary associations and cross-modality message aggregation. Experiments on two real-world datasets demonstrate our method’s consistent improvements and superiority over other baseline models in multi-modal brain disease analysis. Yanwu Yang 0001, Guoqing Cai, Chenfei Ye, Yang Xiang 0003, Heather Ting Ma |
ICASSP | 4 |
| 2023 | TMMDA: A New Token Mixup Multimodal Data Augmentation for Multimodal Sentiment AnalysisabstractExisting methods for Multimodal Sentiment Analysis (MSA) mainly focus on integrating multimodal data effectively on limited multimodal data. Learning more informative multimodal representation often relies on large-scale labeled datasets, which are difficult and unrealistic to obtain. To learn informative multimodal representation on limited labeled datasets as more as possible, we proposed TMMDA for MSA, a new Token Mixup Multimodal Data Augmentation, which first generates new virtual modalities from the mixed token-level representation of raw modalities, and then enhances the representation of raw modalities by utilizing the representation of the generated virtual modalities. To preserve semantics during virtual modality generation, we propose a novel cross-modal token mixup strategy based on the generative adversarial network. Extensive experiments on two benchmark datasets, i.e., CMU-MOSI and CMU-MOSEI, verify the superiority of our model compared with several state-of-the-art baselines. The code is available at https://github.com/xiaobaicaihhh/TMMDA. Xianbing Zhao, Sicen Liu, Xuan Zang, Yang Xiang 0003, Buzhou Tang |
WWW | 5 |
| 2023 | Machine learning-based donor permission extraction from informed consent documentsabstractBACKGROUND: With more clinical trials are offering optional participation in the collection of bio-specimens for biobanking comes the increasing complexity of requirements of informed consent forms. The aim of this study is to develop an automatic natural language processing (NLP) tool to annotate informed consent documents to promote biorepository data regulation, sharing, and decision support. We collected informed consent documents from several publicly available sources, then manually annotated them, covering sentences containing permission information about the sharing of either bio-specimens or donor data, or conducting genetic research or future research using bio-specimens or donor data. RESULTS: We evaluated a variety of machine learning algorithms including random forest (RF) and support vector machine (SVM) for the automatic identification of these sentences. 120 informed consent documents containing 29,204 sentences were annotated, of which 1250 sentences (4.28%) provide answers to a permission question. A support vector machine (SVM) model achieved a F-1 score of 0.95 on classifying the sentences when using a gold standard, which is a prefiltered corpus containing all relevant sentences. CONCLUSIONS: This study provides the feasibility of using machine learning tools to classify permission-related sentences in informed consent documents. Madhuri Sankaranarayanapillai, Jingcheng Du, Yang Xiang 0003, Frank J. Manion, Marcelline R. Harris, Cooper Stansbury, Huy Anh Pham, Cui Tao |
BMC Bioinform. | 4 |
| 2023 | Fine-grained biomedical knowledge negation detection via contrastive learning
Tiantian Zhu 0002, Yang Xiang 0003, Qingcai Chen, Yang Qin 0001, Baotian Hu, Wentai Zhang 0003 |
Knowl. Based Syst. | 2 |
| 2023 | CReg-KD: Model refinement via confidence regularized knowledge distillation for brain imaging
Yanwu Yang 0001, Xutao Guo, Chenfei Ye, Yang Xiang 0003, Heather Ting Ma |
Medical Image Anal. | 4 |
| 2023 | SHAPE: A Sample-Adaptive Hierarchical Prediction Network for Medication RecommendationabstractEffectively medication recommendation with complex multimorbidity conditions is a critical yet challenging task in healthcare. Most existing works predicted medications based on longitudinal records, which assumed the encoding format of intra-visit medical events are serialized and information transmitted patterns of learning longitudinal sequence data are stable. However, the following conditions may have been ignored: 1) A more compact encoder for intra-relationship in the intra-visit medical event is urgent; 2) Strategies for learning accurate representations of the variable longitudinal sequences of patients are different. In this article, we proposed a novel Sample-adaptive Hierarchical medicAtion Prediction nEtwork, termed SHAPE, to tackle the above challenges in the medication recommendation task. Specifically, we design a compact intra-visit set encoder to encode the relationship in the medical event for obtaining visit-level representation and then develop an inter-visit longitudinal encoder to learn the patient-level longitudinal representation efficiently. To endow the model with the capability of modeling the variable visit length, we introduce a soft curriculum learning method to assign the difficulty of each sample automatically by the visit length. Extensive experiments on a benchmark dataset verify the superiority of our model compared with several state-of-the-art baselines. Sicen Liu, Xiaolong Wang 0001, Jingcheng Du, Yongshuai Hou, Xianbing Zhao, Hui Wang 0030, Yang Xiang 0003, Buzhou Tang |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | Multimodal Data Matters: Language Model Pre-Training Over Structured and Unstructured Electronic Health RecordsabstractAs two important textual modalities in electronic health records (EHR), both structured data (clinical codes) and unstructured data (clinical narratives) have recently been increasingly applied to the healthcare domain. Most existing EHR-oriented studies, however, either focus on a particular modality or integrate data from different modalities in a straightforward manner, which usually treats structured and unstructured data as two independent sources of information about patient admission and ignore the intrinsic interactions between them. In fact, the two modalities are documented during the same encounter where structured data inform the documentation of unstructured data and vice versa. In this paper, we proposed a Medical Multimodal Pre-trained Language Model, named MedM-PLM, to learn enhanced EHR representations over structured and unstructured data and explore the interaction of two modalities. In MedM-PLM, two Transformer-based neural network components are firstly adopted to learn representative characteristics from each modality. A cross-modal module is then introduced to model their interactions. We pre-trained MedM-PLM on the MIMIC-III dataset and verified the effectiveness of the model on three downstream clinical tasks, i.e., medication recommendation, 30-day readmission prediction and ICD coding. Extensive experiments demonstrate the power of MedM-PLM compared with state-of-the-art methods. Further analyses and visualizations show the robustness of our model, which could potentially provide more comprehensive interpretations for clinical decision-making. Sicen Liu, Xiaolong Wang 0001, Yongshuai Hou, Ge Li 0002, Hui Wang 0030, Yang Xiang 0003, Buzhou Tang |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | Combining Transfer Learning with Graph Attention Models for HIV Risk Prediction
Evan Yu, Cui Tao, Jingcheng Du, Degui Zhi, Yang Xiang 0003, Kayo Fujimoto, John A. Schneider |
AMIA | 5 |
| 2022 | Modeling Annotator Variation and Annotator Preference for Multiple Annotations Medical Image SegmentationabstractMedical image segmentation annotation suffers from annotator variation due to the inherent differences in annotators’ expertise and the inherent blurriness of medical images. In practice, using opinions from multiple annotators can effectively reduce the impact of such annotator-related biases. Meanwhile, it is common practice in deep learning to fuse multiple annotations through methods such as majority voting, but these methods ignore the rich information of annotator preferences ingrained in the original multi-annotator annotations. To address this issue, we propose a modeling annotator variation and annotator preference (AVAP) framework for multiple annotations medical image segmentation, which consists of three parts. First, the widely used encoder-decoder backbone network use to extract feature maps of the image. Second, an annotator variation modeling (AVM) module is devised to estimate the annotation variation among multiple annotators by modeling multi-annotations as a multi-class segmentation problem. Third, an annotator preference modeling (APM) module estimate each annotator’s preference-involved segmentation by annotator encoding and dynamic filter learning. The experiment on the RIGA benchmark with multiple annotations shows that our AVAP framework outperforms a range of state-of-the-art (SOTA) multiple annotations segmentation methods. Further, we are the first to introduce dynamic filter learning into the annotator preference modeling. Xutao Guo, Shang Lu, Yanwu Yang 0001, Chenfei Ye, Yang Xiang 0003, Heather Ting Ma |
BIBM | 6 |
| 2022 | Multi-modal Dynamic Graph Network: Coupling Structural and Functional Connectome for Disease Diagnosis and ClassificationabstractMulti-modal neuroimaging technology has greatly facilitated the diagnosis efficiency and diagnosis accuracy, and provides complementary information in discovering objective disease biomarkers. Conventional deep learning methods, e.g. convolutional neural networks, overlook relationships between nodes and fail to capture topological properties in graphs. Graph neural networks have been proven to be of great importance in modeling brain connectome networks and relating disease-specific patterns. However, most existing graph methods explicitly require known graph structures, which are not available in the sophisticated brain system. Especially in heterogeneous multi-modal brain networks, there exists a great challenge to model interactions among brain regions in consideration of inter-modal dependencies. In this study, we propose a Multimodal Dynamic Graph Convolution Network (MDGCN) for structural and functional brain network learning. Our method benefits from modeling inter-modal representations and relating attentive multi-model associations into dynamic graphs with a compositional correspondence matrix. Moreover, a bilateral graph convolution layer is proposed to aggregate multi-modal representations in terms of multi-modal associations. Extensive experiments on three datasets demonstrate the superiority of our proposed method in terms of disease classification, with the accuracy of 90.4%, 85.9% and 98.3% in predicting Mild Cognitive Impairment, Parkinson’s Disease, and Schizophrenia respectively. Our statistical evaluations on the correspondence matrix exhibit a high correspondence with previous evidence of biomarkers. Yanwu Yang 0001, Xutao Guo, Zhikai Chang, Chenfei Ye, Yang Xiang 0003, Heather Ting Ma |
BIBM | 5 |
| 2022 | Estimating Brain Age with Global and Local DependenciesabstractThe brain age has been proven to be a phenotype of relevance to cognitive performance and brain disease. Achieving accurate brain age prediction is an essential prerequisite for optimizing the predicted brain-age difference as a biomarker. As a comprehensive biological characteristic, the brain age is hard to be exploited accurately with models using feature engineering and local processing such as local convolution and recurrent operations that process one local neighborhood at a time. Instead, Vision Transformers learn global attentive interaction of patch tokens, introducing less inductive bias and modeling long-range dependencies. In terms of this, we proposed a novel network for learning brain age interpreting with global and local dependencies, where the corresponding representations are captured by Successive Permuted Transformer (SPT) and convolution blocks. The SPT brings computation efficiency and locates the 3D spatial information indirectly via continuously encoding 2D slices from different views. Finally, we collect a large cohort of 22645 subjects with ages ranging from 14 to 97 and our network performed the best among a series of deep learning methods, yielding a mean absolute error (MAE) of 2.855 in validation set, and 2.911 in an independent test set. Yanwu Yang 0001, Xutao Guo, Zhikai Chang, Chenfei Ye, Yang Xiang 0003, Haiyan Lv, Heather Ting Ma |
ICIP | 5 |
| 2022 | Enhancing Entity Representations with Prompt Learning for Biomedical Entity LinkingabstractBiomedical entity linking aims to map mentions in biomedical text to standardized concepts or entities in a curated knowledge base (KB) such as Unified Medical Language System (UMLS). The latest research tends to solve this problem in a unified framework solely based on surface form matching between mentions and entities. Specifically, these methods focus on addressing the variety challenge of the heterogeneous naming of biomedical concepts. Yet, the ambiguity challenge that the same word under different contexts may refer to distinct entities is usually ignored. To address this challenge, we propose a two-stage linking algorithm to enhance the entity representations based on prompt learning. The first stage includes a coarser-grained retrieval from a representation space defined by a bi-encoder that independently embeds the mention and entity’s surface forms. Unlike previous one-model-fits-all systems, each candidate is then re-ranked with a finer-grained encoder based on prompt-tuning that utilizes the contextual information. Extensive experiments show that our model achieves promising performance improvements compared with several state-of-the-art techniques on the largest biomedical public dataset MedMentions and the NCBI disease corpus. We also observe by cases that the proposed prompt-tuning strategy is effective in solving both the variety and ambiguity challenges in the linking task. Tiantian Zhu 0002, Yang Qin 0001, Qingcai Chen, Baotian Hu, Yang Xiang 0003 |
IJCAI | 5 |
| 2022 | CATNet: Cross-event attention-based time-aware network for medical event prediction
Sicen Liu, Xiaolong Wang 0001, Yang Xiang 0003, Hui Wang 0030, Buzhou Tang |
Artif. Intell. Medicine | 3 |
| 2022 | Multi-channel fusion LSTM for medical event prediction using EHRs
Sicen Liu, Xiaolong Wang 0001, Yang Xiang 0003, Hui Wang 0030, Buzhou Tang |
J. Biomed. Informatics | 3 |
| 2022 | Biomedical named entity normalization via interaction-based synonym marginalization
Yang Xiang 0003, Hui Wang 0030, Buzhou Tang |
J. Biomed. Informatics | 3 |
| 2022 | Leveraging Multi-source knowledge for Chinese clinical named entity recognition via relational graph convolutional network
Yang Xiang 0003, Ka-Chun Wong, Qingcai Chen, Jun Yan 0010, Buzhou Tang |
J. Biomed. Informatics | 3 |
| 2021 | Kalman prediction-based virtual network experimental platform for smart living
Desheng Wang 0002, Weizhe Zhang, Yang Xiang 0003, Yu-Chu Tian |
Comput. Commun. | 4 |
| 2021 | Extracting postmarketing adverse events from safety reports in the vaccine adverse event reporting system (VAERS) using deep learningabstractOBJECTIVE: Automated analysis of vaccine postmarketing surveillance narrative reports is important to understand the progression of rare but severe vaccine adverse events (AEs). This study implemented and evaluated state-of-the-art deep learning algorithms for named entity recognition to extract nervous system disorder-related events from vaccine safety reports. MATERIALS AND METHODS: We collected Guillain-Barré syndrome (GBS) related influenza vaccine safety reports from the Vaccine Adverse Event Reporting System (VAERS) from 1990 to 2016. VAERS reports were selected and manually annotated with major entities related to nervous system disorders, including, investigation, nervous_AE, other_AE, procedure, social_circumstance, and temporal_expression. A variety of conventional machine learning and deep learning algorithms were then evaluated for the extraction of the above entities. We further pretrained domain-specific BERT (Bidirectional Encoder Representations from Transformers) using VAERS reports (VAERS BERT) and compared its performance with existing models. RESULTS AND CONCLUSIONS: Ninety-one VAERS reports were annotated, resulting in 2512 entities. The corpus was made publicly available to promote community efforts on vaccine AEs identification. Deep learning-based methods (eg, bi-long short-term memory and BERT models) outperformed conventional machine learning-based methods (ie, conditional random fields with extensive features). The BioBERT large model achieved the highest exact match F-1 scores on nervous_AE, procedure, social_circumstance, and temporal_expression; while VAERS BERT large models achieved the highest exact match F-1 scores on investigation and other_AE. An ensemble of these 2 models achieved the highest exact match microaveraged F-1 score at 0.6802 and the second highest lenient match microaveraged F-1 score at 0.8078 among peer models. Jingcheng Du, Yang Xiang 0003, Madhuri Sankaranarayanapillai, Yuqi Si, Huy Anh Pham, Hua Xu 0001, Yong Chen 0016, Cui Tao |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | COVID-19 trial graph: a linked graph for COVID-19 clinical trialsabstractOBJECTIVE: Clinical trials are an essential part of the effort to find safe and effective prevention and treatment for COVID-19. Given the rapid growth of COVID-19 clinical trials, there is an urgent need for a better clinical trial information retrieval tool that supports searching by specifying criteria, including both eligibility criteria and structured trial information. MATERIALS AND METHODS: We built a linked graph for registered COVID-19 clinical trials: the COVID-19 Trial Graph, to facilitate retrieval of clinical trials. Natural language processing tools were leveraged to extract and normalize the clinical trial information from both their eligibility criteria free texts and structured information from ClinicalTrials.gov. We linked the extracted data using the COVID-19 Trial Graph and imported it to a graph database, which supports both querying and visualization. We evaluated trial graph using case queries and graph embedding. RESULTS: The graph currently (as of October 5, 2020) contains 3392 registered COVID-19 clinical trials, with 17 480 nodes and 65 236 relationships. Manual evaluation of case queries found high precision and recall scores on retrieving relevant clinical trials searching from both eligibility criteria and trial-structured information. We observed clustering in clinical trials via graph embedding, which also showed superiority over the baseline (0.870 vs 0.820) in evaluating whether a trial can complete its recruitment successfully. CONCLUSIONS: The COVID-19 Trial Graph is a novel representation of clinical trials that allows diverse search queries and provides a graph-based visualization of COVID-19 clinical trials. High-dimensional vectors mapped by graph embedding for clinical trials would be potentially beneficial for many downstream applications, such as trial end recruitment status prediction and trial similarity comparison. Our methodology also is generalizable to other clinical trials. Jingcheng Du, Prerana Ramesh, Yang Xiang 0003, Xiaoqian Jiang, Cui Tao |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Distantly supervised biomedical relation extraction using piecewise attentive convolutional neural network and reinforcement learningabstractOBJECTIVE: There have been various methods to deal with the erroneous training data in distantly supervised relation extraction (RE), however, their performance is still far from satisfaction. We aimed to deal with the insufficient modeling problem on instance-label correlations for predicting biomedical relations using deep learning and reinforcement learning. MATERIALS AND METHODS: In this study, a new computational model called piecewise attentive convolutional neural network and reinforcement learning (PACNN+RL) was proposed to perform RE on distantly supervised data generated from Unified Medical Language System with MEDLINE abstracts and benchmark datasets. In PACNN+RL, PACNN was introduced to encode semantic information of biomedical text, and the RL method with memory backtracking mechanism was leveraged to alleviate the erroneous data issue. Extensive experiments were conducted on 4 biomedical RE tasks. RESULTS: The proposed PACNN+RL model achieved competitive performance on 8 biomedical corpora, outperforming most baseline systems. Specifically, PACNN+RL outperformed all baseline methods with the F1-score of 0.5592 on the may-prevent dataset, 0.6666 on the may-treat dataset, and 0.3838 on the DDI corpus, 2011. For the protein-protein interaction RE task, we obtained new state-of-the-art performance on 4 out of 5 benchmark datasets. CONCLUSIONS: The performance on many distantly supervised biomedical RE tasks was substantially improved, primarily owing to the denoising effect of the proposed model. It is anticipated that PACNN+RL will become a useful tool for large-scale RE and other downstream tasks to facilitate biomedical knowledge acquisition. We also made the demonstration program and source code publicly available at http://112.74.48.115:9000/. Tiantian Zhu 0002, Yang Qin 0001, Yang Xiang 0003, Baotian Hu, Qingcai Chen, Weihua Peng |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Neural data-to-text generation with dynamic content planning
Kai Chen 0020, Fayuan Li, Baotian Hu, Weihua Peng, Qingcai Chen, Hong Yu 0001, Yang Xiang 0003 |
Knowl. Based Syst. | 7 |
| 2020 | Learning to Generate Diverse Questions from KeywordsabstractDiverse text generation has been emerging as an important topic of natural language generation. Traditional studies on question generation mainly investigate how to generate one question based on a given input (one-to-one). In this paper, we focus on a more complex question generation task, i.e., generating a series of questions for each set of keywords (one-to-many). As an effort towards this, we propose a novel neural generative model, which incorporates context information and control signal to produce multiple diverse questions from a given fixed set of keywords. The control signal is designed to increase the diversity of questions by capturing the diverse patterns from the entire dataset. The context information is used to guarantee the generated questions are highly related to the given keywords. To evaluate the effectiveness of the proposed model, we collect a dataset which contains 62835 questions with respect to 12567 sets of keywords.1To the best of our knowledge, it's the first Chinese financial dataset for diverse question generation. The experimental results show that our model outperforms the competitor methods in terms of BLEU and Distinct. The qualitative evaluation indicates that our model is able to generate diverse and meaningful questions. Youcheng Pan, Baotian Hu, Qingcai Chen, Yang Xiang 0003, Xiaolong Wang 0001 |
ICASSP | 4 |
| 2020 | Exploiting sequence labeling framework to extract document-level relations from biomedical textsabstractBACKGROUND: Both intra- and inter-sentential semantic relations in biomedical texts provide valuable information for biomedical research. However, most existing methods either focus on extracting intra-sentential relations and ignore inter-sentential ones or fail to extract inter-sentential relations accurately and regard the instances containing entity relations as being independent, which neglects the interactions between relations. We propose a novel sequence labeling-based biomedical relation extraction method named Bio-Seq. In the method, sequence labeling framework is extended by multiple specified feature extractors so as to facilitate the feature extractions at different levels, especially at the inter-sentential level. Besides, the sequence labeling framework enables Bio-Seq to take advantage of the interactions between relations, and thus, further improves the precision of document-level relation extraction. RESULTS: Our proposed method obtained an F1-score of 63.5% on BioCreative V chemical disease relation corpus, and an F1-score of 54.4% on inter-sentential relations, which was 10.5% better than the document-level classification baseline. Also, our method achieved an F1-score of 85.1% on n2c2-ADE sub-dataset. CONCLUSION: Sequence labeling method can be successfully used to extract document-level relations, especially for boosting the performance on inter-sentential relation extraction. Our work can facilitate the research on document-level biomedical text mining. Zhiheng Li 0004, Yang Xiang 0003, Ling Luo 0001, Yuanyuan Sun 0002, Hongfei Lin |
BMC Bioinform. | 3 |
| 2020 | Time event ontology (TEO): to support semantic representation and reasoning of complex temporal relations of clinical eventsabstractOBJECTIVE: The goal of this study is to develop a robust Time Event Ontology (TEO), which can formally represent and reason both structured and unstructured temporal information. MATERIALS AND METHODS: Using our previous Clinical Narrative Temporal Relation Ontology 1.0 and 2.0 as a starting point, we redesigned concept primitives (clinical events and temporal expressions) and enriched temporal relations. Specifically, 2 sets of temporal relations (Allen's interval algebra and a novel suite of basic time relations) were used to specify qualitative temporal order relations, and a Temporal Relation Statement was designed to formalize quantitative temporal relations. Moreover, a variety of data properties were defined to represent diversified temporal expressions in clinical narratives. RESULTS: TEO has a rich set of classes and properties (object, data, and annotation). When evaluated with real electronic health record data from the Mayo Clinic, it could faithfully represent more than 95% of the temporal expressions. Its reasoning ability was further demonstrated on a sample drug adverse event report annotated with respect to TEO. The results showed that our Java-based TEO reasoner could answer a set of frequently asked time-related queries, demonstrating that TEO has a strong capability of reasoning complex temporal relations. CONCLUSION: TEO can support flexible temporal relation representation and reasoning. Our next step will be to apply TEO to the natural language processing field to facilitate automated temporal information annotation, extraction, and timeline reasoning to better support time-based clinical decision-making. Fang Li 0011, Jingcheng Du, Yongqun He, Hsing-yi Song, Mohcine Madkour, Guozheng Rao, Yang Xiang 0003, Henry W. Chen, Sijia Liu 0002, Liwei Wang 0010, Hua Xu 0001, Cui Tao |
J. Am. Medical Informatics Assoc. | 7 |
| 2020 | Representation of EHR data for predictive modeling: a comparison between UMLS and other terminologiesabstractOBJECTIVE: Predictive disease modeling using electronic health record data is a growing field. Although clinical data in their raw form can be used directly for predictive modeling, it is a common practice to map data to standard terminologies to facilitate data aggregation and reuse. There is, however, a lack of systematic investigation of how different representations could affect the performance of predictive models, especially in the context of machine learning and deep learning. MATERIALS AND METHODS: We projected the input diagnoses data in the Cerner HealthFacts database to Unified Medical Language System (UMLS) and 5 other terminologies, including CCS, CCSR, ICD-9, ICD-10, and PheWAS, and evaluated the prediction performances of these terminologies on 2 different tasks: the risk prediction of heart failure in diabetes patients and the risk prediction of pancreatic cancer. Two popular models were evaluated: logistic regression and a recurrent neural network. RESULTS: For logistic regression, using UMLS delivered the optimal area under the receiver operating characteristics (AUROC) results in both dengue hemorrhagic fever (81.15%) and pancreatic cancer (80.53%) tasks. For recurrent neural network, UMLS worked best for pancreatic cancer prediction (AUROC 82.24%), second only (AUROC 85.55%) to PheWAS (AUROC 85.87%) for dengue hemorrhagic fever prediction. DISCUSSION/CONCLUSION: In our experiments, terminologies with larger vocabularies and finer-grained representations were associated with better prediction performances. In particular, UMLS is consistently 1 of the best-performing ones. We believe that our work may help to inform better designs of predictive models, although further investigation is warranted. Laila Rasmy, Firat Tiryaki, Yujia Zhou 0003, Yang Xiang 0003, Cui Tao, Hua Xu 0001, Degui Zhi |
J. Am. Medical Informatics Assoc. | 4 |
| 2020 | A study of deep learning approaches for medication and adverse drug event extraction from clinical textabstractOBJECTIVE: This article presents our approaches to extraction of medications and associated adverse drug events (ADEs) from clinical documents, which is the second track of the 2018 National NLP Clinical Challenges (n2c2) shared task. MATERIALS AND METHODS: The clinical corpus used in this study was from the MIMIC-III database and the organizers annotated 303 documents for training and 202 for testing. Our system consists of 2 components: a named entity recognition (NER) and a relation classification (RC) component. For each component, we implemented deep learning-based approaches (eg, BI-LSTM-CRF) and compared them with traditional machine learning approaches, namely, conditional random fields for NER and support vector machines for RC, respectively. In addition, we developed a deep learning-based joint model that recognizes ADEs and their relations to medications in 1 step using a sequence labeling approach. To further improve the performance, we also investigated different ensemble approaches to generating optimal performance by combining outputs from multiple approaches. RESULTS: Our best-performing systems achieved F1 scores of 93.45% for NER, 96.30% for RC, and 89.05% for end-to-end evaluation, which ranked #2, #1, and #1 among all participants, respectively. Additional evaluations show that the deep learning-based approaches did outperform traditional machine learning algorithms in both NER and RC. The joint model that simultaneously recognizes ADEs and their relations to medications also achieved the best performance on RC, indicating its promise for relation extraction. CONCLUSION: In this study, we developed deep learning approaches for extracting medications and their attributes such as ADEs, and demonstrated its superior performance compared with traditional machine learning algorithms, indicating its uses in broader NER and RC tasks in the medical domain. Qiang Wei 0002, Zongcheng Ji, Zhiheng Li 0004, Jingcheng Du, Jun Xu 0007, Yang Xiang 0003, Firat Tiryaki, Stephen Wu 0004, Yaoyun Zhang, Cui Tao, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 7 |
| 2020 | Deep learning in clinical natural language processing: a methodical reviewabstractOBJECTIVE: This article methodically reviews the literature on deep learning (DL) for natural language processing (NLP) in the clinical domain, providing quantitative analysis to answer 3 research questions concerning methods, scope, and context of current research. MATERIALS AND METHODS: We searched MEDLINE, EMBASE, Scopus, the Association for Computing Machinery Digital Library, and the Association for Computational Linguistics Anthology for articles using DL-based approaches to NLP problems in electronic health records. After screening 1,737 articles, we collected data on 25 variables across 212 papers. RESULTS: DL in clinical NLP publications more than doubled each year, through 2018. Recurrent neural networks (60.8%) and word2vec embeddings (74.1%) were the most popular methods; the information extraction tasks of text classification, named entity recognition, and relation extraction were dominant (89.2%). However, there was a "long tail" of other methods and specific tasks. Most contributions were methodological variants or applications, but 20.8% were new methods of some kind. The earliest adopters were in the NLP community, but the medical informatics community was the most prolific. DISCUSSION: Our analysis shows growing acceptance of deep learning as a baseline for NLP research, and of DL-based NLP in the medical community. A number of common associations were substantiated (eg, the preference of recurrent neural networks for sequence-labeling named entity recognition), while others were surprisingly nuanced (eg, the scarcity of French language clinical NLP with deep learning). CONCLUSION: Deep learning has not yet fully penetrated clinical NLP and is growing rapidly. This review highlighted both the popular and unique trends in this active field. Stephen Wu 0004, Kirk Roberts, Surabhi Datta, Jingcheng Du, Zongcheng Ji, Yuqi Si, Sarvesh Soni, Qiang Wei 0002, Yang Xiang 0003, Bo Zhao 0001, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 10 |
| 2020 | Exploiting adversarial transfer learning for adverse drug reaction detection from texts
Zhiheng Li 0004, Ling Luo 0001, Yang Xiang 0003, Hongfei Lin |
J. Biomed. Informatics | 4 |
| 2019 | ML-Net: multi-label classification of biomedical texts with deep neural networksabstractOBJECTIVE: In multi-label text classification, each textual document is assigned 1 or more labels. As an important task that has broad applications in biomedicine, a number of different computational methods have been proposed. Many of these methods, however, have only modest accuracy or efficiency and limited success in practical use. We propose ML-Net, a novel end-to-end deep learning framework, for multi-label classification of biomedical texts. MATERIALS AND METHODS: ML-Net combines a label prediction network with an automated label count prediction mechanism to provide an optimal set of labels. This is accomplished by leveraging both the predicted confidence score of each label and the deep contextual information (modeled by ELMo) in the target document. We evaluate ML-Net on 3 independent corpora in 2 text genres: biomedical literature and clinical notes. For evaluation, we use example-based measures, such as precision, recall, and the F measure. We also compare ML-Net with several competitive machine learning and deep learning baseline models. RESULTS: Our benchmarking results show that ML-Net compares favorably to state-of-the-art methods in multi-label classification of biomedical text. ML-Net is also shown to be robust when evaluated on different text genres in biomedicine. CONCLUSION: ML-Net is able to accuractely represent biomedical document context and dynamically estimate the label count in a more systematic and accurate manner. Unlike traditional machine learning methods, ML-Net does not require human effort for feature engineering and is a highly efficient and scalable approach to tasks with a large set of labels, so there is no need to build individual classifiers for each separate label. Jingcheng Du, Qingyu Chen 0001, Yifan Peng 0002, Yang Xiang 0003, Cui Tao, Zhiyong Lu |
J. Am. Medical Informatics Assoc. | 4 |
| 2019 | Network context matters: graph convolutional network model over social networks improves the detection of unknown HIV infections among young men who have sex with menabstractOBJECTIVE: HIV infection risk can be estimated based on not only individual features but also social network information. However, there have been insufficient studies using n machine learning methods that can maximize the utility of such information. Leveraging a state-of-the-art network topology modeling method, graph convolutional networks (GCN), our main objective was to include network information for the task of detecting previously unknown HIV infections. MATERIALS AND METHODS: We used multiple social network data (peer referral, social, sex partners, and affiliation with social and health venues) that include 378 young men who had sex with men in Houston, TX, collected between 2014 and 2016. Due to the limited sample size, an ensemble approach was engaged by integrating GCN for modeling information flow and statistical machine learning methods, including random forest and logistic regression, to efficiently model sparse features in individual nodes. RESULTS: Modeling network information using GCN effectively increased the prediction of HIV status in the social network. The ensemble approach achieved 96.6% on accuracy and 94.6% on F1 measure, which outperformed the baseline methods (GCN, logistic regression, and random forest: 79.0%, 90.5%, 94.4% on accuracy, respectively; and 57.7%, 80.2%, 90.4% on F1). In the networks with missing HIV status, the ensemble also produced promising results. CONCLUSION: Network context is a necessary component in modeling infectious disease transmissions such as HIV. GCN, when combined with traditional machine learning approaches, achieved promising performance in detecting previously unknown HIV infections, which may provide a useful tool for combatting the HIV epidemic. Yang Xiang 0003, Kayo Fujimoto, John A. Schneider, Yuxi Jia, Degui Zhi, Cui Tao |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Mining Human Papillomavirus Vaccination Health Beliefs from Twitter Using Attentive Recurrent Neural Network
Jingcheng Du, Fang Li 0011, Yuxi Jia, Yang Xiang 0003, Sahiti Myneni, Cui Tao |
AMIA | 4 |
| 2018 | Asthma Onset Prediction Using Structured EMR Data
Yang Xiang 0003, Jun Xu 0007, Jingcheng Du, Degui Zhi, Cui Tao |
AMIA | 1 |
| 2018 | X-A-BiLSTM: a Deep Learning Approach for Depression Detection in Imbalanced Data
Qing Cong, Zhiyong Feng 0002, Fang Li 0011, Yang Xiang 0003, Guozheng Rao, Cui Tao |
BIBM | 4 |
| 2017 | Answer Selection in Community Question Answering by Normalizing Support Answers
Zhihui Zheng, Daohe Lu, Qingcai Chen, Yang Xiang 0003, Youcheng Pan |
NLPCC | 5 |
| 2017 | Answer Selection in Community Question Answering via Attentive Neural NetworksabstractAnswer selection in community question answering (cQA) is a challenging task in natural language processing. The difficulty lies in that it not only needs the consideration of semantic matching between question answer pairs but also requires a serious modeling of contextual factors. In this letter, we propose an attentive deep neural network architecture so as to learn the deterministic information for answer selection. The architecture can support various input formats through the organization of convolutional neural networks, attention-based long short-term memory, and conditional random fields. Experiments are carried out on the SemEval-2015 cQA dataset. We attain 58.35% on macroaveraged F1, which outperforms the Top-1 system in the shared task by 1.16% and improves the state-of-the-art deep-neural-network-based method by 2.21%. Yang Xiang 0003, Qingcai Chen, Xiaolong Wang 0001, Yang Qin 0001 |
IEEE Signal Process. Lett. | 1 |
| 2016 | Incorporating Label Dependency for Answer Quality Tagging in Community Question Answering via CNN-LSTM-CRFabstractIn community question answering (cQA), the quality of answers are determined by the matching degree between question-answer pairs and the correlation among the answers. In this paper, we show that the dependency between the answer quality labels also plays a pivotal role. To validate the effectiveness of label dependency, we propose two neural network-based models, with different combination modes of Convolutional Neural Net-works, Long Short Term Memory and Conditional Random Fields. Extensive experi-ments are taken on the dataset released by the SemEval-2015 cQA shared task. The first model is a stacked ensemble of the networks. It achieves 58.96% on macro averaged F1, which improves the state-of-the-art neural network-based method by 2.82% and outper-forms the Top-1 system in the shared task by 1.77%. The second is a simple attention-based model whose input is the connection of the question and its corresponding answers. It produces promising results with 58.29% on overall F1 and gains the best performance on the Good and Bad categories. Yang Xiang 0003, Xiaoqiang Zhou, Qingcai Chen, Zhihui Zheng, Buzhou Tang, Xiaolong Wang 0001, Yang Qin 0001 |
COLING | 1 |
| 2015 | Distant Supervision for Relation Extraction via Group Selection
Yang Xiang 0003, Xiaolong Wang 0001, Yaoyun Zhang, Yang Qin 0001, Shixi Fan |
ICONIP (2) | 1 |
| 2013 | Grammatical Error Correction Using Feature Selection and Confidence Tuning
Yang Xiang 0003, Yaoyun Zhang, Xiaolong Wang 0001, Chongqiang Wei, Xiaoqiang Zhou, Yuxiu Hu, Yang Qin 0001 |
IJCNLP | 1 |