VLDB 2026 Research / reviewers in the wild / expert
Fei Xia 0004
dblp:79/1081-4
· DBLP profile ↗
54ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0001-6015-6064ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 10 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identifying Imaging Follow-Up in Radiology Reports: A Comparative Analysis of Traditional ML and LLM ApproachesabstractLarge language models (LLMs) have shown considerable promise in clinical natural language processing, yet few domain-specific datasets exist to rigorously evaluate their performance on radiology tasks. In this work, we introduce an annotated corpus of 6,393 radiology reports from 586 patients, each labeled for follow-up imaging status, to support the development and benchmarking of follow-up adherence detection systems. Using this corpus, we systematically compared traditional machine-learning classifiers, including logistic regression (LR), support vector machines (SVM), Longformer, and a fully fine-tuned Llama3-8B-Instruct, with recent generative LLMs. To evaluate generative LLMs, we tested GPT-4o and the open-source GPT-OSS-20B under two configurations: a baseline (Base) and a task-optimized (Advanced) setting that focused inputs on metadata, recommendation sentences, and their surrounding context. A refined prompt for GPT-OSS-20B further improved reasoning accuracy. Performance was assessed using precision, recall, and F1 scores with 95% confidence intervals estimated via non-parametric bootstrapping. Inter-annotator agreement was high (F1 = 0.846). GPT-4o (Advanced) achieved the best performance (F1 = 0.832), followed closely by GPT-OSS-20B (Advanced; F1 = 0.828). LR and SVM also performed strongly (F1 = 0.776 and 0.775), underscoring that while LLMs approach human-level agreement through prompt optimization, interpretable and resource-efficient models remain valuable baselines. Namu Park, Giridhar Kaushik Ramachandran, Kevin Lybarger, Fei Xia 0004, Özlem Uzuner, Martin L. Gunn, Meliha Yetisgen |
LREC | 4 |
| 2026 | MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question AnsweringabstractEvaluating natural language generation (NLG) systems in the medical domain presents unique challenges due to the critical demands for accuracy, relevance, and domain-specific expertise. Traditional automatic evaluation metrics, such as BLEU, ROUGE, and BERTScore, often fall short in distinguishing between high-quality outputs, especially given the open-ended nature of medical question answering (QA) tasks where multiple valid responses may exist. In this work, we introduce MORQA (Medical Open-Response QA), a new multilingual benchmark designed to assess the effectiveness of NLG evaluation metrics across three medical visual and text-based QA datasets in English and Chinese. Unlike prior resources, our datasets feature 2-4+ gold-standard answers authored by medical professionals, along with expert human ratings for three English and Chinese subsets. We benchmark both traditional metrics and large language model (LLM)-based evaluators, such as GPT-4 and Gemini, finding that LLM-based approaches significantly outperform traditional metrics in correlating with expert judgments. We further analyze factors driving this improvement, including LLMs' sensitivity to semantic nuances and robustness to variability among reference answers. Our results provide the first comprehensive, multilingual qualitative study of NLG evaluation in the medical domain, highlighting the need for human-aligned evaluation methods. All datasets and annotations will be publicly released to support future research. Wen-Wai Yim, Asma Ben Abacha, Zixuan Yu, Robert Doerning, Fei Xia 0004, Meliha Yetisgen |
LREC | 5 |
| 2025 | A scoping review of natural language processing in addressing medically inaccurate information: Errors, misinformation, and hallucination
Zhaoyi Sun, Wen-Wai Yim, Özlem Uzuner, Fei Xia 0004, Meliha Yetisgen |
J. Biomed. Informatics | 4 |
| 2025 | WoundcareVQA: A multilingual visual question answering benchmark dataset for wound careabstractOBJECTIVE: Introduce the task of wound care multimodal multilingual visual question answering, provide baseline performances, and identify areas of future study. METHODS: A dataset of wound care multimodal multilingual visual question answering (VQA) was created using consumer health questions asked online. Practicing US medical doctors were tasked with providing metadata and expert responses labels. Several instruct-enabled, multilingual visual question answering models (GPT-4o, Gemini-1.5-Pro, and Qwen-VL) were tested to benchmark performances. Finally, automatic evaluations were tested against domain expert response ratings. RESULTS: A multilingual dataset of 477 wound care cases, 768 responses, 748 images, 3k structured data labels, 1362 translation instances, and 10k judgments was constructed (https://osf.io/xsj5u/). Metadata scores ranged from 0.32-0.78 accuracy depending on classification type; response generation performances 0.06 BLEU, 0.66 BERTScore, 0.45 ROUGE-L in English and 0.12 BLEU, 0.69 BERTScore, and 0.50 ROUGE-L in Chinese. CONCLUSION: We construct and explore the tasks of multimodal, multilingual VQA. We hope the work here can inspire further research in wound care metadata classification, VQA response generation, and open response automatic evaluation. Wen-Wai Yim, Asma Ben Abacha, Robert Doerning, Chia-Yu Chen, Jiaying Xu, Anita Subbarao, Zixuan Yu, Fei Xia 0004, M. Kennedy Hall, Meliha Yetisgen |
J. Biomed. Informatics | 8 |
| 2024 | Large Language Models Are No Longer Shallow ParsersabstractThe development of large language models (LLMs) brings significant changes to the field of natural language processing (NLP), enabling remarkable performance in various high-level tasks, such as machine translation, questionanswering, dialogue generation, etc., under endto-end settings without requiring much training data.Meanwhile, fundamental NLP tasks, particularly syntactic parsing, are also essential for language study as well as evaluating the capability of LLMs for instruction understanding and usage.In this paper, we focus on analyzing and improving the capability of current state-of-the-art LLMs on a classic fundamental task, namely constituency parsing, which is the representative syntactic task in both linguistics and natural language processing.We observe that these LLMs are effective in shallow parsing but struggle with creating correct full parse trees.To improve the performance of LLMs on deep syntactic parsing, we propose a three-step approach that firstly prompts LLMs for chunking, then filters out low-quality chunks, and finally adds the remaining chunks to prompts to instruct LLMs for parsing, with later enhancement by chain-of-thought prompting.Experimental results on English and Chinese benchmark datasets demonstrate the effectiveness of our approach on improving LLMs' performance on constituency parsing. Yuanhe Tian, Fei Xia 0004, Yan Song 0001 |
ACL (1) | 2 |
| 2024 | Dialogue Summarization with Mixture of Experts based on Large Language ModelsabstractDialogue summarization is an important task that requires to generate highlights for a conversation from different aspects (e.g., content of various speakers).While several studies successfully employ large language models (LLMs) and achieve satisfying results, they are limited by using one model at a time or treat it as a black box, which makes it hard to discriminatively learn essential content in a dialogue from different aspects, therefore may lead to anticipation bias and potential loss of information in the produced summaries.In this paper, we propose an LLM-based approach with roleoriented routing and fusion generation to utilize mixture of experts (MoE) for dialogue summarization.Specifically, the role-oriented routing is an LLM-based module that selects appropriate experts to process different information; fusion generation is another LLM-based module to locate salient information and produce finalized dialogue summaries.The proposed approach offers an alternative solution to employing multiple LLMs for dialogue summarization by leveraging their capabilities of in-context processing and generation in an effective manner.We run experiments on widely used benchmark datasets for this task, where the results demonstrate the superiority of our approach in producing informative and accurate dialogue summarization. 1 Yuanhe Tian, Fei Xia 0004, Yan Song 0001 |
ACL (1) | 2 |
| 2024 | Extracting Social Determinants of Health from Pediatric Patient Notes Using Large Language Models: Novel Corpus and MethodsabstractSocial determinants of health (SDoH) play a critical role in shaping health outcomes, particularly in pediatric populations where interventions can have long-term implications. SDoH are frequently studied in the Electronic Health Record (EHR), which provides a rich repository for diverse patient data. In this work, we present a novel annotated corpus, the Pediatric Social History Annotation Corpus (PedSHAC), and evaluate the automatic extraction of detailed SDoH representations using fine-tuned and in-context learning methods with Large Language Models (LLMs). PedSHAC comprises annotated social history sections from 1,260 clinical notes obtained from pediatric patients within the University of Washington (UW) hospital system. Employing an event-based annotation scheme, PedSHAC captures ten distinct health determinants to encompass living and economic stability, prior trauma, education access, substance use history, and mental health with an overall annotator agreement of 81.9 F1. Our proposed fine-tuning LLM-based extractors achieve high performance at 78.4 F1 for event arguments. In-context learning approaches with GPT-4 demonstrate promise for reliable SDoH extraction with limited annotated examples, with extraction performance at 82.3 F1 for event triggers. Yujuan Fu, Giridhar Kaushik Ramachandran, Nicholas J. Dobbins, Namu Park, Michael Leu, Abby R. Rosenberg, Kevin Lybarger, Fei Xia 0004, Özlem Uzuner, Meliha Yetisgen |
LREC/COLING | 8 |
| 2024 | DermaVQA: A Multilingual Visual Question Answering Dataset for Dermatology
Wen-Wai Yim, Yujuan Fu, Zhaoyi Sun, Asma Ben Abacha, Meliha Yetisgen, Fei Xia 0004 |
MICCAI (5) | 6 |
| 2024 | Diffusion Networks with Task-Specific Noise Control for Radiology Report GenerationabstractExisting radiology report generation (RRG) studies mostly adopt autoregressive (AR) approaches to produce textual descriptions token-by-token for specific clinical radiographs, where they are susceptible to error propagation problems if irrelevant contents are half-way generated, leading to potential ill-presenting of precise diagnoses, especially when there exist complicated abnormalities in radiographs. Although the non-AR paradigm, e.g., diffusion model, provides an alternative solution to tackle the problem from AR by generating all contents in parallel, the mechanism of using Gaussian noise in existing diffusion models still has significant room to improve when such models are used in particular circumstances, i.e., providing proper guidance in controlling noises in the diffusive process to ensure precise report generation. In this paper, we propose to conduct RRG with diffusion networks by controlling the noise with task-specific features, which leverages irrelevant visual and textual information as noise rather than the stochastic Gaussian noise, and allows the diffusion networks to filter particular information through iterative denoising, thus performing a precise and controlled report generation process. Experiments on IU X-Ray and MIMIC-CXR demonstrate the superiority of our approach compared to strong baselines and state-of-the-art solutions. Human evaluation and noise type analysis show that comprehensive noise control greatly helps diffusion networks to refine the generation of global and local report contents. Yuanhe Tian, Fei Xia 0004, Yan Song 0003 |
ACM Multimedia | 2 |
| 2024 | Emotion Cause Extraction in Conversations with Response Graphing
Yuanhe Tian, Pengsen Cheng, Fei Xia 0004, Yongdong Zhang 0001, Yan Song 0004 |
NLPCC (5) | 3 |
| 2024 | CACER: Clinical concept Annotations for Cancer Events and RelationsabstractOBJECTIVE: Clinical notes contain unstructured representations of patient histories, including the relationships between medical problems and prescription drugs. To investigate the relationship between cancer drugs and their associated symptom burden, we extract structured, semantic representations of medical problem and drug information from the clinical narratives of oncology notes. MATERIALS AND METHODS: We present Clinical concept Annotations for Cancer Events and Relations (CACER), a novel corpus with fine-grained annotations for over 48 000 medical problems and drug events and 10 000 drug-problem and problem-problem relations. Leveraging CACER, we develop and evaluate transformer-based information extraction models such as Bidirectional Encoder Representations from Transformers (BERT), Fine-tuned Language Net Text-To-Text Transfer Transformer (Flan-T5), Large Language Model Meta AI (Llama3), and Generative Pre-trained Transformers-4 (GPT-4) using fine-tuning and in-context learning (ICL). RESULTS: In event extraction, the fine-tuned BERT and Llama3 models achieved the highest performance at 88.2-88.0 F1, which is comparable to the inter-annotator agreement (IAA) of 88.4 F1. In relation extraction, the fine-tuned BERT, Flan-T5, and Llama3 achieved the highest performance at 61.8-65.3 F1. GPT-4 with ICL achieved the worst performance across both tasks. DISCUSSION: The fine-tuned models significantly outperformed GPT-4 in ICL, highlighting the importance of annotated training data and model optimization. Furthermore, the BERT models performed similarly to Llama3. For our task, large language models offer no performance advantage over the smaller BERT models. CONCLUSIONS: We introduce CACER, a novel corpus with fine-grained annotations for medical problems, drugs, and their relationships in clinical narratives of oncology notes. State-of-the-art transformer models achieved performance comparable to IAA for several extraction tasks. Yujuan Fu, Giridhar Kaushik Ramachandran, Ahmad Halwani, Bridget T. McInnes, Fei Xia 0004, Kevin Lybarger, Meliha Yetisgen, Özlem Uzuner |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | Enhancing Structure-aware Encoder with Extremely Limited Data for Graph-based Dependency ParsingabstractDependency parsing is an important fundamental natural language processing task which analyzes the syntactic structure of an input sentence by illustrating the syntactic relations between words. To improve dependency parsing, leveraging existing dependency parsers and extra data (e.g., through semi-supervised learning) has been demonstrated to be effective, even though the final parsers are trained on inaccurate (but massive) data. In this paper, we propose a frustratingly easy approach to improve graph-based dependency parsing, where a structure-aware encoder is pre-trained on auto-parsed data by predicting the word dependencies and then fine-tuned on gold dependency trees, which differs from the usual pre-training process that aims to predict the context words along dependency paths. Experimental results and analyses demonstrate the effectiveness and robustness of our approach to benefit from the data (even with noise) processed by different parsers, where our approach outperforms strong baselines under different settings with different dependency standards and model architectures used in pre-training and fine-tuning. More importantly, further analyses find that only 2K auto-parsed sentences are required to obtain improvement when pre-training vanilla BERT-large based parser without requiring extra parameters. Yuanhe Tian, Yan Song 0003, Fei Xia 0004 |
COLING | 3 |
| 2022 | Complementary Learning of Aspect Terms for Aspect-based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) aims to predict the sentiment polarity towards a given aspect term in a sentence on the fine-grained level, which usually requires a good understanding of contextual information, especially appropriately distinguishing of a given aspect and its contexts, to achieve good performance. However, most existing ABSA models pay limited attention to the modeling of the given aspect terms and thus result in inferior results when a sentence contains multiple aspect terms with contradictory sentiment polarities. In this paper, we propose to improve ABSA by complementary learning of aspect terms, which serves as a supportive auxiliary task to enhance ABSA by explicitly recovering the aspect terms from each input sentence so as to better understand aspects and their contexts. Particularly, a discriminator is also introduced to further improve the learning process by appropriately balancing the impact of aspect recovery to sentiment prediction. Experimental results on five widely used English benchmark datasets for ABSA demonstrate the effectiveness of our approach, where state-of-the-art performance is observed on all datasets. Han Qin, Yuanhe Tian, Fei Xia 0004, Yan Song 0003 |
LREC | 3 |
| 2022 | ChiMST: A Chinese Medical Corpus for Word Segmentation and Medical Term RecognitionabstractChinese word segmentation (CWS) and named entity recognition (NER) are two important tasks in Chinese natural language processing. To achieve good model performance on these tasks, existing neural approaches normally require a large amount of labeled training data, which is often unavailable for specific domains such as the Chinese medical domain due to privacy and legal issues. To address this problem, we have developed a Chinese medical corpus named ChiMST which consists of question-answer pairs collected from an online medical healthcare platform and is annotated with word boundary and medical term information. For word boundary, we mainly follow the word segmentation guidelines for the Penn Chinese Treebank (Xia, 2000); for medical terms, we define 9 categories and 18 sub-categories after consulting medical experts. To provide baselines on this corpus, we train existing state-of-the-art models on it and achieve good performance. We believe that the corpus and the baseline systems will be a valuable resource for CWS and NER research on the medical domain. Yuanhe Tian, Han Qin, Fei Xia 0004, Yan Song 0003 |
LREC | 3 |
| 2022 | Syntax-driven Approach for Semantic Role LabelingabstractAs an important task to analyze the semantic structure of a sentence, semantic role labeling (SRL) aims to locate the semantic role (e.g., agent) of noun phrases with respect to a given predicate and thus plays an important role in downstream tasks such as dialogue systems. To achieve a better performance in SRL, a model is always required to have a good understanding of the context information. Although one can use advanced text encoder (e.g., BERT) to capture the context information, extra resources are also required to further improve the model performance. Considering that there are correlations between the syntactic structure and the semantic structure of the sentence, many previous studies leverage auto-generated syntactic knowledge, especially the dependencies, to enhance the modeling of context information through graph-based architectures, where limited attention is paid to other types of auto-generated knowledge. In this paper, we propose map memories to enhance SRL by encoding different types of auto-generated syntactic knowledge (i.e., POS tags, syntactic constituencies, and word dependencies) obtained from off-the-shelf toolkits. Experimental results on two English benchmark datasets for span-style SRL (i.e., CoNLL-2005 and CoNLL-2012) demonstrate the effectiveness of our approach, which outperforms strong baselines and achieves state-of-the-art results on CoNLL-2005. Yuanhe Tian, Han Qin, Fei Xia 0004, Yan Song 0003 |
LREC | 3 |
| 2020 | Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-way Attentions of Auto-analyzed KnowledgeabstractChinese word segmentation (CWS) and partof-speech (POS) tagging are important fundamental tasks for Chinese language processing, where joint learning of them is an effective one-step solution for both tasks.Previous studies for joint CWS and POS tagging mainly follow the character-based tagging paradigm with introducing contextual information such as n-gram features or sentential representations from recurrent neural models.However, for many cases, the joint tagging needs not only modeling from context features but also knowledge attached to them (e.g., syntactic relations among words); limited efforts have been made by existing research to meet such needs.In this paper, we propose a neural model named TWASP for joint CWS and POS tagging following the character-based sequence labeling paradigm, where a two-way attention mechanism is used to incorporate both context feature and their corresponding syntactic knowledge for each input character.Particularly, we use existing language processing toolkits to obtain the auto-analyzed syntactic knowledge for the context, and the proposed attention module can learn and benefit from them although their quality may not be perfect.Our experiments illustrate the effectiveness of the two-way attentions for joint CWS and POS tagging, where state-of-the-art performance is achieved on five benchmark datasets.1 Yuanhe Tian, Yan Song 0003, Xiang Ao 0001, Fei Xia 0004, Xiaojun Quan, Tong Zhang 0001 |
ACL | 4 |
| 2020 | Improving Chinese Word Segmentation with Wordhood Memory NetworksabstractContextual features always play an important role in Chinese word segmentation (CWS).Wordhood information, being one of the contextual features, is proved to be useful in many conventional character-based segmenters.However, this feature receives less attention in recent neural models and it is also challenging to design a framework that can properly integrate wordhood information from different wordhood measures to existing neural frameworks.In this paper, we therefore propose a neural framework, WMSEG, which uses memory networks to incorporate wordhood information with several popular encoder-decoder combinations for CWS.Experimental results on five benchmark datasets indicate the memory mechanism successfully models wordhood information for neural segmenters and helps WMSEG achieve state-ofthe-art performance on all those datasets.Further experiments and analyses also demonstrate the robustness of our proposed framework with respect to different wordhood measures and the efficiency of wordhood information in cross-domain experiments.1 Yuanhe Tian, Yan Song 0003, Fei Xia 0004, Tong Zhang 0001 |
ACL | 3 |
| 2020 | Summarizing Medical Conversations via Identifying Important UtterancesabstractSummarization is an important natural language processing (NLP) task in identifying key information from text.For conversations, the summarization systems need to extract salient contents from spontaneous utterances by multiple speakers.In a special task-oriented scenario, namely medical conversations between patients and doctors, the symptoms, diagnoses, and treatments could be highly important because the nature of such conversation is to find a medical solution to the problem proposed by the patients.Especially consider that current online medical platforms provide millions of public available conversations between real patients and doctors, where the patients propose their medical problems and the registered doctors offer diagnosis and treatment, a conversation in most cases could be too long and the key information is hard to be located.Therefore, summarizations to the patients' problems and the doctors' treatments in the conversations can be highly useful, in terms of helping other patients with similar problems have a precise reference for potential medical solutions.In this paper, we focus on medical conversation summarization, using a dataset of medical conversations and corresponding summaries which were crawled from a well-known online healthcare service provider in China.We propose a hierarchical encoder-tagger model (HET) to generate summaries by identifying important utterances (with respect to problem proposing and solving) in the conversations.For the particular dataset used in this study, we show that high-quality summaries can be generated by extracting two types of utterances, namely, problem statements and treatment recommendations.Experimental results demonstrate that HET outperforms strong baselines and models from previous studies, and adding conversation-related features can further improve system performance.1 * Equal contribution. 1 Our code, models, and the dataset are released at https://github.com/cuhksz-nlp/HET-MC. 2 E.g., in China, the number of outpatient visits exceeded 7 billions and inpatient visits Yan Song 0003, Yuanhe Tian, Fei Xia 0004 |
COLING | 4 |
| 2020 | Joint Chinese Word Segmentation and Part-of-speech Tagging via Multi-channel Attention of Character N-gramsabstractChinese word segmentation (CWS) and part-of-speech (POS) tagging are two fundamental tasks for Chinese language processing.Previous studies have demonstrated that jointly performing them can be an effective one-step solution to both tasks and this joint task can benefit from a good modeling of contextual features such as n-grams.However, their work on modeling such contextual features is limited to concatenating the features or their embeddings directly with the input embeddings without distinguishing whether the contextual features are important for the joint task in the specific context.Therefore, their models for the joint task could be misled by unimportant contextual information.In this paper, we propose a character-based neural model for the joint task enhanced by multi-channel attention of n-grams.In the attention module, n-gram features are categorized into different groups according to several criteria, and n-grams in each group are weighted and distinguished according to their importance for the joint task in the specific context.To categorize n-grams, we try two criteria in this study, i.e., n-gram frequency and length, so that n-grams having different capabilities of carrying contextual information are discriminatively learned by our proposed attention module.Experimental results on five benchmark datasets for CWS and POS tagging demonstrate that our approach outperforms strong baseline models and achieves state-of-the-art performance on all five datasets.1 Yuanhe Tian, Yan Song 0003, Fei Xia 0004 |
COLING | 3 |
| 2020 | Supertagging Combinatory Categorial Grammar with Attentive Graph Convolutional NetworksabstractSupertagging is conventionally regarded as an important task for combinatory categorial grammar (CCG) parsing, where effective modeling of contextual information is highly important to this task.However, existing studies have made limited efforts to leverage contextual features except for applying powerful encoders (e.g., bi-LSTM).In this paper, we propose attentive graph convolutional networks to enhance neural CCG supertagging through a novel solution of leveraging contextual information.Specifically, we build the graph from chunks (n-grams) extracted from a lexicon and apply attention over the graph, so that different word pairs from the contexts within and across chunks are weighted in the model and facilitate the supertagging accordingly.The experiments performed on the CCGbank demonstrate that our approach outperforms all previous studies in terms of both supertagging and parsing.Further analyses illustrate the effectiveness of each component in our approach to discriminatively learn from word pairs to enhance CCG supertagging. 1 Yuanhe Tian, Yan Song 0003, Fei Xia 0004 |
EMNLP (1) | 3 |
| 2020 | Improving biomedical named entity recognition with syntactic informationabstractBACKGROUND: Biomedical named entity recognition (BioNER) is an important task for understanding biomedical texts, which can be challenging due to the lack of large-scale labeled training data and domain knowledge. To address the challenge, in addition to using powerful encoders (e.g., biLSTM and BioBERT), one possible method is to leverage extra knowledge that is easy to obtain. Previous studies have shown that auto-processed syntactic information can be a useful resource to improve model performance, but their approaches are limited to directly concatenating the embeddings of syntactic information to the input word embeddings. Therefore, such syntactic information is leveraged in an inflexible way, where inaccurate one may hurt model performance. RESULTS: In this paper, we propose BIOKMNER, a BioNER model for biomedical texts with key-value memory networks (KVMN) to incorporate auto-processed syntactic information. We evaluate BIOKMNER on six English biomedical datasets, where our method with KVMN outperforms the strong baseline method, namely, BioBERT, from the previous study on all datasets. Specifically, the F1 scores of our best performing model are 85.29% on BC2GM, 77.83% on JNLPBA, 94.22% on BC5CDR-chemical, 90.08% on NCBI-disease, 89.24% on LINNAEUS, and 76.33% on Species-800, where state-of-the-art performance is obtained on four of them (i.e., BC2GM, BC5CDR-chemical, NCBI-disease, and Species-800). CONCLUSION: The experimental results on six English benchmark datasets demonstrate that auto-processed syntactic information can be a useful resource for BioNER and our method with KVMN can appropriately leverage such information to improve model performance. Yuanhe Tian, Wang Shen, Yan Song 0003, Fei Xia 0004, Kenli Li 0001 |
BMC Bioinform. | 4 |
| 2018 | Constructing a Chinese Medical Conversation Corpus Annotated with Conversational Structures and Actions
Yan Song 0003, Fei Xia 0004 |
LREC | 3 |
| 2017 | Learning Word Representations with Regularization from Prior KnowledgeabstractConventional word embeddings are trained with specific criteria (e.g., based on language modeling or co-occurrence) inside a single information source, disregarding the opportunity for further calibration using external knowledge.This paper presents a unified framework that leverages pre-learned or external priors, in the form of a regularizer, for enhancing conventional language model-based embedding learning.We consider two types of regularizers.The first type is derived from topic distribution by running latent Dirichlet allocation on unlabeled data.The second type is based on dictionaries that are created with human annotation efforts.To effectively learn with the regularizers, we propose a novel data structure, trajectory softmax, in this paper.The resulting embeddings are evaluated by word similarity and sentiment classification.Experimental results show that our learning framework with regularization from prior knowledge improves embedding quality across multiple datasets, compared to a diverse collection of baseline methods. Yan Song 0003, Fei Xia 0004 |
CoNLL | 3 |
| 2016 | Annotating and Detecting Medical Events in Clinical Notes
Prescott Klassen, Fei Xia 0004, Meliha Yetisgen |
LREC | 2 |
| 2015 | Enriching, Editing, and Representing Interlinear Glossed Text
Fei Xia 0004, Michael Wayne Goodman, Ryan Georgi, Glenn Slayden, William D. Lewis |
CICLing (1) | 1 |
| 2014 | Annotating Clinical Events in Text Snippets for Phenotype Detection
Prescott Klassen, Fei Xia 0004, Lucy Vanderwende, Meliha Yetisgen |
LREC | 2 |
| 2014 | Modern Chinese Helps Archaic Chinese Processing: Finding and Exploiting the Shared Properties
Yan Song 0003, Fei Xia 0004 |
LREC | 2 |
| 2014 | Enriching ODIN
Fei Xia 0004, William D. Lewis, Michael Wayne Goodman, Joshua Crowgey, Emily M. Bender |
LREC | 1 |
| 2013 | A Common Case of Jekyll and Hyde: The Synergistic Effect of Using Divided Source Training Data for Feature Augmentation
Yan Song 0003, Fei Xia 0004 |
IJCNLP | 2 |
| 2013 | Large-scale evaluation of automated clinical note de-identification and its impact on information extractionabstractOBJECTIVE: (1) To evaluate a state-of-the-art natural language processing (NLP)-based approach to automatically de-identify a large set of diverse clinical notes. (2) To measure the impact of de-identification on the performance of information extraction algorithms on the de-identified documents. MATERIAL AND METHODS: A cross-sectional study that included 3503 stratified, randomly selected clinical notes (over 22 note types) from five million documents produced at one of the largest US pediatric hospitals. Sensitivity, precision, F value of two automated de-identification systems for removing all 18 HIPAA-defined protected health information elements were computed. Performance was assessed against a manually generated 'gold standard'. Statistical significance was tested. The automated de-identification performance was also compared with that of two humans on a 10% subsample of the gold standard. The effect of de-identification on the performance of subsequent medication extraction was measured. RESULTS: The gold standard included 30 815 protected health information elements and more than one million tokens. The most accurate NLP method had 91.92% sensitivity (R) and 95.08% precision (P) overall. The performance of the system was indistinguishable from that of human annotators (annotators' performance was 92.15%(R)/93.95%(P) and 94.55%(R)/88.45%(P) overall while the best system obtained 92.91%(R)/95.73%(P) on same text). The impact of automated de-identification was minimal on the utility of the narrative notes for subsequent information extraction as measured by the sensitivity and precision of medication name extraction. DISCUSSION AND CONCLUSION: NLP-based de-identification shows excellent performance that rivals the performance of human annotators. Furthermore, unlike manual de-identification, the automated approach scales up to millions of documents quickly and inexpensively. Louise Deléger, Katalin Molnár, Guergana K. Savova, Fei Xia 0004, Todd Lingren, Qi Li 0004, Keith Marsolo, Anil G. Jegga, Megan Kaiser, Laura Stoutenborough, Imre Solti |
J. Am. Medical Informatics Assoc. | 4 |
| 2013 | Assertion modeling and its role in clinical phenotype identification
Cosmin Adrian Bejan, Lucy Vanderwende, Fei Xia 0004, Meliha Yetisgen |
J. Biomed. Informatics | 3 |
| 2013 | A text processing pipeline to extract recommendations from radiology reports
Meliha Yetisgen, Martin L. Gunn, Fei Xia 0004, Thomas H. Payne |
J. Biomed. Informatics | 3 |
| 2012 | Smoking Status Detection Across Domains
Michael Tepper, Fei Xia 0004, Meliha Yetisgen |
AMIA | 2 |
| 2012 | Measuring the Divergence of Dependency Structures Cross-Linguistically to Improve Syntactic Projection Algorithms
Ryan Georgi, Fei Xia 0004, William D. Lewis |
LREC | 2 |
| 2012 | Using a Goodness Measurement for Domain Adaptation: A Case Study on Chinese Word Segmentation
Yan Song 0003, Fei Xia 0004 |
LREC | 2 |
| 2012 | Statistical Section Segmentation in Free-Text Clinical Records
Michael Tepper, Daniel Capurro, Fei Xia 0004, Lucy Vanderwende, Meliha Yetisgen |
LREC | 3 |
| 2012 | Pneumonia identification using statistical feature selectionabstractOBJECTIVE: This paper describes a natural language processing system for the task of pneumonia identification. Based on the information extracted from the narrative reports associated with a patient, the task is to identify whether or not the patient is positive for pneumonia. DESIGN: A binary classifier was employed to identify pneumonia from a dataset of multiple types of clinical notes created for 426 patients during their stay in the intensive care unit. For this purpose, three types of features were considered: (1) word n-grams, (2) Unified Medical Language System (UMLS) concepts, and (3) assertion values associated with pneumonia expressions. System performance was greatly increased by a feature selection approach which uses statistical significance testing to rank features based on their association with the two categories of pneumonia identification. RESULTS: Besides testing our system on the entire cohort of 426 patients (unrestricted dataset), we also used a smaller subset of 236 patients (restricted dataset). The performance of the system was compared with the results of a baseline previously proposed for these two datasets. The best results achieved by the system (85.71 and 81.67 F1-measure) are significantly better than the baseline results (50.70 and 49.10 F1-measure) on the restricted and unrestricted datasets, respectively. CONCLUSION: Using a statistical feature selection approach that allows the feature extractor to consider only the most informative features from the feature space significantly improves the performance over a baseline that uses all the features from the same feature space. Extracting the assertion value for pneumonia expressions further improves the system performance. Cosmin Adrian Bejan, Fei Xia 0004, Lucy Vanderwende, Mark M. Wurfel, Meliha Yetisgen |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | Linguistic Phenomena, Analyses, and Representations: Understanding Conversion between Treebanks
Rajesh Bhatt, Owen Rambow, Fei Xia 0004 |
IJCNLP | 3 |
| 2010 | Comparing Language Similarity across Genetic and Typologically-Based Groupings
Ryan Georgi, Fei Xia 0004, William D. Lewis |
COLING | 2 |
| 2010 | Empty Categories in a Hindi Treebank
Archna Bhatia, Rajesh Bhatt, Bhuvana Narasimhan, Martha Palmer, Owen Rambow, Dipti Misra Sharma, Michael Tepper, Ashwini Vaidya, Fei Xia 0004 |
LREC | 9 |
| 2010 | The Problems of Language Identification within Hugely Multilingual Data Sets
Fei Xia 0004, Carrie Lewis, William D. Lewis |
LREC | 1 |
| 2010 | Community annotation experiment for ground truth generation for the i2b2 medication challengeabstractOBJECTIVE: Within the context of the Third i2b2 Workshop on Natural Language Processing Challenges for Clinical Records, the authors (also referred to as 'the i2b2 medication challenge team' or 'the i2b2 team' for short) organized a community annotation experiment. DESIGN: For this experiment, the authors released annotation guidelines and a small set of annotated discharge summaries. They asked the participants of the Third i2b2 Workshop to annotate 10 discharge summaries per person; each discharge summary was annotated by two annotators from two different teams, and a third annotator from a third team resolved disagreements. MEASUREMENTS: In order to evaluate the reliability of the annotations thus produced, the authors measured community inter-annotator agreement and compared it with the inter-annotator agreement of expert annotators when both the community and the expert annotators generated ground truth based on pooled system outputs. For this purpose, the pool consisted of the three most densely populated automatic annotations of each record. The authors also compared the community inter-annotator agreement with expert inter-annotator agreement when the experts annotated raw records without using the pool. Finally, they measured the quality of the community ground truth by comparing it with the expert ground truth. RESULTS AND CONCLUSIONS: The authors found that the community annotators achieved comparable inter-annotator agreement to expert annotators, regardless of whether the experts annotated from the pool. Furthermore, the ground truth generated by the community obtained F-measures above 0.90 against the ground truth of the experts, indicating the value of the community as a source of high-quality ground truth even on intricate and domain-specific annotation tasks. Özlem Uzuner, Imre Solti, Fei Xia 0004, Eithon Cadag |
J. Am. Medical Informatics Assoc. | 3 |
| 2010 | Inducing Morphemes Using Light KnowledgeabstractAllomorphic variation, or form variation among morphs with the same meaning, is a stumbling block to morphological induction (MI). To address this problem, we present a hybrid approach that uses a small amount of linguistic knowledge in the form of orthographic rewrite rules to help refine an existing MI-produced segmentation. Using rules, we derive underlying analyses of morphs---generalized with respect to contextual spelling differences---from an existing surface morph segmentation, and from these we learn a morpheme-level segmentation. To learn morphemes, we have extended the Morfessor segmentation algorithm [Creutz and Lagus 2004; 2005; 2006] by using rules to infer possible underlying analyses from surface segmentations. A segmentation produced by Morfessor Categories-MAP Software v. 0.9.2 is used as input to our procedure and as a baseline that we evaluate against. To suggest analyses for our procedure, a set of language-specific orthographic rules is needed. Our procedure has yielded promising improvements for English and Turkish over the baseline approach when tested on the Morpho Challenge 2005 and 2007 style evaluations. On the Morpho Challenge 2007 test evaluation, we report gains over the current best unsupervised contestant for Turkish, where our technique shows a 2.5% absolute F -score improvement. Michael Tepper, Fei Xia 0004 |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2009 | Language ID in the Context of Harvesting Language Data off the Web
Fei Xia 0004, William D. Lewis, Hoifung Poon |
EACL | 1 |
| 2008 | Automatically Identifying Computationally Relevant Typological Features
William D. Lewis, Fei Xia 0004 |
IJCNLP | 2 |
| 2008 | A Hybrid Approach to the Induction of Underlying Morphology
Michael Tepper, Fei Xia 0004 |
IJCNLP | 2 |
| 2008 | Repurposing Theoretical Linguistic Data for Tool Development and Search
Fei Xia 0004, William D. Lewis |
IJCNLP | 1 |
| 2007 | Multilingual Structural Projection across Interlinear Text
Fei Xia 0004, William D. Lewis |
HLT-NAACL | 1 |
| 2005 | Automatically Generating Tree Adjoining Grammars from Abstract SpecificationsabstractThe paper describes a system that can automatically generate tree adjoining grammars from abstract specifications. Our system is based on the use of tree descriptions to specify a grammar by separately defining pieces of tree structure that encode independent syntactic principles. Various individual specifications are then combined to form the elementary trees of the grammar. The system enables efficient development and maintenance of a grammar, and also allows underlying linguistic constructions (such as wh-movement) to be expressed explicitly. We have carefully designed our system to be as language independent as possible and tested its performance by constructing both English and Chinese grammars, with significant reductions in grammar development time. Provably consistent abstract specifications for different languages also offer unique opportunities for investigating how languages relate to themselves and to each other. For instance, the impact of a linguistic structure such as wh-movement can be traced from its specification to the descriptions that it combines with, to its actual realization in trees. By focusing on syntactic properties at a higher level, our approach allowed a unique comparison of our English and Chinese grammars. Fei Xia 0004, Martha Palmer, K. Vijay-Shanker |
Comput. Intell. | 1 |
| 2005 | The Penn Chinese TreeBank: Phrase structure annotation of a large corpusabstractWith growing interest in Chinese Language Processing, numerous NLP tools (e.g., word segmenters, part-of-speech taggers, and parsers) for Chinese have been developed all over the world. However, since no large-scale bracketed corpora are available to the public, these tools are trained on corpora with different segmentation criteria, part-of-speech tagsets and bracketing guidelines, and therefore, comparisons are difficult. As a first step towards addressing this issue, we have been preparing a large bracketed corpus since late 1998. The first two installments of the corpus, 250 thousand words of data, fully segmented, POS-tagged and syntactically bracketed, have been released to the public via LDC ( www.ldc.upenn.edu ). In this paper, we discuss several Chinese linguistic issues and their implications for our treebanking efforts and how we address these issues when developing our annotation guidelines. We also describe our engineering strategies to improve speed while ensuring annotation quality. Nianwen Xue, Fei Xia 0004, Fu-Dong Chiou, Martha Palmer |
Nat. Lang. Eng. | 2 |
| 2001 | Automatically Extracting and Comparing Lexicalized Grammars for Different Languages
Fei Xia 0004, Chung-hye Han, Martha Palmer, Aravind K. Joshi |
IJCAI | 1 |
| 2000 | A Uniform Method of Grammar Extraction and Its ApplicationsabstractGrammars are core elements of many NLP applications. In this paper, we present a system that automatically extracts lexicalized grammars from annotated corpora. The data produced by this system have been used in several tasks, such as training NLP tools (such as Supertaggers) and estimating the coverage of hand-crafted grammars. We report experimental results on two of those tasks and compare our approaches with related work. Fei Xia 0004, Martha Palmer, Aravind K. Joshi |
EMNLP | 1 |
| 2000 | Developing Guidelines and Ensuring Consistency for Chinese Text Annotation
Fei Xia 0004, Martha Palmer, Nianwen Xue, Mary Ellen Okurowski, John Kovarik, Fu-Dong Chiou, Shizhe Huang, Tony Kroch, Mitchell P. Marcus |
LREC | 1 |
| 1997 | A Comparison of Head Transducers and Transfer for a Limited Domain Translation ApplicationabstractWe compare the effectiveness of two related machine translation models applied to the same limited-domain task. One is a transfer model with monolingual head automata for analysis and generation; the other is a direct transduction model based on bilingual head transducers. We conclude that the head transducer model is more effective according to measures of accuracy, computational requirements, model size, and development effort. Hiyan Alshawi, Adam L. Buchsbaum, Fei Xia 0004 |
ACL | 3 |