Yuanhe Tian

dblp:246/0133 · DBLP profile ↗
← Back
31ranked-venue papers
20as first author
23since 2021 · last 2026
0000-0001-6841-2341ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 15 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Feature Decomposition via Shared Low-Rank Matrix Recovery for CT Report Generation
abstract
Generating reports for medical images is an important task in medical automation that not only provides valuable objective diagnostic evidence but also alleviates the workload of radiologists. Many existing studies focus on chest X-rays that typically consist of one or a few images, where less attention is paid to other medical image types, such as computed tomography (CT) that contain a large number of continuous images. Many studies on CT report generation (CTRG) rely on convolutional networks or standard Transformers to model CT slice representation and combine them to obtain CT features, yet relatively little research has focused on subtle lesion features and volumetric continuity. In this paper, we propose shared low-rank matrix recovery (S-LMR) to decompose CT slices into shared anatomical patterns and lesion-focused features, together with continuous slice encoding (CSE) to explicitly model inter-slice continuity and capture progressive changes across adjacent slices, which are subsequently integrated with a large language model (LLM) for report generation. Specifically, the S-LMR separates the common patterns from the sparse lesion-focused features to highlight clinically significant information. Based on the outputs of S-LMR, CSE captures inter-slice relationships within a dedicated Transformer encoder and aligns the resulting visual features with textual information, thereby instructing the LLM to produce a CT report. Experiment results on benchmark datasets for CTRG show that our approach outperforms strong baselines and existing models, demonstrating state-of-the-art performance. Analyses further confirm that S-LMR and CSE effectively capture key evidence, leading to more accurate CTRG.
Yuanhe Tian, Yan Song 0003
IEEE Trans. Medical Imaging1
2026 Extractive Radiology Reporting With Memory-Based Cross-Modal Representations
abstract
Radiology report generation (RRG) produces detailed textual descriptions for radiographs, serving as a crucial task for medical analysis and diagnosis. Most existing RRG approaches naturally follow the multimodal text generation paradigm, where autoregressive models are utilized to perform token-by-token report generation and thus are potentially risky in generating invalid content while being limited in low information processing speed. Although advanced architectures, such as pre-trained models and large language models (LLMs), are applied for RRG and achieve good performance, they still face the aforementioned risk and speed limitation, especially that LLMs may introduce hallucinations. Consider that radiology reports are highly patternized, sentences in them convey specific meanings independently and are frequently reused, we propose a new extractive radiograph reporting (ERR) workflow and design a dedicated framework that efficiently and accurately extracts appropriate sentences from existing radiological cases for report generation. Our approach employs a memory module to store important medical information and enhance the encoding for input radiograph with better cross-modal representations, which are used to match sentences for the extraction process. We conducted experiments on two widely used benchmark datasets, with the results demonstrating that our approach outperforms strong baselines and achieves comparable results with existing state-of-the-art generative models. Analyses further confirm that our ERR approach not only produces reports with reliable content but also ensures high training and inference efficiency.
Yuanhe Tian, Zexuan Yan, Nenan Lyu, Yan Song 0003
IEEE Trans. Medical Imaging1
2026 Multimodal Aspect-Based Sentiment Analysis With Plugin-Enhanced Large Language Models
abstract
Multimodal aspect-based sentiment analysis (MABSA) is a challenging task that predicts sentiment polarity for specific aspect terms based on inputs across modalities. Existing approaches typically employ advanced visual and textual encoders to extract multimodal features and align them for MABSA prediction, yet they still face challenges in handling complex connections between multiple modalities. Recent bloom of large language models (LLMs), as well as their multimodal counterparts, has shown significant promise in various tasks, which offer a promising solution for MABSA, with potential limitations such as semantic mismatch between images and texts, and their high computational cost of fine-tuning for specific tasks. To address these limitations, in this article, we propose a novel plugin-based approach for MABSA, which uses plugins to encode key knowledge instances, such as salient objects in images and word relationships in texts, with an attentive graph convolutional network (A-GCN). We further utilize a memory-based hub to integrate the encoded multimodal knowledge and align the knowledge representations with the LLM, guiding it to better understand the intricate connections between modalities. We evaluate our approach on two benchmark MABSA datasets, which outperforms baselines and achieves state-of-the-art performance over existing studies. Further analysis shows that our approach enables efficient and scalable adaptation of multimodal LLMs to specific tasks, making it a promising solution for related tasks. The code is available at https://github.com/synlp/MABSA-LLMPlug.
Yuanhe Tian, Yan Song 0003, Yongdong Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation
abstract
Generating reports for computed tomography (CT) images is a challenging task. Although it is related to existing work on medical image report generation, it exhibits several unique characteristics, including the spatial encoding of multiple images and the alignment between image volumes and text. Existing solutions typically use general 2D or 3D image processing techniques to extract features from a CT volume, where they firstly compress the volume and then divide the compressed CT slices into patches for visual encoding. These approaches do not explicitly account for the transformations among CT slices, nor do they effectively integrate multi-level image features, particularly those containing specific organ lesions, to instruct CT report generation (CTRG). In considering the strong correlation among consecutive slices in CT scans, in this paper, we propose a large language model (LLM) based CTRG method with recurrent visual feature extraction and stereo attentions for hierarchical feature modeling. Specifically, we use a vision Transformer to recurrently process each slice in a CT volume, and employ a set of attentions over the encoded slices from different perspectives to selectively obtain important visual information and align it with textual features, so as to better instruct an LLM for CTRG. Experiment results and further analysis on the benchmark M3DCap dataset show that our method outperforms strong baseline models and achieves state-of-the-art results, demonstrating its validity and effectiveness11Code is available at https://github.com/synlp/SA-CTRG
Yuanhe Tian, Yan Song 0001
BIBM1
2024 Bootstrapping Large Language Models for Radiology Report Generation
abstract
Radiology report generation (RRG) aims to automatically generate a free-text description from a specific clinical radiograph, e.g., chest X-Ray images. Existing approaches tend to perform RRG with specific models trained on the public yet limited data from scratch, where they often lead to inferior performance owing to the problem of inefficient capabilities in both aligning visual and textual features and generating informative reports accordingly. Currently, large language models (LLMs) offered a promising solution to text generation with their power in learning from big data, especially for cross-modal scenarios such as RRG. However, most existing LLMs are pre-trained on general data, and suffer from the same problem of conventional approaches caused by knowledge gap between general and medical domain if they are applied to RRG. Therefore in this paper, we propose an approach to bootstrapping LLMs for RRG with a in-domain instance induction and a coarse-to-fine decoding process. Specifically, the in-domain instance induction process learns to align the LLM to radiology reports from general texts through contrastive learning. The coarse-to-fine decoding performs a text elevating process for those reports from the ranker, further enhanced with visual features and refinement prompts. Experimental results on two prevailing RRG datasets, namely, IU X-Ray and MIMIC-CXR, demonstrate the superiority of our approach to previous state-of-the-art solutions. Further analyses illustrate that, for the LLM, the induction process enables it to better align with the medical domain and the coarse-to-fine generation allows it to conduct more precise text generation.
Yuanhe Tian, Weidong Chen 0013, Yan Song 0004, Yongdong Zhang 0001
AAAI2
2024 ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences
abstract
Recently, the increasing demand for superior medical services has highlighted the discrepancies in the medical infrastructure.With big data, especially texts, forming the foundation of medical services, there is an exigent need for effective natural language processing (NLP) solutions tailored to the healthcare domain.Conventional approaches leveraging pre-trained models present promising results in this domain and current large language models (LLMs) offer advanced foundation for medical text processing.However, most medical LLMs are trained only with supervised fine-tuning (SFT), even though it efficiently empowers LLMs to understand and respond to medical instructions but is ineffective in learning domain knowledge and aligning with human preference.In this work, we propose CHIMED-GPT, a new benchmark LLM designed explicitly for Chinese medical domain, and undergoes a comprehensive training regime with pre-training, SFT, and RLHF.Evaluations on tasks including information extraction, question answering, and dialogue generation demonstrate CHIMED-GPT's superior performance over general domain LLMs.Furthermore, we analyze possible biases through prompting CHIMED-GPT to perform attitude scales regarding discrimination of patients, so as to contribute to further responsible development of LLMs in the medical domain. 1
Yuanhe Tian, Ruyi Gan, Yan Song 0004, Jiaxing Zhang 0001, Yongdong Zhang 0001
ACL (1)1
2024 Large Language Models Are No Longer Shallow Parsers
abstract
The development of large language models (LLMs) brings significant changes to the field of natural language processing (NLP), enabling remarkable performance in various high-level tasks, such as machine translation, questionanswering, dialogue generation, etc., under endto-end settings without requiring much training data.Meanwhile, fundamental NLP tasks, particularly syntactic parsing, are also essential for language study as well as evaluating the capability of LLMs for instruction understanding and usage.In this paper, we focus on analyzing and improving the capability of current state-of-the-art LLMs on a classic fundamental task, namely constituency parsing, which is the representative syntactic task in both linguistics and natural language processing.We observe that these LLMs are effective in shallow parsing but struggle with creating correct full parse trees.To improve the performance of LLMs on deep syntactic parsing, we propose a three-step approach that firstly prompts LLMs for chunking, then filters out low-quality chunks, and finally adds the remaining chunks to prompts to instruct LLMs for parsing, with later enhancement by chain-of-thought prompting.Experimental results on English and Chinese benchmark datasets demonstrate the effectiveness of our approach on improving LLMs' performance on constituency parsing.
Yuanhe Tian, Fei Xia 0004, Yan Song 0001
ACL (1)1
2024 Dialogue Summarization with Mixture of Experts based on Large Language Models
abstract
Dialogue summarization is an important task that requires to generate highlights for a conversation from different aspects (e.g., content of various speakers).While several studies successfully employ large language models (LLMs) and achieve satisfying results, they are limited by using one model at a time or treat it as a black box, which makes it hard to discriminatively learn essential content in a dialogue from different aspects, therefore may lead to anticipation bias and potential loss of information in the produced summaries.In this paper, we propose an LLM-based approach with roleoriented routing and fusion generation to utilize mixture of experts (MoE) for dialogue summarization.Specifically, the role-oriented routing is an LLM-based module that selects appropriate experts to process different information; fusion generation is another LLM-based module to locate salient information and produce finalized dialogue summaries.The proposed approach offers an alternative solution to employing multiple LLMs for dialogue summarization by leveraging their capabilities of in-context processing and generation in an effective manner.We run experiments on widely used benchmark datasets for this task, where the results demonstrate the superiority of our approach in producing informative and accurate dialogue summarization. 1
Yuanhe Tian, Fei Xia 0004, Yan Song 0001
ACL (1)1
2024 Improving Radiology Report Generation with D2-Net: When Diffusion Meets Discriminator
abstract
Radiology report generation (RRG) aims to automatically provide observations and insight into a patient’s condition based on radiology images, which is able to greatly reduce the workload of physicians on the premise of ensuring the quality of medical treatment. Existing works leverage the Transformer decoder to generate reports word-by-wordly. However, unlike image captioning, radiology reports are long text containing many semantic words. The autoregressive method, such as the Transformer-base method, will accumulate errors in the generation process and generate unsatisfied reports. Benefiting from the recent success of Diffusion, we propose a novel Diffusion-based paradigm for RRG, which leverages visual information as a condition, making the generation process focus on pathological features within the radiology image. Meanwhile, we integrate a discriminator into each layer of the Diffusion to actively judge whether the generated words are meaningful, which, on the one hand, controls the length of predicted reports and, on the other hand, calibrates confidence scores and token generation results, improving the quality of the generated reports. Extensive experiment results demonstrate the superiority of our proposed method. Source code is available at: https://github.com/Yuda-Jin/D-2-Net.
Yuda Jin, Weidong Chen 0013, Yuanhe Tian, Yan Song 0004, Chenggang Yan 0001, Zhendong Mao 0001
ICASSP3
2024 Diffusion Networks with Task-Specific Noise Control for Radiology Report Generation
abstract
Existing radiology report generation (RRG) studies mostly adopt autoregressive (AR) approaches to produce textual descriptions token-by-token for specific clinical radiographs, where they are susceptible to error propagation problems if irrelevant contents are half-way generated, leading to potential ill-presenting of precise diagnoses, especially when there exist complicated abnormalities in radiographs. Although the non-AR paradigm, e.g., diffusion model, provides an alternative solution to tackle the problem from AR by generating all contents in parallel, the mechanism of using Gaussian noise in existing diffusion models still has significant room to improve when such models are used in particular circumstances, i.e., providing proper guidance in controlling noises in the diffusive process to ensure precise report generation. In this paper, we propose to conduct RRG with diffusion networks by controlling the noise with task-specific features, which leverages irrelevant visual and textual information as noise rather than the stochastic Gaussian noise, and allows the diffusion networks to filter particular information through iterative denoising, thus performing a precise and controlled report generation process. Experiments on IU X-Ray and MIMIC-CXR demonstrate the superiority of our approach compared to strong baselines and state-of-the-art solutions. Human evaluation and noise type analysis show that comprehensive noise control greatly helps diffusion networks to refine the generation of global and local report contents.
Yuanhe Tian, Fei Xia 0004, Yan Song 0003
ACM Multimedia1
2024 Emotion Cause Extraction in Conversations with Response Graphing
Yuanhe Tian, Pengsen Cheng, Fei Xia 0004, Yongdong Zhang 0001, Yan Song 0004
NLPCC (5)1
2024 Improving radiology report generation with multi-grained abnormality prediction
Yuda Jin, Yuanhe Tian, Yan Song 0001
Neurocomputing3
2023 Improving Image Captioning via Predicting Structured Concepts
abstract
Having the difficulty of solving the semantic gap between images and texts for the image captioning task, conventional studies in this area paid some attention to treating semantic concepts as a bridge between the two modalities and improved captioning performance accordingly.Although promising results on concept prediction were obtained, the aforementioned studies normally ignore the relationship among concepts, which relies on not only objects in the image, but also word dependencies in the text, so that offers a considerable potential for improving the process of generating good descriptions.In this paper, we propose a structured concept predictor (SCP) to predict concepts and their structures, then we integrate them into captioning, so as to enhance the contribution of visual signals in this task via concepts and further use their relations to distinguish cross-modal semantics for better description generation.Particularly, we design weighted graph convolutional networks (W-GCN) to depict concept relations driven by word dependencies, and then learns differentiated contributions from these concepts for following decoding process.Therefore, our approach captures potential relations among concepts and discriminatively learns different concepts, so that effectively facilitates image captioning with inherited information across modalities.Extensive experiments and their results demonstrate the effectiveness of our approach as well as each proposed module in this work.Source code is available
Weidong Chen 0013, Yuanhe Tian, Yan Song 0004, Zhendong Mao 0001
EMNLP3
2022 Enhancing Structure-aware Encoder with Extremely Limited Data for Graph-based Dependency Parsing
abstract
Dependency parsing is an important fundamental natural language processing task which analyzes the syntactic structure of an input sentence by illustrating the syntactic relations between words. To improve dependency parsing, leveraging existing dependency parsers and extra data (e.g., through semi-supervised learning) has been demonstrated to be effective, even though the final parsers are trained on inaccurate (but massive) data. In this paper, we propose a frustratingly easy approach to improve graph-based dependency parsing, where a structure-aware encoder is pre-trained on auto-parsed data by predicting the word dependencies and then fine-tuned on gold dependency trees, which differs from the usual pre-training process that aims to predict the context words along dependency paths. Experimental results and analyses demonstrate the effectiveness and robustness of our approach to benefit from the data (even with noise) processed by different parsers, where our approach outperforms strong baselines under different settings with different dependency standards and model architectures used in pre-training and fine-tuning. More importantly, further analyses find that only 2K auto-parsed sentences are required to obtain improvement when pre-training vanilla BERT-large based parser without requiring extra parameters.
Yuanhe Tian, Yan Song 0003, Fei Xia 0004
COLING1
2022 Enhancing Relation Extraction via Adversarial Multi-task Learning
abstract
Relation extraction (RE) is a sub-field of information extraction, which aims to extract the relation between two given named entities (NEs) in a sentence and thus requires a good understanding of contextual information, especially the entities and their surrounding texts. However, limited attention is paid by most existing studies to re-modeling the given NEs and thus lead to inferior RE results when NEs are sometimes ambiguous. In this paper, we propose a RE model with two training stages, where adversarial multi-task learning is applied to the first training stage to explicitly recover the given NEs so as to enhance the main relation extractor, which is trained alone in the second stage. In doing so, the RE model is optimized by named entity recognition (NER) and thus obtains a detailed understanding of entity-aware context. We further propose the adversarial mechanism to enhance the process, which controls the effect of NER on the main relation extractor and allows the extractor to benefit from NER while keep focusing on RE rather than the entire multi-task learning. Experimental results on two English benchmark datasets for RE demonstrate the effectiveness of our approach, where state-of-the-art performance is observed on both datasets.
Han Qin, Yuanhe Tian, Yan Song 0003
LREC2
2022 Complementary Learning of Aspect Terms for Aspect-based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) aims to predict the sentiment polarity towards a given aspect term in a sentence on the fine-grained level, which usually requires a good understanding of contextual information, especially appropriately distinguishing of a given aspect and its contexts, to achieve good performance. However, most existing ABSA models pay limited attention to the modeling of the given aspect terms and thus result in inferior results when a sentence contains multiple aspect terms with contradictory sentiment polarities. In this paper, we propose to improve ABSA by complementary learning of aspect terms, which serves as a supportive auxiliary task to enhance ABSA by explicitly recovering the aspect terms from each input sentence so as to better understand aspects and their contexts. Particularly, a discriminator is also introduced to further improve the learning process by appropriately balancing the impact of aspect recovery to sentiment prediction. Experimental results on five widely used English benchmark datasets for ABSA demonstrate the effectiveness of our approach, where state-of-the-art performance is observed on all datasets.
Han Qin, Yuanhe Tian, Fei Xia 0004, Yan Song 0003
LREC2
2022 ChiMST: A Chinese Medical Corpus for Word Segmentation and Medical Term Recognition
abstract
Chinese word segmentation (CWS) and named entity recognition (NER) are two important tasks in Chinese natural language processing. To achieve good model performance on these tasks, existing neural approaches normally require a large amount of labeled training data, which is often unavailable for specific domains such as the Chinese medical domain due to privacy and legal issues. To address this problem, we have developed a Chinese medical corpus named ChiMST which consists of question-answer pairs collected from an online medical healthcare platform and is annotated with word boundary and medical term information. For word boundary, we mainly follow the word segmentation guidelines for the Penn Chinese Treebank (Xia, 2000); for medical terms, we define 9 categories and 18 sub-categories after consulting medical experts. To provide baselines on this corpus, we train existing state-of-the-art models on it and achieve good performance. We believe that the corpus and the baseline systems will be a valuable resource for CWS and NER research on the medical domain.
Yuanhe Tian, Han Qin, Fei Xia 0004, Yan Song 0003
LREC1
2022 Syntax-driven Approach for Semantic Role Labeling
abstract
As an important task to analyze the semantic structure of a sentence, semantic role labeling (SRL) aims to locate the semantic role (e.g., agent) of noun phrases with respect to a given predicate and thus plays an important role in downstream tasks such as dialogue systems. To achieve a better performance in SRL, a model is always required to have a good understanding of the context information. Although one can use advanced text encoder (e.g., BERT) to capture the context information, extra resources are also required to further improve the model performance. Considering that there are correlations between the syntactic structure and the semantic structure of the sentence, many previous studies leverage auto-generated syntactic knowledge, especially the dependencies, to enhance the modeling of context information through graph-based architectures, where limited attention is paid to other types of auto-generated knowledge. In this paper, we propose map memories to enhance SRL by encoding different types of auto-generated syntactic knowledge (i.e., POS tags, syntactic constituencies, and word dependencies) obtained from off-the-shelf toolkits. Experimental results on two English benchmark datasets for span-style SRL (i.e., CoNLL-2005 and CoNLL-2012) demonstrate the effectiveness of our approach, which outperforms strong baselines and achieves state-of-the-art results on CoNLL-2005.
Yuanhe Tian, Han Qin, Fei Xia 0004, Yan Song 0003
LREC1
2021 Dependency-driven Relation Extraction with Attentive Graph Convolutional Networks
abstract
Yuanhe Tian, Guimin Chen, Yan Song, Xiang Wan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yuanhe Tian, Guimin Chen, Yan Song 0003
ACL/IJCNLP (1)1
2021 Enhancing Aspect-level Sentiment Analysis with Word Dependencies
abstract
Aspect-level sentiment analysis (ASA) has received much attention in recent years.Most existing approaches tried to leverage syntactic information, such as the dependency parsing results of the input text, to improve sentiment analysis on different aspects.Although these approaches achieved satisfying results, their main focus is to leverage the dependency arcs among words where the dependency type information is omitted; and they model different dependencies equally where the noisy dependency results may hurt model performance.In this paper, we propose an approach to enhance aspect-level sentiment analysis with word dependencies, where the type information is modeled by key-value memory networks and different dependency results are selectively leveraged.Experimental results on five benchmark datasets demonstrate the effectiveness of our approach, where it outperforms baseline models on all datasets and achieves state-of-the-art performance on three of them. 1 * Equal contribution.
Yuanhe Tian, Guimin Chen, Yan Song 0003
EACL1
2021 Improving Federated Learning for Aspect-based Sentiment Analysis via Topic Memories
abstract
Aspect-based sentiment analysis (ABSA) predicts the sentiment polarity towards a particular aspect term in a sentence, which is an important task in real-world applications.To perform ABSA, the trained model is required to have a good understanding of the contextual information, especially the particular patterns that suggest the sentiment polarity.However, these patterns typically vary in different sentences, especially when the sentences come from different sources (domains), which makes ABSA still very challenging.Although combining labeled data across different sources (domains) is a promising solution to address the challenge, in practical applications, these labeled data are usually stored at different locations and might be inaccessible to each other due to privacy or legal concerns (e.g., the data are owned by different companies).To address this issue and make the best use of all labeled data, we propose a novel ABSA model with federated learning (FL) adopted to overcome the data isolation limitations and incorporate topic memory (TM) proposed to take the cases of data from diverse sources (domains) into consideration.Particularly, TM aims to identify different isolated data sources due to data inaccessibility by providing useful categorical information for localized predictions.Experimental results on a simulated environment for FL with three nodes demonstrate the effectiveness of our approach, where TM-FL outperforms different baselines including some well-designed FL frameworks. 1 * Equal contribution.
Han Qin, Guimin Chen, Yuanhe Tian, Yan Song 0003
EMNLP (1)3
2021 Relation Extraction with Word Graphs from N-grams
abstract
Most recent studies for relation extraction (RE) leverage the dependency tree of the input sentence to incorporate syntax-driven contextual information to improve model performance, with little attention paid to the limitation where high-quality dependency parsers in most cases unavailable, especially for indomain scenarios.To address this limitation, in this paper, we propose attentive graph convolutional networks (A-GCN) to improve neural RE methods with an unsupervised manner to build the context graph, without relying on the existence of a dependency parser.Specifically, we construct the graph from n-grams extracted from a lexicon built from pointwise mutual information (PMI) and apply attention over the graph.Therefore, different word pairs from the contexts within and across n-grams are weighted in the model and facilitate RE accordingly.Experimental results with further analyses on two English benchmark datasets for RE demonstrate the effectiveness of our approach, where state-of-the-art performance is observed on both datasets.1
Han Qin, Yuanhe Tian, Yan Song 0003
EMNLP (1)2
2021 Aspect-based Sentiment Analysis with Type-aware Graph Convolutional Networks and Layer Ensemble
abstract
It is popular that neural graph-based models are applied in existing aspect-based sentiment analysis (ABSA) studies for utilizing word relations through dependency parses to facilitate the task with better semantic guidance for analyzing context and aspect words.However, most of these studies only leverage dependency relations without considering their dependency types, and are limited in lacking efficient mechanisms to distinguish the important relations as well as learn from different layers of graph based models.To address such limitations, in this paper, we propose an approach to explicitly utilize dependency types for ABSA with type-aware graph convolutional networks (T-GCN), where attention is used in T-GCN to distinguish different edges (relations) in the graph and attentive layer ensemble is proposed to comprehensively learn from different layers of T-GCN.The validity and effectiveness of our approach are demonstrated in the experimental results, where state-of-the-art performance is achieved on six English benchmark datasets.Further experiments are conducted to analyze the contributions of each component in our approach and illustrate how different layers in T-GCN help ABSA with quantitative and qualitative analysis.1
Yuanhe Tian, Guimin Chen, Yan Song 0003
NAACL-HLT1
2020 Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-way Attentions of Auto-analyzed Knowledge
abstract
Chinese word segmentation (CWS) and partof-speech (POS) tagging are important fundamental tasks for Chinese language processing, where joint learning of them is an effective one-step solution for both tasks.Previous studies for joint CWS and POS tagging mainly follow the character-based tagging paradigm with introducing contextual information such as n-gram features or sentential representations from recurrent neural models.However, for many cases, the joint tagging needs not only modeling from context features but also knowledge attached to them (e.g., syntactic relations among words); limited efforts have been made by existing research to meet such needs.In this paper, we propose a neural model named TWASP for joint CWS and POS tagging following the character-based sequence labeling paradigm, where a two-way attention mechanism is used to incorporate both context feature and their corresponding syntactic knowledge for each input character.Particularly, we use existing language processing toolkits to obtain the auto-analyzed syntactic knowledge for the context, and the proposed attention module can learn and benefit from them although their quality may not be perfect.Our experiments illustrate the effectiveness of the two-way attentions for joint CWS and POS tagging, where state-of-the-art performance is achieved on five benchmark datasets.1
Yuanhe Tian, Yan Song 0003, Xiang Ao 0001, Fei Xia 0004, Xiaojun Quan, Tong Zhang 0001
ACL1
2020 Improving Chinese Word Segmentation with Wordhood Memory Networks
abstract
Contextual features always play an important role in Chinese word segmentation (CWS).Wordhood information, being one of the contextual features, is proved to be useful in many conventional character-based segmenters.However, this feature receives less attention in recent neural models and it is also challenging to design a framework that can properly integrate wordhood information from different wordhood measures to existing neural frameworks.In this paper, we therefore propose a neural framework, WMSEG, which uses memory networks to incorporate wordhood information with several popular encoder-decoder combinations for CWS.Experimental results on five benchmark datasets indicate the memory mechanism successfully models wordhood information for neural segmenters and helps WMSEG achieve state-ofthe-art performance on all those datasets.Further experiments and analyses also demonstrate the robustness of our proposed framework with respect to different wordhood measures and the efficiency of wordhood information in cross-domain experiments.1
Yuanhe Tian, Yan Song 0003, Fei Xia 0004, Tong Zhang 0001
ACL1
2020 Joint Aspect Extraction and Sentiment Analysis with Directional Graph Convolutional Networks
abstract
End-to-end aspect-based sentiment analysis (EASA) consists of two sub-tasks: the first extracts the aspect terms in a sentence and the second predicts the sentiment polarities for such terms.For EASA, compared to pipeline and multi-task approaches, joint aspect extraction and sentiment analysis provides a one-step solution to predict both aspect terms and their sentiment polarities through a single decoding process, which avoids the mismatches in between the results of aspect terms and sentiment polarities, as well as error propagation.Previous studies, especially recent ones, for this task focus on using powerful encoders (e.g., Bi-LSTM and BERT) to model contextual information from the input, with limited efforts paid to using advanced neural architectures (such as attentions and graph convolutional networks) or leveraging extra knowledge (such as syntactic information).To extend such efforts, in this paper, we propose directional graph convolutional networks (D-GCN) to jointly perform aspect extraction and sentiment analysis with encoding syntactic information, where dependency among words are integrated into our model to enhance its ability to represent input sentences and help EASA accordingly.Experimental results on three benchmark datasets demonstrate the effectiveness of our approach, where D-GCN achieves state-of-the-art performance on all datasets.1
Guimin Chen, Yuanhe Tian, Yan Song 0003
COLING2
2020 Summarizing Medical Conversations via Identifying Important Utterances
abstract
Summarization is an important natural language processing (NLP) task in identifying key information from text.For conversations, the summarization systems need to extract salient contents from spontaneous utterances by multiple speakers.In a special task-oriented scenario, namely medical conversations between patients and doctors, the symptoms, diagnoses, and treatments could be highly important because the nature of such conversation is to find a medical solution to the problem proposed by the patients.Especially consider that current online medical platforms provide millions of public available conversations between real patients and doctors, where the patients propose their medical problems and the registered doctors offer diagnosis and treatment, a conversation in most cases could be too long and the key information is hard to be located.Therefore, summarizations to the patients' problems and the doctors' treatments in the conversations can be highly useful, in terms of helping other patients with similar problems have a precise reference for potential medical solutions.In this paper, we focus on medical conversation summarization, using a dataset of medical conversations and corresponding summaries which were crawled from a well-known online healthcare service provider in China.We propose a hierarchical encoder-tagger model (HET) to generate summaries by identifying important utterances (with respect to problem proposing and solving) in the conversations.For the particular dataset used in this study, we show that high-quality summaries can be generated by extracting two types of utterances, namely, problem statements and treatment recommendations.Experimental results demonstrate that HET outperforms strong baselines and models from previous studies, and adding conversation-related features can further improve system performance.1 * Equal contribution. 1 Our code, models, and the dataset are released at https://github.com/cuhksz-nlp/HET-MC. 2 E.g., in China, the number of outpatient visits exceeded 7 billions and inpatient visits
Yan Song 0003, Yuanhe Tian, Fei Xia 0004
COLING2
2020 Joint Chinese Word Segmentation and Part-of-speech Tagging via Multi-channel Attention of Character N-grams
abstract
Chinese word segmentation (CWS) and part-of-speech (POS) tagging are two fundamental tasks for Chinese language processing.Previous studies have demonstrated that jointly performing them can be an effective one-step solution to both tasks and this joint task can benefit from a good modeling of contextual features such as n-grams.However, their work on modeling such contextual features is limited to concatenating the features or their embeddings directly with the input embeddings without distinguishing whether the contextual features are important for the joint task in the specific context.Therefore, their models for the joint task could be misled by unimportant contextual information.In this paper, we propose a character-based neural model for the joint task enhanced by multi-channel attention of n-grams.In the attention module, n-gram features are categorized into different groups according to several criteria, and n-grams in each group are weighted and distinguished according to their importance for the joint task in the specific context.To categorize n-grams, we try two criteria in this study, i.e., n-gram frequency and length, so that n-grams having different capabilities of carrying contextual information are discriminatively learned by our proposed attention module.Experimental results on five benchmark datasets for CWS and POS tagging demonstrate that our approach outperforms strong baseline models and achieves state-of-the-art performance on all five datasets.1
Yuanhe Tian, Yan Song 0003, Fei Xia 0004
COLING1
2020 Named Entity Recognition for Social Media Texts with Semantic Augmentation
abstract
Existing approaches for named entity recognition suffer from data sparsity problems when conducted on short and informal texts, especially user-generated social media content.Semantic augmentation is a potential way to alleviate this problem.Given that rich semantic information is implicitly preserved in pre-trained word embeddings, they are potential ideal resources for semantic augmentation.In this paper, we propose a neural-based approach to NER for social media texts where both local (from running text) and augmented semantics are taken into account.In particular, we obtain the augmented semantic information from a large-scale corpus, and propose an attentive semantic augmentation module and a gate module to encode and aggregate such information, respectively.Extensive experiments are performed on three benchmark datasets collected from English and Chinese social media platforms, where the results demonstrate the superiority of our approach to previous studies across all three datasets.1 * Equal contribution.
Yuyang Nie, Yuanhe Tian, Yan Song 0003, Bo Dai 0006
EMNLP (1)2
2020 Supertagging Combinatory Categorial Grammar with Attentive Graph Convolutional Networks
abstract
Supertagging is conventionally regarded as an important task for combinatory categorial grammar (CCG) parsing, where effective modeling of contextual information is highly important to this task.However, existing studies have made limited efforts to leverage contextual features except for applying powerful encoders (e.g., bi-LSTM).In this paper, we propose attentive graph convolutional networks to enhance neural CCG supertagging through a novel solution of leveraging contextual information.Specifically, we build the graph from chunks (n-grams) extracted from a lexicon and apply attention over the graph, so that different word pairs from the contexts within and across chunks are weighted in the model and facilitate the supertagging accordingly.The experiments performed on the CCGbank demonstrate that our approach outperforms all previous studies in terms of both supertagging and parsing.Further analyses illustrate the effectiveness of each component in our approach to discriminatively learn from word pairs to enhance CCG supertagging. 1
Yuanhe Tian, Yan Song 0003, Fei Xia 0004
EMNLP (1)1
2020 Improving biomedical named entity recognition with syntactic information
abstract
BACKGROUND: Biomedical named entity recognition (BioNER) is an important task for understanding biomedical texts, which can be challenging due to the lack of large-scale labeled training data and domain knowledge. To address the challenge, in addition to using powerful encoders (e.g., biLSTM and BioBERT), one possible method is to leverage extra knowledge that is easy to obtain. Previous studies have shown that auto-processed syntactic information can be a useful resource to improve model performance, but their approaches are limited to directly concatenating the embeddings of syntactic information to the input word embeddings. Therefore, such syntactic information is leveraged in an inflexible way, where inaccurate one may hurt model performance. RESULTS: In this paper, we propose BIOKMNER, a BioNER model for biomedical texts with key-value memory networks (KVMN) to incorporate auto-processed syntactic information. We evaluate BIOKMNER on six English biomedical datasets, where our method with KVMN outperforms the strong baseline method, namely, BioBERT, from the previous study on all datasets. Specifically, the F1 scores of our best performing model are 85.29% on BC2GM, 77.83% on JNLPBA, 94.22% on BC5CDR-chemical, 90.08% on NCBI-disease, 89.24% on LINNAEUS, and 76.33% on Species-800, where state-of-the-art performance is obtained on four of them (i.e., BC2GM, BC5CDR-chemical, NCBI-disease, and Species-800). CONCLUSION: The experimental results on six English benchmark datasets demonstrate that auto-processed syntactic information can be a useful resource for BioNER and our method with KVMN can appropriately leverage such information to improve model performance.
Yuanhe Tian, Wang Shen, Yan Song 0003, Fei Xia 0004, Kenli Li 0001
BMC Bioinform.1