Yan Song 0003

dblp:09/1398-3 · DBLP profile ↗
← Back
59ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0002-2849-2962ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Feature Decomposition via Shared Low-Rank Matrix Recovery for CT Report Generation
abstract
Generating reports for medical images is an important task in medical automation that not only provides valuable objective diagnostic evidence but also alleviates the workload of radiologists. Many existing studies focus on chest X-rays that typically consist of one or a few images, where less attention is paid to other medical image types, such as computed tomography (CT) that contain a large number of continuous images. Many studies on CT report generation (CTRG) rely on convolutional networks or standard Transformers to model CT slice representation and combine them to obtain CT features, yet relatively little research has focused on subtle lesion features and volumetric continuity. In this paper, we propose shared low-rank matrix recovery (S-LMR) to decompose CT slices into shared anatomical patterns and lesion-focused features, together with continuous slice encoding (CSE) to explicitly model inter-slice continuity and capture progressive changes across adjacent slices, which are subsequently integrated with a large language model (LLM) for report generation. Specifically, the S-LMR separates the common patterns from the sparse lesion-focused features to highlight clinically significant information. Based on the outputs of S-LMR, CSE captures inter-slice relationships within a dedicated Transformer encoder and aligns the resulting visual features with textual information, thereby instructing the LLM to produce a CT report. Experiment results on benchmark datasets for CTRG show that our approach outperforms strong baselines and existing models, demonstrating state-of-the-art performance. Analyses further confirm that S-LMR and CSE effectively capture key evidence, leading to more accurate CTRG.
Yuanhe Tian, Yan Song 0003
IEEE Trans. Medical Imaging2
2026 Extractive Radiology Reporting With Memory-Based Cross-Modal Representations
abstract
Radiology report generation (RRG) produces detailed textual descriptions for radiographs, serving as a crucial task for medical analysis and diagnosis. Most existing RRG approaches naturally follow the multimodal text generation paradigm, where autoregressive models are utilized to perform token-by-token report generation and thus are potentially risky in generating invalid content while being limited in low information processing speed. Although advanced architectures, such as pre-trained models and large language models (LLMs), are applied for RRG and achieve good performance, they still face the aforementioned risk and speed limitation, especially that LLMs may introduce hallucinations. Consider that radiology reports are highly patternized, sentences in them convey specific meanings independently and are frequently reused, we propose a new extractive radiograph reporting (ERR) workflow and design a dedicated framework that efficiently and accurately extracts appropriate sentences from existing radiological cases for report generation. Our approach employs a memory module to store important medical information and enhance the encoding for input radiograph with better cross-modal representations, which are used to match sentences for the extraction process. We conducted experiments on two widely used benchmark datasets, with the results demonstrating that our approach outperforms strong baselines and achieves comparable results with existing state-of-the-art generative models. Analyses further confirm that our ERR approach not only produces reports with reliable content but also ensures high training and inference efficiency.
Yuanhe Tian, Zexuan Yan, Nenan Lyu, Yan Song 0003
IEEE Trans. Medical Imaging4
2026 Multimodal Aspect-Based Sentiment Analysis With Plugin-Enhanced Large Language Models
abstract
Multimodal aspect-based sentiment analysis (MABSA) is a challenging task that predicts sentiment polarity for specific aspect terms based on inputs across modalities. Existing approaches typically employ advanced visual and textual encoders to extract multimodal features and align them for MABSA prediction, yet they still face challenges in handling complex connections between multiple modalities. Recent bloom of large language models (LLMs), as well as their multimodal counterparts, has shown significant promise in various tasks, which offer a promising solution for MABSA, with potential limitations such as semantic mismatch between images and texts, and their high computational cost of fine-tuning for specific tasks. To address these limitations, in this article, we propose a novel plugin-based approach for MABSA, which uses plugins to encode key knowledge instances, such as salient objects in images and word relationships in texts, with an attentive graph convolutional network (A-GCN). We further utilize a memory-based hub to integrate the encoded multimodal knowledge and align the knowledge representations with the LLM, guiding it to better understand the intricate connections between modalities. We evaluate our approach on two benchmark MABSA datasets, which outperforms baselines and achieves state-of-the-art performance over existing studies. Further analysis shows that our approach enables efficient and scalable adaptation of multimodal LLMs to specific tasks, making it a promising solution for related tasks. The code is available at https://github.com/synlp/MABSA-LLMPlug.
Yuanhe Tian, Yan Song 0003, Yongdong Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Learning Macroeconomic Policies Through Dynamic Stackelberg Mean-Field Games
abstract
Macroeconomic outcomes emerge from individuals’ decisions, making it essential to model how agents interact with macro policy via consumption, investment, and labor choices. We formulate this as a dynamic Stackelberg game: the government (leader) sets policies, and agents (followers) respond by optimizing their behavior over time. Unlike static models, this dynamic formulation captures temporal dependencies and strategic feedback critical to policy design. However, as the number of agents increases, explicitly simulating all agent–agent and agent–government interactions becomes computationally infeasible. To address this, we propose the Dynamic Stackelberg Mean Field Game (DSMFG) framework, which approximates these complex interactions via agent–population and government–population couplings. This approximation preserves individual-level feedback while ensuring scalability, enabling DSMFG to jointly model three core features of real-world policy-making: dynamic feedback, asymmetry, and large-scale. We further introduce Stackelberg Mean Field Reinforcement Learning (SMFRL), a data-driven algorithm that learns the leader’s optimal policies while maintaining personalized responses for individual agents. Empirically, we validate our approach in a large-scale simulated economy, where it scales to 1,000 agents (vs. 100 in prior work) and achieves a 4× GDP gain over classical economic methods and a 19× improvement over the static 2022 U.S. federal income tax policy.
Qirui Mi, Chengdong Ma, Si-Yu Xia, Yan Song 0003, Mengyue Yang, Jun Wang 0012, Haifeng Zhang 0002
ECAI5
2025 ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
abstract
Evaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to robustly evaluate the reasoning capability of LLMs. ThinkBench proposes a dynamic data generation method for constructing out-of-distribution (OOD) datasets and offers an OOD dataset that contains 2,912 samples drawn from reasoning tasks. ThinkBench unifies the evaluation of reasoning models and non-reasoning models. We evaluate 16 LLMs and 4 PRMs under identical experimental conditions and show that most of the LLMs' performance are far from robust and they face a certain level of data leakage. By dynamically generating OOD datasets, ThinkBench effectively provides a reliable evaluation of LLMs and reduces data contamination impact. Our data and codes are available at https://github.com/huangshulin123/ThinkBench.
Shulin Huang, Linyi Yang, Yan Song 0003, Shawn Chen, Leyang Cui, Ziyu Wan, Qingcheng Zeng, Ying Wen 0001, Kun Shao, Weinan Zhang 0001, Jun Wang 0012, Yue Zhang 0004
NeurIPS3
2025 ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement Learning
abstract
Recent research on Reasoning of Large Language Models (LLMs) has sought to further enhance their performance by integrating meta-thinking—enabling models to monitor, evaluate, and control their reasoning processes for more adaptive and effective problem-solving. However, current single-agent work lacks a specialized design for acquiring meta-thinking, resulting in low efficacy. To address this challenge, we introduce Reinforced Meta-thinking Agents (ReMA), a novel framework that leverages Multi-Agent Reinforcement Learning (MARL) to elicit meta-thinking behaviors, encouraging LLMs to think about thinking. ReMA decouples the reasoning process into two hierarchical agents: a high-level meta-thinking agent responsible for generating strategic oversight and plans, and a low-level reasoning agent for detailed executions. Through iterative reinforcement learning with aligned objectives, these agents explore and learn collaboration, leading to improved generalization and robustness. Empirical results from single-turn experiments demonstrate that ReMA outperforms single-agent RL baselines on complex reasoning tasks, including competitive-level mathematical benchmarks and LLM-as-a-Judge benchmarks. Additionally, we further extend ReMA to multi-turn interaction settings, leveraging turn-level ratio and parameter sharing to improve efficiency. Comprehensive ablation studies further illustrate the evolving dynamics of each distinct agent, providing valuable insights into how the meta-thinking reasoning process enhances the reasoning capabilities of LLMs.
Ziyu Wan, Xiaoyu Wen 0001, Yan Song 0003, Hanjing Wang, Linyi Yang, Mark Schmidt 0001, Jun Wang 0012, Weinan Zhang 0001, Shuyue Hu, Ying Wen 0001
NeurIPS4
2025 A Bilevel Reinforcement Learning Framework with Language Prior Knowledge
Yan Song 0003, Filippos Christianos, David Mguni
ECML/PKDD (6)2
2024 Diffusion Networks with Task-Specific Noise Control for Radiology Report Generation
abstract
Existing radiology report generation (RRG) studies mostly adopt autoregressive (AR) approaches to produce textual descriptions token-by-token for specific clinical radiographs, where they are susceptible to error propagation problems if irrelevant contents are half-way generated, leading to potential ill-presenting of precise diagnoses, especially when there exist complicated abnormalities in radiographs. Although the non-AR paradigm, e.g., diffusion model, provides an alternative solution to tackle the problem from AR by generating all contents in parallel, the mechanism of using Gaussian noise in existing diffusion models still has significant room to improve when such models are used in particular circumstances, i.e., providing proper guidance in controlling noises in the diffusive process to ensure precise report generation. In this paper, we propose to conduct RRG with diffusion networks by controlling the noise with task-specific features, which leverages irrelevant visual and textual information as noise rather than the stochastic Gaussian noise, and allows the diffusion networks to filter particular information through iterative denoising, thus performing a precise and controlled report generation process. Experiments on IU X-Ray and MIMIC-CXR demonstrate the superiority of our approach compared to strong baselines and state-of-the-art solutions. Human evaluation and noise type analysis show that comprehensive noise control greatly helps diffusion networks to refine the generation of global and local report contents.
Yuanhe Tian, Fei Xia 0004, Yan Song 0003
ACM Multimedia3
2023 Hashtag-Guided Low-Resource Tweet Classification
abstract
Social media classification tasks (e.g., tweet sentiment analysis, tweet stance detection) are challenging because social media posts are typically short, informal, and ambiguous. Thus, training on tweets is challenging and demands large-scale human-annotated labels, which are time-consuming and costly to obtain. In this paper, we find that providing hashtags to social media tweets can help alleviate this issue because hashtags can enrich short and ambiguous tweets in terms of various information, such as topic, sentiment, and stance. This motivates us to propose a novel Hashtag-guided Tweet Classification model (HashTation), which automatically generates meaningful hashtags for the input tweet to provide useful auxiliary signals for tweet classification. To generate high-quality and insightful hashtags, our hashtag generation model retrieves and encodes the post-level and entity-level information across the whole corpus. Experiments show that HashTation achieves significant improvements on seven low-resource tweet classification tasks, in which only a limited amount of training data is provided, showing that automatically enriching tweets with model-generated hashtags could significantly reduce the demand for large-scale human-labeled data. Further analysis demonstrates that HashTation is able to generate high-quality hashtags that are consistent with the tweets and their labels. The code is available at https://github.com/shizhediao/HashTation.
Shizhe Diao, Sedrick Keh, Liangming Pan, Zhiliang Tian, Yan Song 0003, Tong Zhang 0001
WWW5
2022 Enhancing Structure-aware Encoder with Extremely Limited Data for Graph-based Dependency Parsing
abstract
Dependency parsing is an important fundamental natural language processing task which analyzes the syntactic structure of an input sentence by illustrating the syntactic relations between words. To improve dependency parsing, leveraging existing dependency parsers and extra data (e.g., through semi-supervised learning) has been demonstrated to be effective, even though the final parsers are trained on inaccurate (but massive) data. In this paper, we propose a frustratingly easy approach to improve graph-based dependency parsing, where a structure-aware encoder is pre-trained on auto-parsed data by predicting the word dependencies and then fine-tuned on gold dependency trees, which differs from the usual pre-training process that aims to predict the context words along dependency paths. Experimental results and analyses demonstrate the effectiveness and robustness of our approach to benefit from the data (even with noise) processed by different parsers, where our approach outperforms strong baselines under different settings with different dependency standards and model architectures used in pre-training and fine-tuning. More importantly, further analyses find that only 2K auto-parsed sentences are required to obtain improvement when pre-training vanilla BERT-large based parser without requiring extra parameters.
Yuanhe Tian, Yan Song 0003, Fei Xia 0004
COLING2
2022 Enhancing Relation Extraction via Adversarial Multi-task Learning
abstract
Relation extraction (RE) is a sub-field of information extraction, which aims to extract the relation between two given named entities (NEs) in a sentence and thus requires a good understanding of contextual information, especially the entities and their surrounding texts. However, limited attention is paid by most existing studies to re-modeling the given NEs and thus lead to inferior RE results when NEs are sometimes ambiguous. In this paper, we propose a RE model with two training stages, where adversarial multi-task learning is applied to the first training stage to explicitly recover the given NEs so as to enhance the main relation extractor, which is trained alone in the second stage. In doing so, the RE model is optimized by named entity recognition (NER) and thus obtains a detailed understanding of entity-aware context. We further propose the adversarial mechanism to enhance the process, which controls the effect of NER on the main relation extractor and allows the extractor to benefit from NER while keep focusing on RE rather than the entire multi-task learning. Experimental results on two English benchmark datasets for RE demonstrate the effectiveness of our approach, where state-of-the-art performance is observed on both datasets.
Han Qin, Yuanhe Tian, Yan Song 0003
LREC3
2022 Complementary Learning of Aspect Terms for Aspect-based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) aims to predict the sentiment polarity towards a given aspect term in a sentence on the fine-grained level, which usually requires a good understanding of contextual information, especially appropriately distinguishing of a given aspect and its contexts, to achieve good performance. However, most existing ABSA models pay limited attention to the modeling of the given aspect terms and thus result in inferior results when a sentence contains multiple aspect terms with contradictory sentiment polarities. In this paper, we propose to improve ABSA by complementary learning of aspect terms, which serves as a supportive auxiliary task to enhance ABSA by explicitly recovering the aspect terms from each input sentence so as to better understand aspects and their contexts. Particularly, a discriminator is also introduced to further improve the learning process by appropriately balancing the impact of aspect recovery to sentiment prediction. Experimental results on five widely used English benchmark datasets for ABSA demonstrate the effectiveness of our approach, where state-of-the-art performance is observed on all datasets.
Han Qin, Yuanhe Tian, Fei Xia 0004, Yan Song 0003
LREC4
2022 ChiMST: A Chinese Medical Corpus for Word Segmentation and Medical Term Recognition
abstract
Chinese word segmentation (CWS) and named entity recognition (NER) are two important tasks in Chinese natural language processing. To achieve good model performance on these tasks, existing neural approaches normally require a large amount of labeled training data, which is often unavailable for specific domains such as the Chinese medical domain due to privacy and legal issues. To address this problem, we have developed a Chinese medical corpus named ChiMST which consists of question-answer pairs collected from an online medical healthcare platform and is annotated with word boundary and medical term information. For word boundary, we mainly follow the word segmentation guidelines for the Penn Chinese Treebank (Xia, 2000); for medical terms, we define 9 categories and 18 sub-categories after consulting medical experts. To provide baselines on this corpus, we train existing state-of-the-art models on it and achieve good performance. We believe that the corpus and the baseline systems will be a valuable resource for CWS and NER research on the medical domain.
Yuanhe Tian, Han Qin, Fei Xia 0004, Yan Song 0003
LREC4
2022 Syntax-driven Approach for Semantic Role Labeling
abstract
As an important task to analyze the semantic structure of a sentence, semantic role labeling (SRL) aims to locate the semantic role (e.g., agent) of noun phrases with respect to a given predicate and thus plays an important role in downstream tasks such as dialogue systems. To achieve a better performance in SRL, a model is always required to have a good understanding of the context information. Although one can use advanced text encoder (e.g., BERT) to capture the context information, extra resources are also required to further improve the model performance. Considering that there are correlations between the syntactic structure and the semantic structure of the sentence, many previous studies leverage auto-generated syntactic knowledge, especially the dependencies, to enhance the modeling of context information through graph-based architectures, where limited attention is paid to other types of auto-generated knowledge. In this paper, we propose map memories to enhance SRL by encoding different types of auto-generated syntactic knowledge (i.e., POS tags, syntactic constituencies, and word dependencies) obtained from off-the-shelf toolkits. Experimental results on two English benchmark datasets for span-style SRL (i.e., CoNLL-2005 and CoNLL-2012) demonstrate the effectiveness of our approach, which outperforms strong baselines and achieves state-of-the-art results on CoNLL-2005.
Yuanhe Tian, Han Qin, Fei Xia 0004, Yan Song 0003
LREC4
2021 Cross-modal Memory Networks for Radiology Report Generation
abstract
Zhihong Chen, Yaling Shen, Yan Song, Xiang Wan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yaling Shen, Yan Song 0003
ACL/IJCNLP (1)3
2021 Taming Pre-trained Language Models with N-gram Representations for Low-Resource Domain Adaptation
abstract
Shizhe Diao, Ruijia Xu, Hongjin Su, Yilei Jiang, Yan Song, Tong Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Shizhe Diao, Ruijia Xu, Hongjin Su, Yilei Jiang, Yan Song 0003, Tong Zhang 0001
ACL/IJCNLP (1)5
2021 Dependency-driven Relation Extraction with Attentive Graph Convolutional Networks
abstract
Yuanhe Tian, Guimin Chen, Yan Song, Xiang Wan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yuanhe Tian, Guimin Chen, Yan Song 0003
ACL/IJCNLP (1)3
2021 Discovering Protagonist of Sentiment with Aspect Reconstructed Capsule Network
Guoxin Yu, Min Yang 0007, Xiting Wang, Yan Song 0003, Xiang Ao 0001
DASFAA (2)6
2021 Enhancing Aspect-level Sentiment Analysis with Word Dependencies
abstract
Aspect-level sentiment analysis (ASA) has received much attention in recent years.Most existing approaches tried to leverage syntactic information, such as the dependency parsing results of the input text, to improve sentiment analysis on different aspects.Although these approaches achieved satisfying results, their main focus is to leverage the dependency arcs among words where the dependency type information is omitted; and they model different dependencies equally where the noisy dependency results may hurt model performance.In this paper, we propose an approach to enhance aspect-level sentiment analysis with word dependencies, where the type information is modeled by key-value memory networks and different dependency results are selectively leveraged.Experimental results on five benchmark datasets demonstrate the effectiveness of our approach, where it outperforms baseline models on all datasets and achieves state-of-the-art performance on three of them. 1 * Equal contribution.
Yuanhe Tian, Guimin Chen, Yan Song 0003
EACL3
2021 Improving Federated Learning for Aspect-based Sentiment Analysis via Topic Memories
abstract
Aspect-based sentiment analysis (ABSA) predicts the sentiment polarity towards a particular aspect term in a sentence, which is an important task in real-world applications.To perform ABSA, the trained model is required to have a good understanding of the contextual information, especially the particular patterns that suggest the sentiment polarity.However, these patterns typically vary in different sentences, especially when the sentences come from different sources (domains), which makes ABSA still very challenging.Although combining labeled data across different sources (domains) is a promising solution to address the challenge, in practical applications, these labeled data are usually stored at different locations and might be inaccessible to each other due to privacy or legal concerns (e.g., the data are owned by different companies).To address this issue and make the best use of all labeled data, we propose a novel ABSA model with federated learning (FL) adopted to overcome the data isolation limitations and incorporate topic memory (TM) proposed to take the cases of data from diverse sources (domains) into consideration.Particularly, TM aims to identify different isolated data sources due to data inaccessibility by providing useful categorical information for localized predictions.Experimental results on a simulated environment for FL with three nodes demonstrate the effectiveness of our approach, where TM-FL outperforms different baselines including some well-designed FL frameworks. 1 * Equal contribution.
Han Qin, Guimin Chen, Yuanhe Tian, Yan Song 0003
EMNLP (1)4
2021 Relation Extraction with Word Graphs from N-grams
abstract
Most recent studies for relation extraction (RE) leverage the dependency tree of the input sentence to incorporate syntax-driven contextual information to improve model performance, with little attention paid to the limitation where high-quality dependency parsers in most cases unavailable, especially for indomain scenarios.To address this limitation, in this paper, we propose attentive graph convolutional networks (A-GCN) to improve neural RE methods with an unsupervised manner to build the context graph, without relying on the existence of a dependency parser.Specifically, we construct the graph from n-grams extracted from a lexicon built from pointwise mutual information (PMI) and apply attention over the graph.Therefore, different word pairs from the contexts within and across n-grams are weighted in the model and facilitate RE accordingly.Experimental results with further analyses on two English benchmark datasets for RE demonstrate the effectiveness of our approach, where state-of-the-art performance is observed on both datasets.1
Han Qin, Yuanhe Tian, Yan Song 0003
EMNLP (1)3
2021 Aspect-based Sentiment Analysis with Type-aware Graph Convolutional Networks and Layer Ensemble
abstract
It is popular that neural graph-based models are applied in existing aspect-based sentiment analysis (ABSA) studies for utilizing word relations through dependency parses to facilitate the task with better semantic guidance for analyzing context and aspect words.However, most of these studies only leverage dependency relations without considering their dependency types, and are limited in lacking efficient mechanisms to distinguish the important relations as well as learn from different layers of graph based models.To address such limitations, in this paper, we propose an approach to explicitly utilize dependency types for ABSA with type-aware graph convolutional networks (T-GCN), where attention is used in T-GCN to distinguish different edges (relations) in the graph and attentive layer ensemble is proposed to comprehensively learn from different layers of T-GCN.The validity and effectiveness of our approach are demonstrated in the experimental results, where state-of-the-art performance is achieved on six English benchmark datasets.Further experiments are conducted to analyze the contributions of each component in our approach and illustrate how different layers in T-GCN help ABSA with quantitative and qualitative analysis.1
Yuanhe Tian, Guimin Chen, Yan Song 0003
NAACL-HLT3
2020 Coordinated Reasoning for Cross-Lingual Knowledge Graph Alignment
abstract
Existing entity alignment methods mainly vary on the choices of encoding the knowledge graph, but they typically use the same decoding method, which independently chooses the local optimal match for each source entity. This decoding method may not only cause the “many-to-one” problem but also neglect the coordinated nature of this task, that is, each alignment decision may highly correlate to the other decisions. In this paper, we introduce two coordinated reasoning methods, i.e., the Easy-to-Hard decoding strategy and joint entity alignment algorithm. Specifically, the Easy-to-Hard strategy first retrieves the model-confident alignments from the predicted results and then incorporates them as additional knowledge to resolve the remaining model-uncertain alignments. To achieve this, we further propose an enhanced alignment model that is built on the current state-of-the-art baseline. In addition, to address the many-to-one problem, we propose to jointly predict entity alignments so that the one-to-one constraint can be naturally incorporated into the alignment prediction. Experimental results show that our model achieves the state-of-the-art performance and our reasoning methods can also significantly improve existing baselines.
Kun Xu 0005, Linfeng Song, Yansong Feng 0002, Yan Song 0003, Dong Yu 0001
AAAI4
2020 Conditional Augmentation for Aspect Term Extraction via Masked Sequence-to-Sequence Generation
abstract
Aspect term extraction aims to extract aspect terms from review texts as opinion targets for sentiment analysis.One of the big challenges with this task is the lack of sufficient annotated data.While data augmentation is potentially an effective technique to address the above issue, it is uncontrollable as it may change aspect words and aspect labels unexpectedly.In this paper, we formulate the data augmentation as a conditional generation task: generating a new sentence while preserving the original opinion targets and labels.We propose a masked sequence-to-sequence method for conditional augmentation of aspect term extraction.Unlike existing augmentation approaches, ours is controllable and allows us to generate more diversified sentences.Experimental results confirm that our method alleviates the data scarcity problem significantly.It also effectively boosts the performances of several current models for aspect term extraction.
Kun Li 0003, Chengbo Chen, Xiaojun Quan, Yan Song 0003
ACL5
2020 Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-way Attentions of Auto-analyzed Knowledge
abstract
Chinese word segmentation (CWS) and partof-speech (POS) tagging are important fundamental tasks for Chinese language processing, where joint learning of them is an effective one-step solution for both tasks.Previous studies for joint CWS and POS tagging mainly follow the character-based tagging paradigm with introducing contextual information such as n-gram features or sentential representations from recurrent neural models.However, for many cases, the joint tagging needs not only modeling from context features but also knowledge attached to them (e.g., syntactic relations among words); limited efforts have been made by existing research to meet such needs.In this paper, we propose a neural model named TWASP for joint CWS and POS tagging following the character-based sequence labeling paradigm, where a two-way attention mechanism is used to incorporate both context feature and their corresponding syntactic knowledge for each input character.Particularly, we use existing language processing toolkits to obtain the auto-analyzed syntactic knowledge for the context, and the proposed attention module can learn and benefit from them although their quality may not be perfect.Our experiments illustrate the effectiveness of the two-way attentions for joint CWS and POS tagging, where state-of-the-art performance is achieved on five benchmark datasets.1
Yuanhe Tian, Yan Song 0003, Xiang Ao 0001, Fei Xia 0004, Xiaojun Quan, Tong Zhang 0001
ACL2
2020 Improving Chinese Word Segmentation with Wordhood Memory Networks
abstract
Contextual features always play an important role in Chinese word segmentation (CWS).Wordhood information, being one of the contextual features, is proved to be useful in many conventional character-based segmenters.However, this feature receives less attention in recent neural models and it is also challenging to design a framework that can properly integrate wordhood information from different wordhood measures to existing neural frameworks.In this paper, we therefore propose a neural framework, WMSEG, which uses memory networks to incorporate wordhood information with several popular encoder-decoder combinations for CWS.Experimental results on five benchmark datasets indicate the memory mechanism successfully models wordhood information for neural segmenters and helps WMSEG achieve state-ofthe-art performance on all those datasets.Further experiments and analyses also demonstrate the robustness of our proposed framework with respect to different wordhood measures and the efficiency of wordhood information in cross-domain experiments.1
Yuanhe Tian, Yan Song 0003, Fei Xia 0004, Tong Zhang 0001
ACL2
2020 Joint Aspect Extraction and Sentiment Analysis with Directional Graph Convolutional Networks
abstract
End-to-end aspect-based sentiment analysis (EASA) consists of two sub-tasks: the first extracts the aspect terms in a sentence and the second predicts the sentiment polarities for such terms.For EASA, compared to pipeline and multi-task approaches, joint aspect extraction and sentiment analysis provides a one-step solution to predict both aspect terms and their sentiment polarities through a single decoding process, which avoids the mismatches in between the results of aspect terms and sentiment polarities, as well as error propagation.Previous studies, especially recent ones, for this task focus on using powerful encoders (e.g., Bi-LSTM and BERT) to model contextual information from the input, with limited efforts paid to using advanced neural architectures (such as attentions and graph convolutional networks) or leveraging extra knowledge (such as syntactic information).To extend such efforts, in this paper, we propose directional graph convolutional networks (D-GCN) to jointly perform aspect extraction and sentiment analysis with encoding syntactic information, where dependency among words are integrated into our model to enhance its ability to represent input sentences and help EASA accordingly.Experimental results on three benchmark datasets demonstrate the effectiveness of our approach, where D-GCN achieves state-of-the-art performance on all datasets.1
Guimin Chen, Yuanhe Tian, Yan Song 0003
COLING3
2020 Meet Changes with Constancy: Learning Invariance in Multi-Source Translation
abstract
Multi-source neural machine translation aims to translate from parallel sources of information (e.g.languages, images, etc.) to a single target language, which has shown better performance than most one-to-one systems.Despite the remarkable success of existing models, they usually neglect the fact that multiple source inputs may have inconsistencies.Such differences might bring noise to the task and limit the performance of existing multi-source NMT approaches due to their indiscriminate usage of input sources for target word predictions.In this paper, we attempt to leverage the potential complementary information among distinct sources and alleviate the occasional conflicts of them.To accomplish that, we propose a source invariance network to learn the invariant information of parallel sources.Such network can be easily integrated with multi-encoder based multi-source NMT methods (e.g.multi-encoder RNN and transformer) to enhance the translation results.Extensive experiments on two multi-source translation tasks demonstrate that the proposed approach not only achieves clear gains in translation quality but also captures implicit invariance between different sources.
Xiang Ao 0001, Yan Song 0003
COLING4
2020 Summarizing Medical Conversations via Identifying Important Utterances
abstract
Summarization is an important natural language processing (NLP) task in identifying key information from text.For conversations, the summarization systems need to extract salient contents from spontaneous utterances by multiple speakers.In a special task-oriented scenario, namely medical conversations between patients and doctors, the symptoms, diagnoses, and treatments could be highly important because the nature of such conversation is to find a medical solution to the problem proposed by the patients.Especially consider that current online medical platforms provide millions of public available conversations between real patients and doctors, where the patients propose their medical problems and the registered doctors offer diagnosis and treatment, a conversation in most cases could be too long and the key information is hard to be located.Therefore, summarizations to the patients' problems and the doctors' treatments in the conversations can be highly useful, in terms of helping other patients with similar problems have a precise reference for potential medical solutions.In this paper, we focus on medical conversation summarization, using a dataset of medical conversations and corresponding summaries which were crawled from a well-known online healthcare service provider in China.We propose a hierarchical encoder-tagger model (HET) to generate summaries by identifying important utterances (with respect to problem proposing and solving) in the conversations.For the particular dataset used in this study, we show that high-quality summaries can be generated by extracting two types of utterances, namely, problem statements and treatment recommendations.Experimental results demonstrate that HET outperforms strong baselines and models from previous studies, and adding conversation-related features can further improve system performance.1 * Equal contribution. 1 Our code, models, and the dataset are released at https://github.com/cuhksz-nlp/HET-MC. 2 E.g., in China, the number of outpatient visits exceeded 7 billions and inpatient visits
Yan Song 0003, Yuanhe Tian, Fei Xia 0004
COLING1
2020 Joint Chinese Word Segmentation and Part-of-speech Tagging via Multi-channel Attention of Character N-grams
abstract
Chinese word segmentation (CWS) and part-of-speech (POS) tagging are two fundamental tasks for Chinese language processing.Previous studies have demonstrated that jointly performing them can be an effective one-step solution to both tasks and this joint task can benefit from a good modeling of contextual features such as n-grams.However, their work on modeling such contextual features is limited to concatenating the features or their embeddings directly with the input embeddings without distinguishing whether the contextual features are important for the joint task in the specific context.Therefore, their models for the joint task could be misled by unimportant contextual information.In this paper, we propose a character-based neural model for the joint task enhanced by multi-channel attention of n-grams.In the attention module, n-gram features are categorized into different groups according to several criteria, and n-grams in each group are weighted and distinguished according to their importance for the joint task in the specific context.To categorize n-grams, we try two criteria in this study, i.e., n-gram frequency and length, so that n-grams having different capabilities of carrying contextual information are discriminatively learned by our proposed attention module.Experimental results on five benchmark datasets for CWS and POS tagging demonstrate that our approach outperforms strong baseline models and achieves state-of-the-art performance on all five datasets.1
Yuanhe Tian, Yan Song 0003, Fei Xia 0004
COLING2
2020 Generating Radiology Reports via Memory-driven Transformer
abstract
Medical imaging is frequently used in clinical practice and trials for diagnosis and treatment.Writing imaging reports is time-consuming and can be error-prone for inexperienced radiologists.Therefore, automatically generating radiology reports is highly desired to lighten the workload of radiologists and accordingly promote clinical automation, which is an essential task to apply artificial intelligence to the medical domain.In this paper, we propose to generate radiology reports with memorydriven Transformer, where a relational memory is designed to record key information of the generation process and a memory-driven conditional layer normalization is applied to incorporating the memory into the decoder of Transformer.Experimental results on two prevailing radiology report datasets, IU X-Ray and MIMIC-CXR, show that our proposed approach outperforms previous models with respect to both language generation metrics and clinical evaluations.Particularly, this is the first work reporting the generation results on MIMIC-CXR to the best of our knowledge.Further analyses also demonstrate that our approach is able to generate long reports with necessary medical terms as well as meaningful image-text attention mappings.1
Yan Song 0003, Tsung-Hui Chang
EMNLP (1)2
2020 Named Entity Recognition for Social Media Texts with Semantic Augmentation
abstract
Existing approaches for named entity recognition suffer from data sparsity problems when conducted on short and informal texts, especially user-generated social media content.Semantic augmentation is a potential way to alleviate this problem.Given that rich semantic information is implicitly preserved in pre-trained word embeddings, they are potential ideal resources for semantic augmentation.In this paper, we propose a neural-based approach to NER for social media texts where both local (from running text) and augmented semantics are taken into account.In particular, we obtain the augmented semantic information from a large-scale corpus, and propose an attentive semantic augmentation module and a gate module to encode and aggregate such information, respectively.Extensive experiments are performed on three benchmark datasets collected from English and Chinese social media platforms, where the results demonstrate the superiority of our approach to previous studies across all three datasets.1 * Equal contribution.
Yuyang Nie, Yuanhe Tian, Yan Song 0003, Bo Dai 0006
EMNLP (1)4
2020 Supertagging Combinatory Categorial Grammar with Attentive Graph Convolutional Networks
abstract
Supertagging is conventionally regarded as an important task for combinatory categorial grammar (CCG) parsing, where effective modeling of contextual information is highly important to this task.However, existing studies have made limited efforts to leverage contextual features except for applying powerful encoders (e.g., bi-LSTM).In this paper, we propose attentive graph convolutional networks to enhance neural CCG supertagging through a novel solution of leveraging contextual information.Specifically, we build the graph from chunks (n-grams) extracted from a lexicon and apply attention over the graph, so that different word pairs from the contexts within and across chunks are weighted in the model and facilitate the supertagging accordingly.The experiments performed on the CCGbank demonstrate that our approach outperforms all previous studies in terms of both supertagging and parsing.Further analyses illustrate the effectiveness of each component in our approach to discriminatively learn from word pairs to enhance CCG supertagging. 1
Yuanhe Tian, Yan Song 0003, Fei Xia 0004
EMNLP (1)2
2020 Improving biomedical named entity recognition with syntactic information
abstract
BACKGROUND: Biomedical named entity recognition (BioNER) is an important task for understanding biomedical texts, which can be challenging due to the lack of large-scale labeled training data and domain knowledge. To address the challenge, in addition to using powerful encoders (e.g., biLSTM and BioBERT), one possible method is to leverage extra knowledge that is easy to obtain. Previous studies have shown that auto-processed syntactic information can be a useful resource to improve model performance, but their approaches are limited to directly concatenating the embeddings of syntactic information to the input word embeddings. Therefore, such syntactic information is leveraged in an inflexible way, where inaccurate one may hurt model performance. RESULTS: In this paper, we propose BIOKMNER, a BioNER model for biomedical texts with key-value memory networks (KVMN) to incorporate auto-processed syntactic information. We evaluate BIOKMNER on six English biomedical datasets, where our method with KVMN outperforms the strong baseline method, namely, BioBERT, from the previous study on all datasets. Specifically, the F1 scores of our best performing model are 85.29% on BC2GM, 77.83% on JNLPBA, 94.22% on BC5CDR-chemical, 90.08% on NCBI-disease, 89.24% on LINNAEUS, and 76.33% on Species-800, where state-of-the-art performance is obtained on four of them (i.e., BC2GM, BC5CDR-chemical, NCBI-disease, and Species-800). CONCLUSION: The experimental results on six English benchmark datasets demonstrate that auto-processed syntactic information can be a useful resource for BioNER and our method with KVMN can appropriately leverage such information to improve model performance.
Yuanhe Tian, Wang Shen, Yan Song 0003, Fei Xia 0004, Kenli Li 0001
BMC Bioinform.3
2019 Reinforced Training Data Selection for Domain Adaptation
abstract
Supervised models suffer from the problem of domain shifting where distribution mismatch in the data across domains greatly affect model performance.To solve the problem, training data selection (TDS) has been proven to be a prospective solution for domain adaptation in leveraging appropriate data.However, conventional TDS methods normally requires a predefined threshold which is neither easy to set nor can be applied across tasks, and models are trained separately with the TDS process.To make TDS self-adapted to data and task, and to combine it with model training, in this paper, we propose a reinforcement learning (RL) framework that synchronously searches for training instances relevant to the target domain and learns better representations for them.A selection distribution generator (SDG) is designed to perform the selection and is updated according to the rewards computed from the selected data, where a predictor is included in the framework to ensure a taskspecific model can be trained on the selected data and provides feedback to rewards.Experimental results from part-of-speech tagging, dependency parsing, and sentiment analysis, as well as ablation studies, illustrate that the proposed framework is not only effective in data selection and representation, but also generalized to accommodate different NLP tasks.
Miaofeng Liu, Yan Song 0003, Hongbin Zou, Tong Zhang 0001
ACL (1)2
2019 Cross-lingual Knowledge Graph Alignment via Graph Matching Neural Network
abstract
Previous cross-lingual knowledge graph (KG) alignment studies rely on entity embeddings derived only from monolingual KG structural information, which may fail at matching entities that have different facts in two KGs. In this paper, we introduce the topic entity graph, a local sub-graph of an entity, to represent entities with their contextual information in KG. From this view, the KB-alignment task can be formulated as a graph matching problem; and we further propose a graph-attention based solution, which first matches all entities in two topic entity graphs, and then jointly model the local matching information to derive a graph-level matching vector. Experiments show that our model outperforms previous state-of-the-art methods by a large margin.
Kun Xu 0005, Liwei Wang 0009, Mo Yu, Yansong Feng 0002, Yan Song 0003, Zhiguo Wang 0006, Dong Yu 0001
ACL (1)5
2019 Knowledge-aware Pronoun Coreference Resolution
abstract
Resolving pronoun coreference requires knowledge support, especially for particular domains (e.g., medicine).In this paper, we explore how to leverage different types of knowledge to better resolve pronoun coreference with a neural model.To ensure the generalization ability of our model, we directly incorporate knowledge in the format of triplets, which is the most common format of modern knowledge graphs, instead of encoding it with features or rules as that in conventional approaches.Moreover, since not all knowledge is helpful in certain contexts, to selectively use them, we propose a knowledge attention module, which learns to select and use informative knowledge based on contexts, to enhance our model.Experimental results on two datasets from different domains prove the validity and effectiveness of our model, where it outperforms state-of-the-art baselines by a large margin.Moreover, since our model learns to use external knowledge rather than only fitting the training data, it also demonstrates superior performance to baselines in the cross-domain setting.
Hongming Zhang 0009, Yan Song 0003, Yangqiu Song, Dong Yu 0001
ACL (1)2
2019 Reading Like HER: Human Reading Inspired Extractive Summarization
abstract
Ling Luo, Xiang Ao, Yan Song, Feiyang Pan, Min Yang, Qing He. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xiang Ao 0001, Yan Song 0003, Feiyang Pan, Min Yang 0007, Qing He 0003
EMNLP/IJCNLP (1)3
2019 What You See is What You Get: Visual Pronoun Coreference Resolution in Dialogues
abstract
Xintong Yu, Hongming Zhang, Yangqiu Song, Yan Song, Changshui Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xintong Yu 0002, Hongming Zhang 0009, Yangqiu Song, Yan Song 0003, Changshui Zhang
EMNLP/IJCNLP (1)4
2019 Multiplex Word Embeddings for Selectional Preference Acquisition
abstract
Hongming Zhang, Jiaxin Bai, Yan Song, Kun Xu, Changlong Yu, Yangqiu Song, Wilfred Ng, Dong Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Hongming Zhang 0009, Jiaxin Bai, Yan Song 0003, Kun Xu 0005, Changlong Yu, Yangqiu Song, Wilfred Ng, Dong Yu 0001
EMNLP/IJCNLP (1)3
2019 Unsupervised Neural Aspect Extraction with Sememes
abstract
Aspect extraction relies on identifying aspects by discovering coherence among words, which is challenging when word meanings are diversified and processing on short texts. To enhance the performance on aspect extraction, leveraging lexical semantic resources is a possible solution to such challenge. In this paper, we present an unsupervised neural framework that leverages sememes to enhance lexical semantics. The overall framework is analogous to an autoenoder which reconstructs sentence representations and learns aspects by latent variables. Two models that form sentence representations are proposed by exploiting sememes via (1) a hierarchical attention; (2) a context-enhanced attention. Experiments on two real-world datasets demonstrate the validity and the effectiveness of our models, which significantly outperforms existing baselines.
Xiang Ao 0001, Yan Song 0003, Jinyao Li, Qing He 0003, Dong Yu 0001
IJCAI3
2018 hyperdoc2vec: Distributed Representations of Hypertext Documents
abstract
Hypertext documents, such as web pages and academic papers, are of great importance in delivering information in our daily life.Although being effective on plain documents, conventional text embedding methods suffer from information loss if directly adapted to hyper-documents.In this paper, we propose a general embedding approach for hyper-documents, namely, hyperdoc2vec, along with four criteria characterizing necessary information that hyper-document embedding models should preserve.Systematic comparisons are conducted between hyperdoc2vec and several competitors on two tasks, i.e., paper classification and citation recommendation, in the academic paper domain.Analyses and experiments both validate the superiority of hyperdoc2vec to other models w.r.t. the four criteria.
Jialong Han, Yan Song 0003, Wayne Xin Zhao, Shuming Shi 0001, Haisong Zhang
ACL (1)2
2018 Iterative Document Representation Learning Towards Summarization with Polishing
abstract
In this paper, we introduce Iterative Text Summarization (ITS), an iteration-based model for supervised extractive text summarization, inspired by the observation that it is often necessary for a human to read an article multiple times in order to fully understand and summarize its contents.Current summarization approaches read through a document only once to generate a document representation, resulting in a sub-optimal representation.To address this issue we introduce a model which iteratively polishes the document representation on many passes through the document.As part of our model, we also introduce a selective reading mechanism that decides more accurately the extent to which each sentence in the model should be updated.Experimental results on the CNN/DailyMail and DUC2002 datasets demonstrate that our model significantly outperforms state-of-the-art extractive systems when evaluated by machines and by humans.
Xiuying Chen, Shen Gao, Chongyang Tao, Yan Song 0003, Dongyan Zhao 0001, Rui Yan 0001
EMNLP4
2018 Generating Classical Chinese Poems via Conditional Variational Autoencoder and Adversarial Training
abstract
It is a challenging task to automatically compose poems with not only fluent expressions but also aesthetic wording.Although much attention has been paid to this task and promising progress is made, there exist notable gaps between automatically generated ones with those created by humans, especially on the aspects of term novelty and thematic consistency.Towards filling the gap, in this paper, we propose a conditional variational autoencoder with adversarial training for classical Chinese poem generation, where the autoencoder part generates poems with novel terms and a discriminator is applied to adversarially learn their thematic consistency with their titles.Experimental results on a large poetry corpus confirm the validity and effectiveness of our model, where its automatic and human evaluation scores outperform existing models.
Juntao Li 0005, Yan Song 0003, Haisong Zhang, Dongmin Chen, Shuming Shi 0001, Dongyan Zhao 0001, Rui Yan 0001
EMNLP2
2018 A Hybrid Approach to Automatic Corpus Generation for Chinese Spelling Check
abstract
Chinese spelling check (CSC) is a challenging yet meaningful task, which not only serves as a preprocessing in many natural language processing (NLP) applications, but also facilitates reading and understanding of running texts in peoples' daily lives.However, to utilize datadriven approaches for CSC, there is one major limitation that annotated corpora are not enough in applying algorithms and building models.In this paper, we propose a novel approach of constructing CSC corpus with automatically generated spelling errors, which are either visually or phonologically resembled characters, corresponding to the OCRand ASR-based methods, respectively.Upon the constructed corpus, different models are trained and evaluated for CSC with respect to three standard test sets.Experimental results demonstrate the effectiveness of the corpus, therefore confirm the validity of our approach.* This work was conducted during Dingmin Wang's internship in Tencent AI Lab. SentenceCorrection
Dingmin Wang, Yan Song 0003, Jing Li 0049, Jialong Han, Haisong Zhang
EMNLP2
2018 Topic Memory Networks for Short Text Classification
abstract
Many classification models work poorly on short texts due to data sparsity.To address this issue, we propose topic memory networks for short text classification with a novel topic memory mechanism to encode latent topic representations indicative of class labels.Different from most prior work that focuses on extending features with external knowledge or pre-trained topics, our model jointly explores topic inference and text classification with memory networks in an end-to-end manner.Experimental results on four benchmark datasets show that our model outperforms state-of-the-art models on short text classification, meanwhile generates coherent topics.* This work was mainly conducted when Jichuan Zeng was an intern in Tencent AI Lab.† Jing Li is the corresponding author.Training instances R1: [SuperBowl] I'll do anything to see the Steelers win.R2: [New.Music.Live] Please give wristbands, she have major Bieber Fever.
Jichuan Zeng, Jing Li 0049, Yan Song 0003, Cuiyun Gao 0001, Michael R. Lyu, Irwin King
EMNLP3
2018 Complementary Learning of Word Embeddings
abstract
Continuous bag-of-words (CB) and skip-gram (SG) models are popular approaches to training word embeddings. Conventionally they are two standing-alone techniques used individually. However, with the same goal of building embeddings by leveraging surrounding words, they are in fact a pair of complementary tasks where the output of one model can be used as input of the other, and vice versa. In this paper, we propose complementary learning of word embeddings based on the CB and SG model. Specifically, one round of learning first integrates the predicted output of a SG model with existing context, then forms an enlarged context as input to the CB model. Final models are obtained through several rounds of parameter updating. Experimental results indicate that our approach can effectively improve the quality of initial embeddings, in terms of intrinsic and extrinsic evaluations.
Yan Song 0003, Shuming Shi 0001
IJCAI1
2018 Joint Learning Embeddings for Chinese Words and their Components via Ladder Structured Networks
abstract
The components, such as characters and radicals, of a Chinese word are important sources to help in capturing semantic information of the word. In this paper, we propose a novel framework, namely, ladder structured networks (LSN), which contains three layers representing word, character and radical and learns their embeddings synchronously. LSN captures not only the relations among words, but also the relations among their component characters and radicals, as well as the relations across layers. Each layer in LSN is pluggable so that any particular type of unit (word, character, radical) can be removed and the LSN is thus adjusted for particular types of inputs. In evaluating our framework, we use word similarity as the intrinsic evaluation and part-of-speech tagging and document classification as extrinsic evaluations. Experimental results confirm the validity of our approach and show superiority of our approach over previous work.
Yan Song 0003, Shuming Shi 0001, Jing Li 0049
IJCAI1
2018 Constructing a Chinese Medical Conversation Corpus Annotated with Conversational Structures and Actions
Yan Song 0003, Fei Xia 0004
LREC2
2018 Encoding Conversation Context for Neural Keyphrase Extraction from Microblog Posts
abstract
Yingyi Zhang, Jing Li, Yan Song, Chengzhi Zhang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Jing Li 0049, Yan Song 0003
NAACL-HLT3
2018 When Less Is More: Using Less Context Information to Generate Better Utterances in Group Conversations
Haisong Zhang, Zhangming Chan, Yan Song 0003, Dongyan Zhao 0001, Rui Yan 0001
NLPCC (1)3
2018 A Joint Model of Conversational Discourse and Latent Topics on Microblogs
abstract
Conventional topic models are ineffective for topic extraction from microblog messages, because the data sparseness exhibited in short messages lacking structure and contexts results in poor message-level word co-occurrence patterns. To address this issue, we organize microblog messages as conversation trees based on their reposting and replying relations, and propose an unsupervised model that jointly learns word distributions to represent: (1) different roles of conversational discourse, and (2) various latent topics in reflecting content information. By explicitly distinguishing the probabilities of messages with varying discourse roles in containing topical words, our model is able to discover clusters of discourse words that are indicative of topical content. In an automatic evaluation on large-scale microblog corpora, our joint model yields topics with better coherence scores than competitive topic models from previous studies. Qualitative analysis on model outputs indicates that our model induces meaningful representations for both discourse and topics. We further present an empirical study on microblog summarization based on the outputs of our joint model. The results show that the jointly modeled discourse and topic representations can effectively indicate summary-worthy content in microblog conversations.
Jing Li 0049, Yan Song 0003, Zhongyu Wei, Kam-Fai Wong
Comput. Linguistics2
2017 Learning Word Representations with Regularization from Prior Knowledge
abstract
Conventional word embeddings are trained with specific criteria (e.g., based on language modeling or co-occurrence) inside a single information source, disregarding the opportunity for further calibration using external knowledge.This paper presents a unified framework that leverages pre-learned or external priors, in the form of a regularizer, for enhancing conventional language model-based embedding learning.We consider two types of regularizers.The first type is derived from topic distribution by running latent Dirichlet allocation on unlabeled data.The second type is based on dictionaries that are created with human annotation efforts.To effectively learn with the regularizers, we propose a novel data structure, trajectory softmax, in this paper.The resulting embeddings are evaluated by word similarity and sentiment classification.Experimental results show that our learning framework with regularization from prior knowledge improves embedding quality across multiple datasets, compared to a diverse collection of baseline methods.
Yan Song 0003, Fei Xia 0004
CoNLL1
2014 Modern Chinese Helps Archaic Chinese Processing: Finding and Exploiting the Shared Properties
Yan Song 0003, Fei Xia 0004
LREC1
2013 Non-Monotonic Sentence Alignment via Semisupervised Learning
Xiaojun Quan, Chunyu Kit, Yan Song 0003
ACL (1)3
2013 A Common Case of Jekyll and Hyde: The Synergistic Effect of Using Divided Source Training Data for Feature Augmentation
Yan Song 0003, Fei Xia 0004
IJCNLP1
2012 Using a Goodness Measurement for Domain Adaptation: A Case Study on Chinese Word Segmentation
Yan Song 0003, Fei Xia 0004
LREC1
2010 How Large a Corpus Do We Need: Statistical Method Versus Rule-based Method
Hai Zhao 0001, Yan Song 0003, Chunyu Kit
LREC2
2009 Cross Language Dependency Parsing using a Bilingual Lexicon
Hai Zhao 0001, Yan Song 0003, Chunyu Kit, Guodong Zhou 0001
ACL/IJCNLP2