Sujian Li

dblp:05/4288 · DBLP profile ↗
← Back
119ranked-venue papers
6as first author
38since 2021 · last 2026
0000-0001-7493-0786ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 104 · 3 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 since 2021Databases, data management, data science and information retrieval · 15 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 DocLens: A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
abstract
Comprehending long visual documents, where information is distributed across extensive pages of text and visual elements, is a critical but challenging task for modern Vision-Language Models (VLMs).Existing approaches falter on a fundamental challenge: evidence localization.They struggle to retrieve relevant pages and overlook fine-grained details within visual elements, leading to limited performance and model hallucination.To address this, we propose DOCLENS, a toolaugmented multi-agent framework that effectively "zooms in" on evidence like a lens.It first navigates from the full document to specific visual elements on relevant pages, then employs a sampling-adjudication mechanism to generate a single, reliable answer.Paired with Gemini-2.5-Pro,DOCLENS achieves stateof-the-art performance on MMLongBench-Doc and FinRAGBench-V, surpassing even human experts.The framework's superiority is particularly evident on vision-centric and unanswerable queries, demonstrating the power of its enhanced localization capabilities.
Jiefeng Chen 0001, Sujian Li, Tomas Pfister, Jinsung Yoon
ACL (1)4
2025 ISR: Self-Refining Referring Expressions for Entity Grounding
abstract
Entity grounding, a crucial task in constructing multimodal knowledge graphs, aims to align entities from knowledge graphs with their corresponding images.Unlike conventional visual grounding tasks that use referring expressions (REs) as inputs, entity grounding relies solely on entity names and types, presenting a significant challenge.To address this, we introduce a novel Iterative Self-Refinement (ISR) scheme to enhance the multimodal large language model's capability to generate high quality REs for the given entities as explicit contextual clues.This training scheme, inspired by human learning dynamics and human annotation processes, enables the MLLM to iteratively generate and refine REs by learning from successes and failures, guided by outcome rewards from a visual grounding model.This iterative cycle of self-refinement avoids overfitting to fixed annotations and fosters continued improvement in referring expression generation.Extensive experiments demonstrate that our methods surpasses other methods in entity grounding, highlighting its effectiveness, robustness and potential for broader applications 1 .
Zhuocheng Yu, Bingchan Zhao, Yifan Song 0002, Sujian Li, Zhonghui He
ACL (1)4
2025 Hierarchical Memory Organization for Wikipedia Generation
abstract
Eugene J. Yu, Dawei Zhu, Yifan Song, Xiangyu Wong, Jiebin Zhang, Wenxuan Shi, Xiaoguang Li, Qun Liu, Sujian Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Eugene J. Yu, Yifan Song 0002, Xiangyu Wong, Jiebin Zhang, Qun Liu 0001, Sujian Li
ACL (1)9
2025 EERPD: Leveraging Emotion and Emotion Regulation for Improving Personality Detection
abstract
Personality is a fundamental construct in psychology, reflecting an individual’s behavior, thinking, and emotional patterns. While previous researches have made progress in personality detection, their designed methods generally overlook the important connection between psychological knowledge “emotion regulation” and personality traits. Based on this, we propose a new personality detection method called EERPD. This method introduces the use of emotion regulation, a psychological concept highly correlated with personality, for personality prediction. By combining this concept with emotion features, EERPD retrieves few-shot examples and provides process CoTs for inferring labels from text. This approach enhances the understanding of LLM for personality implicit within text and improves the performance in personality detection. Experimental results demonstrate that EERPD significantly enhances the accuracy and robustness of personality detection, outperforming previous SOTA by 15.05/4.29 in average F1 on the two benchmark datasets.
Sujian Li, Qilong Ma, Weimin Xiong
COLING2
2025 WIKIGENBENCH: Exploring Full-length Wikipedia Generation under Real-World Scenario
abstract
It presents significant challenges to generate comprehensive and accurate Wikipedia articles for newly emerging events under real-world scenario. Existing attempts fall short either by focusing only on short snippets or by using metrics that are insufficient to evaluate real-world scenarios. In this paper, we construct WIKIGENBENCH, a new benchmark consisting of 1,320 entries, designed to align with real-world scenarios in both generation and evaluation. For generation, we explore a real-world scenario where structured, full-length Wikipedia articles with citations are generated for new events using input documents from web sources. For evaluation, we integrate systematic metrics and LLM-based metrics to assess the verifiability, organization, and other aspects aligned with real-world scenarios. Based on this benchmark, we conduct extensive experiments using various models within three commonly used frameworks: direct RAG, hierarchical structure-based RAG, and RAG with fine-tuned generation model. Experimental results show that hierarchical-based methods can generate more comprehensive content, while fine-tuned methods achieve better verifiability. However, even the best methods still show a significant gap compared to existing Wikipedia content, indicating that further research is necessary.
Jiebin Zhang, Eugene J. Yu, Qinyu Chen, Chenhao Xiong, Han Qian, Mingbo Song, Weimin Xiong, Qun Liu 0001, Sujian Li
COLING11
2025 Exploring Fine-Grained Human Motion Video Captioning
abstract
Detailed descriptions of human motion are crucial for effective fitness training, which highlights the importance of research in fine-grained human motion video captioning. Existing video captioning models often fail to capture the nuanced semantics of videos, resulting in the generated descriptions that are coarse and lack details, especially when depicting human motions. To benchmark the Body Fitness Training scenario, in this paper, we construct a fine-grained human motion video captioning dataset named BoFiT and design a state-of-the-art baseline model named BoFiT-Gen (Body Fitness Training Text Generation). BoFiT-Gen makes use of computer vision techniques to extract angular representations of human motions from videos and LLMs to generate fine-grained descriptions of human motions via prompting. Results show that BoFiT-Gen outperforms previous methods on comprehensive metrics. We aim for this dataset to serve as a useful evaluation set for visio-linguistic models and drive further progress in this field. Our dataset is released at https://github.com/colmon46/bofit.
Bingchan Zhao, Zhuocheng Yu, Tongchen Yang, Yifan Song 0002, Mingyu Jin, Sujian Li
COLING7
2025 VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
abstract
Vision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference labels from traditional VL tasks, which can introduce biases and often fail to effectively challenge state-of-the-art models. To address these limitations, we introduce VL-RewardBench, a comprehensive benchmark spanning general multimodal queries, visual hallucination detection, and complex reasoning tasks. Through our AI-assisted annotation pipeline that combines sample selection with human verification, we curate 1,250 high-quality examples specifically designed to probe VL-GenRMs limitations. Comprehensive evaluation across 16 leading large vision-language models demonstrates VL-RewardBench’s effectiveness as a challenging testbed, where even GPT-4o achieves only 65.4% accuracy, and state-of-the-art open-source models such as Qwen2-VL-72B, struggle to surpass random-guessing. Importantly, performance on VL-RewardBench strongly correlates (Pearson’s r > 0.9) with MMMU-Pro accuracy using Best-of-N sampling with VL-GenRMs. Analysis experiments uncover three critical insights for improving VL-GenRMs: (i) models predominantly fail at basic visual perception tasks rather than reasoning tasks; (ii) inference-time scaling benefits vary dramatically by model capacity; and (iii) training VL-GenRMs to learn to judge substantially boosts judgment capability (+14.7% accuracy for a 7B VL-GenRM). We believe VL-RewardBench along with the experimental insights will become a valuable resource for advancing VL-GenRMs. Project page: https://vl-rewardbench.github.io.
Lei Li 0039, Yuancheng Wei, Zhihui Xie 0002, Xuqing Yang, Yifan Song 0002, Peiyi Wang, Chenxin An, Tianyu Liu 0001, Sujian Li, Bill Y. Lin, Lingpeng Kong, Qi Liu 0049
CVPR9
2025 Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
abstract
The Mixture-of-Experts (MoE) models have gained significant attention in deep learning due to their dynamic resource allocation and superior performance across diverse tasks. However, efficiently training these models remains challenging. The MoE upcycling technique has been proposed to reuse and improve existing model components, thereby minimizing training overhead. Despite this, simple routers, such as linear routers, often struggle with complex routing tasks within MoE upcycling. In response, we propose a novel routing technique called Router Upcycling to enhance the performance of MoE upcycling models. Our approach initializes multiple routers from the attention heads of preceding attention layers during upcycling. These routers collaboratively assign tokens to specialized experts in an attention-like manner. Each token is processed into diverse queries and aligned with the experts’ features (serving as keys). Experimental results demonstrate that our method achieves state-of-the-art (SOTA) performance, outperforming other upcycling baselines.
Junfeng Ran, Guangxiang Zhao, Yuhan Wu 0001, Longyun Wu, Yikai Zhao 0001, Tong Yang 0003, Lin Sun 0010, Xiangzheng Zhang, Sujian Li
ECAI10
2025 Realistic Training Data Generation and Rule Enhanced Decoding in LLM for NameGuess
abstract
The wide use of abbreviated column names (derived from English words or Chinese Pinyin) in database tables poses significant challenges for table-centric tasks in natural language processing and database management.Such a column name expansion task, referred to as the NameGuess task, has previously been addressed by fine-tuning Large Language Models (LLMs) on synthetically generated rule-based data.However, the current approaches yield suboptimal performance due to two fundamental limitations: 1) the rule-generated abbreviation data fails to reflect real-world distribution, and 2) the failure of LLMs to follow the rulesensitive patterns in NameGuess persistently.For the data realism issue, we propose a novel approach that integrates a subsequence abbreviation generator trained on human-annotated data and collects non-subsequence abbreviations to improve the training set.For the rule violation issue, we propose a decoding system constrained on an automaton that represents the rules of abbreviation expansion.We extended the original English NameGuess test set to include non-subsequence and PinYin scenarios.Experimental results show that properly tuned 7/8B moderate-size LLMs with a refined decoding system can surpass the few-shot performance of state-of-the-art LLMs, such as the GPT-4 series.The code and data are presented in the supplementary material.
Yikuan Xia, Jiazun Chen, Sujian Li, Jun Gao 0003
EMNLP3
2025 FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain
abstract
Retrieval-Augmented Generation (RAG) plays a vital role in the financial domain, powering applications such as real-time market analysis, trend forecasting, and interest rate computation. However, most existing RAG research in finance focuses predominantly on textual data, overlooking the rich visual content in financial documents, resulting in the loss of key analytical insights. To bridge this gap, we present FinRAGBench-V, a comprehensive visual RAG benchmark tailored for finance. This benchmark effectively integrates multimodal data and provides visual citation to ensure traceability. It includes a bilingual retrieval corpus with 60,780 Chinese and 51,219 English pages, along with a high-quality, human-annotated question-answering (QA) dataset spanning heterogeneous data types and seven question categories. Moreover, we introduce RGenCite, an RAG baseline that seamlessly integrates visual citation with generation. Furthermore, we propose an automatic citation evaluation method to systematically assess the visual citation capabilities of Multimodal Large Language Models (MLLMs). Extensive experiments on RGenCite underscore the challenging nature of FinRAGBench-V, providing valuable insights for the development of multimodal RAG systems in finance.
Suifeng Zhao, Zhuoran Jin, Sujian Li, Jun Gao 0003
EMNLP3
2025 The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
abstract
Yifan Song, Guoyin Wang, Sujian Li, Bill Yuchen Lin. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yifan Song 0002, Sujian Li, Bill Y. Lin
NAACL (Long Papers)3
2024 Trial and Error: Exploration-Based Trajectory Optimization of LLM Agents
abstract
Large Language Models (LLMs) have become integral components in various autonomous agent systems.In this study, we present an exploration-based trajectory optimization approach, referred to as ETO.This learning method is designed to enhance the performance of open LLM agents.Contrary to previous studies that exclusively train on successful expert trajectories, our method allows agents to learn from their exploration failures.This leads to improved performance through an iterative optimization framework.During the exploration phase, the agent interacts with the environment while completing given tasks, gathering failure trajectories to create contrastive trajectory pairs.In the subsequent training phase, the agent utilizes these trajectory preference pairs to update its policy using contrastive learning methods like DPO (Rafailov et al., 2023).This iterative cycle of exploration and training fosters continued improvement in the agents.Our experiments on three complex tasks demonstrate that ETO consistently surpasses baseline performance by a large margin.Furthermore, an examination of task-solving efficiency and potential in scenarios lacking expert trajectory underscores the effectiveness of our approach.1
Yifan Song 0002, Da Yin, Xiang Yue, Sujian Li, Bill Y. Lin
ACL (1)5
2024 FaGANet: An Evidence-Based Fact-Checking Model with Integrated Encoder Leveraging Contextual Information
abstract
In the face of the rapidly growing spread of false and misleading information in the real world, manual evidence-based fact-checking efforts become increasingly challenging and time-consuming. In order to tackle this issue, we propose FaGANet, an automated and accurate fact-checking model that leverages the power of sentence-level attention and graph attention network to enhance performance. This model adeptly integrates encoder-only models with graph attention network, effectively fusing claims and evidence information for accurate identification of even well-disguised data. Experiment results showcase the significant improvement in accuracy achieved by our FaGANet model, as well as its state-of-the-art performance in the evidence-based fact-checking task. We release our code and data in https://github.com/WeiyaoLuo/FaGANet.
Weiyao Luo, Junfeng Ran, Zailong Tian, Sujian Li, Zhifang Sui
LREC/COLING4
2024 TabMedBERT: A Tabular Knowledge Enhanced Biomedical Pretrained Language Model
abstract
Most existing biomedical language models are trained on plain text with general learning goals such as random word infilling, failing to capture the knowledge in the biomedical corpus sufficiently. Since biomedical articles usually contain many tables summarising the main entities and their relations, in the paper, we propose a Tabular knowledge enhanced bioMedical pretrained language model, called TabMedBERT. Specifically, we align entities between table cells, and article text spans with pre-defined rules. Then we add two table-related self-supervised tasks to integrate tabular knowledge into the language model: Entity Infilling (EI) and Table Cloze Test (TCT). While EI masks tokens within aligned entities in the article, TCT converts aligned entities in the table layout into a cloze text by erasing one entity and prompts the model to extract the appropriate span to fill in the blank. Experimental results demonstrate that TabMedBERT surpasses all competing language models without adding additional parameters, establishing a new state-of-the-art performance of 85.59% (+1.29%) on the BLURB biomedical NLP benchmark and 7 additional information extraction datasets. Moreover, the model architecture for TCT provides a straightforward solution to revise information extraction with paired entities.
Lei Geng, Ziqiang Cao, Juntao Li 0005, Wenjie Li 0002, Sujian Li, Yang Yang 0074, Jun Zhang 0069
ECAI6
2024 Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement
abstract
Large language model agents have exhibited exceptional performance across a range of complex interactive tasks.Recent approaches have utilized tuning with expert trajectories to enhance agent performance, yet they primarily concentrate on outcome rewards, which may lead to errors or suboptimal actions due to the absence of process supervision signals.In this paper, we introduce the Iterative step-level Process Refinement (IPR) framework, which provides detailed step-by-step guidance to enhance agent training.Specifically, we adopt the Monte Carlo method to estimate step-level rewards.During each iteration, the agent explores along the expert trajectory and generates new actions.These actions are then evaluated against the corresponding step of expert trajectory using step-level rewards.Such comparison helps identify discrepancies, yielding contrastive action pairs that serve as training data for the agent.Our experiments on three complex agent tasks demonstrate that our framework outperforms a variety of strong baselines.Moreover, our analytical findings highlight the effectiveness of IPR in augmenting action efficiency and its applicability to diverse models † .
Weimin Xiong, Yifan Song 0002, Xiutian Zhao, Cheng Li 0040, Wei Peng 0011, Sujian Li
EMNLP9
2024 LongEmbed: Extending Embedding Models for Long Context Retrieval
abstract
Embedding models play a pivotal role in modern NLP applications such as document retrieval.However, existing embedding models are limited to encoding short documents of typically 512 tokens, restrained from application scenarios requiring long inputs.This paper explores context window extension of existing embedding models, pushing their input length to a maximum of 32,768.We begin by evaluating the performance of existing embedding models using our newly constructed LONGEM-BED benchmark, which includes two synthetic and four real-world tasks, featuring documents of varying lengths and dispersed target information.The benchmarking results highlight huge opportunities for enhancement in current models.Via comprehensive experiments, we demonstrate that training-free context window extension strategies can effectively increase the input length of these models by several folds.Moreover, comparison of models using Absolute Position Encoding (APE) and Rotary Position Encoding (RoPE) reveals the superiority of RoPE-based embedding models in context window extension, offering empirical guidance for future models.Our benchmark, code and trained models will be released to advance the research in long context embedding models.
Liang Wang 0046, Nan Yang 0002, Yifan Song 0002, Furu Wei, Sujian Li
EMNLP7
2024 PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
abstract
Large Language Models (LLMs) are trained with a pre-defined context length, restricting their use in scenarios requiring long inputs. Previous efforts for adapting LLMs to a longer length usually requires fine-tuning with this target length (Full-length fine-tuning), suffering intensive training cost. To decouple train length from target length for efficient context window extension, we propose Positional Skip-wisE (PoSE) training that smartly simulates long inputs using a fixed context window. This is achieved by first dividing the original context window into several chunks, then designing distinct skipping bias terms to manipulate the position indices of each chunk. These bias terms and the lengths of each chunk are altered for every training example, allowing the model to adapt to all positions within target length. Experimental results show that PoSE greatly reduces memory and time overhead compared with Full-length fine-tuning, with minimal impact on performance. Leveraging this advantage, we have successfully extended the LLaMA model to 128k tokens using a 2k training context window. Furthermore, we empirically confirm that PoSE is compatible with all RoPE-based LLMs and position interpolation strategies. Notably, our method can potentially support infinite length, limited only by memory usage in inference. With ongoing progress for efficient inference, we believe PoSE can further scale the context window beyond 128k.
Nan Yang 0002, Liang Wang 0046, Yifan Song 0002, Furu Wei, Sujian Li
ICLR7
2024 Selecting Large Language Model to Fine-tune via Rectified Scaling Law
abstract
The ever-growing ecosystem of LLMs has posed a challenge in selecting the most appropriate pre-trained model to fine-tune amidst a sea of options. Given constrained resources, fine-tuning all models and making selections afterward is unrealistic. In this work, we formulate this resource-constrained selection task into predicting fine-tuning performance and illustrate its natural connection with Scaling Law. Unlike pre-training, we find that the fine-tuning scaling curve includes not just the well-known "power phase" but also the previously unobserved "pre-power phase". We also explain why existing Scaling Law fails to capture this phase transition phenomenon both theoretically and empirically. To address this, we introduce the concept of "pre-learned data size" into our Rectified Scaling Law, which overcomes theoretical limitations and fits experimental results much better. By leveraging our law, we propose a novel LLM selection algorithm that selects the near-optimal model with hundreds of times less resource consumption, while other methods may provide negatively correlated selection. The project page is available at rectified-scaling-law.github.io.
Haowei Lin, Baizhou Huang, Haotian Ye, Qinyu Chen, Sujian Li, Jianzhu Ma, Xiaojun Wan 0001, James Zou 0001, Yitao Liang
ICML6
2024 Shapley Value-based Contrastive Alignment for Multimodal Information Extraction
abstract
The rise of social media and the exponential growth of multimodal communication necessitates advanced techniques for Multimodal Information Extraction (MIE). However, existing methodologies primarily rely on direct Image-Text interactions, a paradigm that often faces significant challenges due to semantic and modality gaps between images and text. In this paper, we introduce a new paradigm of Image-Context-Text interaction, where large multimodal models (LMMs) are utilized to generate descriptive textual context to bridge these gaps. In line with this paradigm, we propose a novel Shapley Value-based Contrastive Alignment (Shap-CA) method, which aligns both context-text and context-image pairs. Shap-CA initially applies the Shapley value concept from cooperative game theory to assess the individual contribution of each element in the set of contexts, texts and images towards total semantic and modality overlaps. Following this quantitative evaluation, a contrastive learning strategy is employed to enhance the interactive contribution within context-text/image pairs, while minimizing the influence across these pairs. Furthermore, we design an adaptive fusion module for selective cross-modal fusion. Extensive experiments across four MIE datasets demonstrate that our method significantly outperforms existing state-of-the-art methods.
Wen Luo 0001, Yu Xia 0024, Tianshu Shen, Sujian Li
ACM Multimedia4
2024 CoUDA: Coherence Evaluation via Unified Data Augmentation
abstract
Dawei Zhu, Wenhao Wu, Yifan Song, Fangwei Zhu, Ziqiang Cao, Sujian Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yifan Song 0002, Fangwei Zhu, Ziqiang Cao, Sujian Li
NAACL-HLT6
2024 Leveraging Two-Stream Cause-Effect Relation for Emotion-Cause Analysis
Junfeng Ran, Zailong Tian, Weiyao Luo, Sujian Li
NLPCC (4)4
2024 Dialogue acts enhanced extract-abstract framework for meeting summarization
Shichao Sun, Ruifeng Yuan, Wenjie Li 0002, Ziqiang Cao, Sujian Li
Inf. Process. Manag.5
2023 WeCheck: Strong Factual Consistency Checker via Weakly Supervised Learning
abstract
A crucial issue of current text generation models is that they often uncontrollably generate text that is factually inconsistent with inputs.Due to lack of annotated data, existing factual consistency metrics usually train evaluation models on synthetic texts or directly transfer from other related tasks, such as question answering (QA) and natural language inference (NLI).Bias in synthetic text or upstream tasks makes them perform poorly on text actually generated by language models, especially for general evaluation for various tasks.To alleviate this problem, we propose a weakly supervised framework named WeCheck that is directly trained on actual generated samples from language models with weakly annotated labels.WeCheck first utilizes a generative model to infer the factual labels of generated samples by aggregating weak labels from multiple resources.Next, we train a simple noise-aware classification model as the target metric using the inferred weakly supervised information.Comprehensive experiments on various tasks demonstrate the strong performance of WeCheck, achieving an average absolute improvement of 3.3% on the TRUE benchmark over 11B state-of-the-art methods using only 435M parameters.Furthermore, it is up to 30× faster than previous evaluation methods, greatly improving the accuracy and efficiency of factual consistency evaluation. 1
Wei Li 0176, Xinyan Xiao, Sujian Li, Yajuan Lyu
ACL (1)5
2023 Rationale-Enhanced Language Models are Better Continual Relation Learners
abstract
Continual relation extraction (CRE) aims to solve the problem of catastrophic forgetting when learning a sequence of newly emerging relations.Recent CRE studies have found that catastrophic forgetting arises from the model's lack of robustness against future analogous relations.To address the issue, we introduce rationale, i.e., the explanations of relation classification results generated by large language models (LLM), into CRE task.Specifically, we design the multi-task rationale tuning strategy to help the model learn current relations robustly.We also conduct contrastive rationale replay to further distinguish analogous relations.Experimental results on two standard benchmarks demonstrate that our method outperforms the state-of-the-art CRE models.Our code is available at https://github.com/WeiminXiong/ RationaleCL
Weimin Xiong, Yifan Song 0002, Peiyi Wang, Sujian Li
EMNLP4
2023 Improving Sentence Similarity Estimation for Unsupervised Extractive Summarization
abstract
Unsupervised extractive summarization aims to extract salient sentences from a document as the summary without labeled data. Recent literatures mostly research how to leverage sentence similarity to rank sentences in the order of salience. However, sentence similarity estimation using pre-trained language models mostly takes little account of document-level information and has a weak correlation with sentence salience ranking. In this paper, we proposed two novel strategies to improve sentence similarity estimation for unsupervised extractive summarization. We use contrastive learning to optimize a document-level objective that sentences from the same document are more similar than those from different documents. Moreover, we use mutual learning to enhance the relationship between sentence similarity estimation and sentence salience ranking, where an extra signal amplifier is used to refine the pivotal information. Experimental results demonstrate the effectiveness of our strategies.1
Shichao Sun, Ruifeng Yuan, Wenjie Li 0002, Sujian Li
ICASSP4
2023 DocRED-FE: A Document-Level Fine-Grained Entity and Relation Extraction Dataset
abstract
Joint entity and relation extraction (JERE) is one of the most important tasks in information extraction. However, most existing works focus on sentence-level coarse-grained JERE, which have limitations in real-world scenarios. In this paper, we construct a large-scale document-level fine-grained JERE dataset DocRED-FE, which improves DocRED with Fine-Grained Entity Type. Specifically, we redesign a hierarchical entity type schema including 11 coarse-grained types and 119 fine-grained types, and then re-annotate DocRED manually according to this schema. Through comprehensive experiments we find that: (1) DocRED-FE is challenging to existing JERE models; (2) Our fine-grained entity types promote relation classification. We make DocRED-FE with instruction and the code for our baselines publicly available at https://github.com/PKU-TANGENT/DOCRED-FE.
Weimin Xiong, Yifan Song 0002, Sujian Li
ICASSP6
2023 Probing Bilingual Guidance for Cross-Lingual Summarization
Sujian Li
NLPCC (1)3
2022 Premise-based Multimodal Reasoning: Conditional Inference on Joint Textual and Visual Clues
abstract
Qingxiu Dong, Ziwei Qin, Heming Xia, Tian Feng, Shoujie Tong, Haoran Meng, Lin Xu, Zhongyu Wei, Weidong Zhan, Baobao Chang, Sujian Li, Tianyu Liu, Zhifang Sui. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Qingxiu Dong, Ziwei Qin, Heming Xia, Shoujie Tong, Haoran Meng, Zhongyu Wei, Weidong Zhan, Baobao Chang, Sujian Li, Tianyu Liu 0001, Zhifang Sui
ACL (1)11
2022 IntTower: The Next Generation of Two-Tower Model for Pre-Ranking System
abstract
Scoring a large number of candidates precisely in several milliseconds is vital for industrial pre-ranking systems. Existing pre-ranking systems primarily adopt the two-tower model since the "user-item decoupling architecture" paradigm is able to balance the efficiency and effectiveness. However, the cost of high efficiency is the neglect of the potential information interaction between user and item towers, hindering the prediction accuracy critically. In this paper, we show it is possible to design a two-tower model that emphasizes both information interactions and inference efficiency. The proposed model, IntTower (short for Interaction enhanced Two-Tower), consists of Light-SE, FE-Block and CIR modules. Specifically, lightweight Light-SE module is used to identify the importance of different features and obtain refined feature representations in each tower. FE-Block module performs fine-grained and early feature interactions to capture the interactive signals between user and item towers explicitly and CIR module leverages a contrastive interaction regularization to further enhance the interactions implicitly. Experimental results on three public datasets show that IntTower outperforms the SOTA pre-ranking models significantly and even achieves comparable performance in comparison with the ranking models. Moreover, we further verify the effectiveness of IntTower on a large-scale advertisement pre-ranking system. The code of IntTower is publicly available https://gitee.com/mindspore/models/tree/master/research/recommend/IntTower.
Xiangyang Li 0004, Bo Chen 0023, Huifeng Guo, Chenxu Zhu, Xiang Long, Sujian Li, Yichao Wang 0002, Wei Guo 0006, Longxia Mao, Zhenhua Dong, Ruiming Tang
CIKM7
2022 A Transition-based Method for Complex Question Understanding
abstract
Complex Question Understanding (CQU) parses complex questions to Question Decomposition Meaning Representation (QDMR) which is a sequence of atomic operators. Existing works are based on end-to-end neural models which do not explicitly model the intermediate states and lack interpretability for the parsing process. Besides, they predict QDMR in a mismatched granularity and do not model the step-wise information which is an essential characteristic of QDMR. To alleviate the issues, we treat QDMR as a computational graph and propose a transition-based method where a decider predicts a sequence of actions to build the graph node-by-node. In this way, the partial graph at each step enables better representation of the intermediate states and better interpretability. At each step, the decider encodes the intermediate state with specially designed encoders and predicts several candidates of the next action and its confidence. For inference, a searcher seeks the optimal graph based on the predictions of the decider to alleviate the error propagation. Experimental results demonstrate the parsing accuracy of our method against several strong baselines. Moreover, our method has transparent and human-readable intermediate results, showing improved interpretability.
Wenbin Jiang 0002, Yajuan Lyu, Sujian Li
COLING4
2022 ConFiguRe: Exploring Discourse-level Chinese Figures of Speech
abstract
Figures of speech, such as metaphor and irony, are ubiquitous in literature works and colloquial conversations. This poses great challenge for natural language understanding since figures of speech usually deviate from their ostensible meanings to express deeper semantic implications. Previous research lays emphasis on the literary aspect of figures and seldom provide a comprehensive exploration from a view of computational linguistics. In this paper, we first propose the concept of figurative unit, which is the carrier of a figure. Then we select 12 types of figures commonly used in Chinese, and build a Chinese corpus for Contextualized Figure Recognition (ConFiguRe). Different from previous token-level or sentence-level counterparts, ConFiguRe aims at extracting a figurative unit from discourse-level context, and classifying the figurative unit into the right figure type. On ConFiguRe, three tasks, i.e., figure extraction, figure type classification and figure recognition, are designed and the state-of-the-art techniques are utilized to implement the benchmarks. We conduct thorough experiments and show that all three tasks are challenging for existing models, thus requiring further research. Our dataset and code are publicly available at https://github.com/pku-tangent/ConFiguRe.
Qiusi Zhan, Zhejian Zhou, Yifan Song 0002, Jiebin Zhang, Sujian Li
COLING6
2022 Learning Robust Representations for Continual Relation Extraction via Adversarial Class Augmentation
abstract
Continual relation extraction (CRE) aims to continually learn new relations from a classincremental data stream.CRE model usually suffers from catastrophic forgetting problem, i.e., the performance of old relations seriously degrades when the model learns new relations.Most previous work attributes catastrophic forgetting to the corruption of the learned representations as new relations come, with an implicit assumption that the CRE models have adequately learned the old relations.In this paper, through empirical studies we argue that this assumption may not hold, and an important reason for catastrophic forgetting is that the learned representations do not have good robustness against the appearance of analogous relations in the subsequent learning process.To address this issue, we encourage the model to learn more precise and robust representations through a simple yet effective adversarial class augmentation mechanism (ACA), which is easy to implement and model-agnostic.Experimental results show that ACA can consistently improve the performance of state-of-theart CRE models on two popular benchmarks.
Peiyi Wang, Yifan Song 0002, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Sujian Li, Zhifang Sui
EMNLP6
2022 Precisely the Point: Adversarial Augmentations for Faithful and Informative Text Generation
abstract
Though model robustness has been extensively studied in language understanding, the robustness of Seq2Seq generation remains understudied.In this paper, we conduct the first quantitative analysis on the robustness of pre-trained Seq2Seq models.We find that even current SOTA pre-trained Seq2Seq model (BART) is still vulnerable, which leads to significant degeneration in faithfulness and informativeness for text generation tasks.This motivated us to further propose a novel adversarial augmentation framework, namely AdvSeq, for generally improving faithfulness and informativeness of Seq2Seq models via enhancing their robustness.AdvSeq automatically constructs two types of adversarial augmentations during training, including implicit adversarial samples by perturbing word representations and explicit adversarial samples by word swapping, both of which effectively improve Seq2Seq robustness.Extensive experiments on three popular text generation tasks demonstrate that AdvSeq significantly improves both the faithfulness and informativeness of Seq2Seq generation under both automatic and human evaluation settings.
Wei Li 0176, Xinyan Xiao, Sujian Li, Yajuan Lyu
EMNLP5
2022 Low Resource Style Transfer via Domain Adaptive Meta Learning
abstract
Text style transfer (TST) without parallel data has achieved some practical success.However, most of the existing unsupervised text style transfer methods suffer from (i) requiring massive amounts of non-parallel data to guide transferring different text styles.(ii) colossal performance degradation when fine-tuning the model in new domains.In this work, we propose DAML-ATM (Domain Adaptive Meta-Learning with Adversarial Transfer Model), which consists of two parts: DAML and ATM.DAML is a domain adaptive meta-learning approach to learn general knowledge in multiple heterogeneous source domains, capable of adapting to new unseen domains with a small amount of data.Moreover, we propose a new unsupervised TST approach Adversarial Transfer Model (ATM), composed of a sequence-to-sequence pre-trained language model and uses adversarial style training for better content preservation and style transfer.Results on multi-domain datasets demonstrate that our approach generalizes well on unseen low-resource domains, achieving state-of-theart results against ten strong baselines.
Xiang Long, Sujian Li
NAACL-HLT4
2022 RODA: Reverse Operation Based Data Augmentation for Solving Math Word Problems
abstract
Automatically solving math word problems is a critical task in the field of natural language processing. Recent models have reached their performance bottleneck and require more high-quality data for training. We propose a novel data augmentation method that reverses the mathematical logic of math word problems to produce new high-quality math problems and introduce new knowledge points that can benefit learning the mathematical reasoning logic. We apply the augmented data on two SOTA math word problem solving models and compare our results with a strong data augmentation baseline. Experimental results show the effectiveness of our approach (we release our code and data athttps://github.com/yiyunya/RODA).
Qianying Liu, Wenyu Guan, Sujian Li, Fei Cheng 0002, Daisuke Kawahara, Sadao Kurohashi
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Guiding the Growth: Difficulty-Controllable Question Generation through Step-by-Step Rewriting
abstract
Yi Cheng, Siyao Li, Bang Liu, Ruihui Zhao, Sujian Li, Chenghua Lin, Yefeng Zheng. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Siyao Li, Bang Liu 0003, Ruihui Zhao, Sujian Li, Chenghua Lin 0002, Yefeng Zheng 0001
ACL/IJCNLP (1)5
2021 BASS: Boosting Abstractive Summarization with Unified Semantic Graph
abstract
Wenhao Wu, Wei Li, Xinyan Xiao, Jiachen Liu, Ziqiang Cao, Sujian Li, Hua Wu, Haifeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wei Li 0176, Xinyan Xiao, Ziqiang Cao, Sujian Li, Hua Wu 0003, Haifeng Wang 0001
ACL/IJCNLP (1)6
2021 Knowledge Enhanced Transformers System for Claim Stance Classification
Xiangyang Li 0004, Zheng Li 0030, Sujian Li, Shimin Yang
NLPCC (2)3
2020 A Robust Adversarial Training Approach to Machine Reading Comprehension
abstract
Lacking robustness is a serious problem for Machine Reading Comprehension (MRC) models. To alleviate this problem, one of the most promising ways is to augment the training dataset with sophisticated designed adversarial examples. Generally, those examples are created by rules according to the observed patterns of successful adversarial attacks. Since the types of adversarial examples are innumerable, it is not adequate to manually design and enrich training data to defend against all types of adversarial attacks. In this paper, we propose a novel robust adversarial training approach to improve the robustness of MRC models in a more generic way. Given an MRC model well-trained on the original dataset, our approach dynamically generates adversarial examples based on the parameters of current model and further trains the model by using the generated examples in an iterative schedule. When applied to the state-of-the-art MRC models, including QANET, BERT and ERNIE2.0, our approach obtains significant and comprehensive improvements on 5 adversarial datasets constructed in different ways, without sacrificing the performance on the original SQuAD development set. Moreover, when coupled with other data augmentation strategy, our approach further boosts the overall performance on adversarial datasets and outperforms the state-of-the-art methods.
Kai Liu 0023, Xin Liu 0066, An Yang, Jing Liu 0022, Jinsong Su, Sujian Li, Qiaoqiao She
AAAI6
2020 Composing Elementary Discourse Units in Abstractive Summarization
abstract
In this paper, we argue that elementary discourse unit (EDU) is a more appropriate textual unit of content selection than the sentence unit in abstractive summarization.To well handle the problem of composing EDUs into an informative and fluent summary, we propose a novel summarization method that first designs an EDU selection model to extract and group informative EDUs and then an EDU fusion model to fuse the EDUs in each group into one sentence.We also design the reinforcement learning mechanism to use EDU fusion results to reward the EDU selection action, boosting the final summarization performance.Experiments on CNN/Daily Mail have demonstrated the effectiveness of our model.
Zhenwen Li, Sujian Li
ACL3
2020 Syntax-Aware Graph Attention Network for Aspect-Level Sentiment Classification
abstract
Aspect-level sentiment classification aims to distinguish the sentiment polarities over aspect terms in a sentence.Existing approaches mostly focus on modeling the relationship between the given aspect words and their contexts with attention, and ignore the use of more elaborate knowledge implicit in the context.In this paper, we exploit syntactic awareness to the model by the graph attention network on the dependency tree structure and external pre-training knowledge by BERT language model, which helps to model the interaction between the context and aspect words better.And the subwords of BERT are integrated into the dependency tree graphs, which can obtain more accurate representations of words by graph attention.Experiments demonstrate the effectiveness of our model.
Lianzhe Huang, Xin Sun 0013, Sujian Li, Linhao Zhang, Houfeng Wang
COLING3
2020 Joint Extraction of Entities and Relations Based on a Novel Decomposition Strategy
abstract
Joint extraction of entities and relations aims to detect entity pairs along with their relations using a single model. Prior work typically solves this task in the extract-then-classify or unified labeling manner. However, these methods either suffer from the redundant entity pairs, or ignore the important inner structure in the process of extracting entities and relations. To address these limitations, in this paper, we first decompose the joint extraction task into two interrelated subtasks, namely HE extraction and TER extraction. The former subtask is to distinguish all head-entities that may be involved with target relations, and the latter is to identify corresponding tail-entities and relations for each extracted head-entity. Next, these two subtasks are further deconstructed into several sequence labeling problems based on our proposed span-based tagging scheme, which are conveniently solved by a hierarchical boundary tagger and a multi-span decoding algorithm. Owing to the reasonable decomposition strategy, our model can fully capture the semantic interdependency between different steps, as well as reduce noise from irrelevant entity pairs. Experimental results show that our method outperforms previous work by 5.2%, 5.9% and 21.5% (F1 score), achieving a new state-of-the-art on three public datasets.
Bowen Yu 0002, Zhenyu Zhang 0006, Xiaobo Shu, Tingwen Liu, Bin Wang 0004, Sujian Li
ECAI7
2020 Evaluating Text Coherence at Sentence and Paragraph Levels
abstract
In this paper, to evaluate text coherence, we propose the paragraph ordering task as well as conducting sentence ordering. We collected four distinct corpora from different domains on which we investigate the adaptation of existing sentence ordering methods to a paragraph ordering task. We also compare the learnability and robustness of existing models by artificially creating mini datasets and noisy datasets respectively and verifying the efficiency of established models under these circumstances. Furthermore, we carry out human evaluation on the rearranged passages from two competitive models and confirm that WLCS-l is a better metric performing significantly higher correlations with human rating than τ , the most prevalent metric used before. Results from these evaluations show that except for certain extreme conditions, the recurrent graph neural network-based model is an optimal choice for coherence modeling.
Sennan Liu, Shuang Zeng, Sujian Li
LREC3
2020 Enhanced-RCNN: An Efficient Method for Learning Sentence Similarity
abstract
Learning sentence similarity is a fundamental research topic and has been explored using various deep learning methods recently. In this paper, we further propose an enhanced recurrent convolutional neural network (Enhanced-RCNN) model for learning sentence similarity. Compared to the state-of-the-art BERT model, the architecture of our proposed model is far less complex. Experimental results show that our similarity learning method outperforms the baselines and achieves the competitive performance on two real-world paraphrase identification datasets.
Shuang Peng 0009, Hengbin Cui, Niantao Xie, Sujian Li, Xiaolong Li 0005
WWW4
2019 Exploring Sequence-to-Sequence Learning in Aspect Term Extraction
abstract
Aspect term extraction (ATE) aims at identifying all aspect terms in a sentence and is usually modeled as a sequence labeling problem.However, sequence labeling based methods cannot make full use of the overall meaning of the whole sentence and have the limitation in processing dependencies between labels.To tackle these problems, we first explore to formalize ATE as a sequence-tosequence (Seq2Seq) learning task where the source sequence and target sequence are composed of words and labels respectively.At the same time, to make Seq2Seq learning suit to ATE where labels correspond to words one by one, we design the gated unit networks to incorporate corresponding word representation into the decoder, and position-aware attention to pay more attention to the adjacent words of a target word.The experimental results on two datasets show that Seq2Seq learning is effective in ATE accompanied with our proposed gated unit networks and position-aware attention mechanism.
Dehong Ma, Sujian Li, Fangzhao Wu, Xing Xie 0001, Houfeng Wang
ACL (1)2
2019 Enhancing Pre-Trained Language Representations with Rich Knowledge for Machine Reading Comprehension
abstract
Machine reading comprehension (MRC) is a crucial and challenging task in NLP.Recently, pre-trained language models (LMs), especially BERT, have achieved remarkable success, presenting new state-of-the-art results in MRC.In this work, we investigate the potential of leveraging external knowledge bases (KBs) to further improve BERT for MRC.We introduce KT-NET, which employs an attention mechanism to adaptively select desired knowledge from KBs, and then fuses selected knowledge with BERT to enable context-and knowledgeaware predictions.We believe this would combine the merits of both deep LMs and curated KBs towards better MRC.Experimental results indicate that KT-NET offers significant and consistent improvements over BERT, outperforming competitive baselines on ReCoRD and SQuAD1.1 benchmarks.Notably, it ranks the 1st place on the ReCoRD leaderboard, and is also the best single model on the SQuAD1.1 leaderboard at the time of submission (March 4th, 2019). 1
An Yang, Quan Wang 0002, Jing Liu 0022, Kai Liu 0023, Yajuan Lyu, Hua Wu 0003, Qiaoqiao She, Sujian Li
ACL (1)8
2019 Text Level Graph Neural Network for Text Classification
abstract
Lianzhe Huang, Dehong Ma, Sujian Li, Xiaodong Zhang, Houfeng Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Lianzhe Huang, Dehong Ma, Sujian Li, Xiaodong Zhang 0022, Houfeng Wang
EMNLP/IJCNLP (1)3
2019 Tree-structured Decoding for Solving Math Word Problems
abstract
Qianying Liu, Wenyv Guan, Sujian Li, Daisuke Kawahara. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Qianying Liu, Wenyu Guan, Sujian Li, Daisuke Kawahara
EMNLP/IJCNLP (1)3
2019 Do NLP Models Know Numbers? Probing Numeracy in Embeddings
abstract
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, Matt Gardner. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh 0001, Matt Gardner 0001
EMNLP/IJCNLP (1)3
2019 Denoising based Sequence-to-Sequence Pre-training for Text Generation
abstract
Liang Wang, Wei Zhao, Ruoyu Jia, Sujian Li, Jingming Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Liang Wang 0046, Ruoyu Jia, Sujian Li, Jingming Liu
EMNLP/IJCNLP (1)4
2019 Beyond Word Attention: Using Segment Attention in Neural Relation Extraction
abstract
Relation extraction studies the issue of predicting semantic relations between pairs of entities in sentences. Attention mechanisms are often used in this task to alleviate the inner-sentence noise by performing soft selections of words independently. Based on the observation that information pertinent to relations is usually contained within segments (continuous words in a sentence), it is possible to make use of this phenomenon for better extraction. In this paper, we aim to incorporate such segment information into neural relation extractor. Our approach views the attention mechanism as linear-chain conditional random fields over a set of latent variables whose edges encode the desired structure, and regards attention weight as the marginal distribution of each word being selected as a part of the relational expression. Experimental results show that our method can attend to continuous relational expressions without explicit annotations, and achieve the state-of-the-art performance on the large-scale TACRED dataset.
Bowen Yu 0002, Zhenyu Zhang 0006, Tingwen Liu, Bin Wang 0004, Sujian Li, Quangang Li
IJCAI5
2019 We Know What You Will Ask: A Dialogue System for Multi-intent Switch and Prediction
Qi Chen 0009, Lei Sha, Hui Xue 0004, Sujian Li, Houfeng Wang
NLPCC (1)5
2018 Faithful to the Original: Fact Aware Neural Abstractive Summarization
abstract
Unlike extractive summarization, abstractive summarization has to fuse different parts of the source text, which inclines to create fake facts. Our preliminary study reveals nearly 30% of the outputs from a state-of-the-art neural summarization system suffer from this problem. While previous abstractive summarization approaches usually focus on the improvement of informativeness, we argue that faithfulness is also a vital prerequisite for a practical abstractive summarization system. To avoid generating fake facts in a summary, we leverage open information extraction and dependency parse technologies to extract actual fact descriptions from the source text. The dual-attention sequence-to-sequence framework is then proposed to force the generation conditioned on both the source text and the extracted fact descriptions. Experiments on the Gigaword benchmark dataset demonstrate that our model can greatly reduce fake summaries by 80%. Notably, the fact descriptions also bring significant improvement on informativeness since they often condense the meaning of the source text.
Ziqiang Cao, Furu Wei, Wenjie Li 0002, Sujian Li
AAAI4
2018 Order-Planning Neural Text Generation From Structured Data
abstract
Generating texts from structured data (e.g., a table) is important for various natural language processing tasks such as question answering and dialog systems. In recent studies, researchers use neural language models and encoder-decoder frameworks for table-to-text generation. However, these neural network-based approaches typically do not model the order of content during text generation. When a human writes a summary based on a given table, he or she would probably consider the content order before wording. In this paper, we propose an order-planning text generation model, where order information is explicitly captured by link-based attention. Then a self-adaptive gate combines the link-based attention with traditional content-based attention. We conducted experiments on the WikiBio dataset and achieve higher performance than previous methods in terms of BLEU, ROUGE, and NIST scores; we also performed ablation tests to analyze each component of our model.
Lei Sha, Lili Mou, Tianyu Liu 0001, Pascal Poupart, Sujian Li, Baobao Chang, Zhifang Sui
AAAI5
2018 Retrieve, Rerank and Rewrite: Soft Template Based Neural Summarization
abstract
Most previous seq2seq summarization systems purely depend on the source text to generate summaries, which tends to work unstably.Inspired by the traditional template-based summarization approaches, this paper proposes to use existing summaries as soft templates to guide the seq2seq model.To this end, we use a popular IR platform to Retrieve proper summaries as candidate templates.Then, we extend the seq2seq framework to jointly conduct template Reranking and templateaware summary generation (Rewriting).Experiments show that, in terms of informativeness, our model significantly outperforms the state-of-the-art methods, and even soft templates themselves demonstrate high competitiveness.In addition, the import of high-quality external summaries improves the stability and readability of generated summaries.
Ziqiang Cao, Wenjie Li 0002, Sujian Li, Furu Wei
ACL (1)3
2018 Multi-Passage Machine Reading Comprehension with Cross-Passage Answer Verification
abstract
Yizhong Wang, Kai Liu, Jing Liu, Wei He, Yajuan Lyu, Hua Wu, Sujian Li, Haifeng Wang. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Yizhong Wang, Kai Liu 0023, Jing Liu 0022, Wei He 0014, Yajuan Lyu, Hua Wu 0003, Sujian Li, Haifeng Wang 0001
ACL (1)7
2018 Multi-Perspective Context Aggregation for Semi-supervised Cloze-style Reading Comprehension
abstract
Cloze-style reading comprehension has been a popular task for measuring the progress of natural language understanding in recent years. In this paper, we design a novel multi-perspective framework, which can be seen as the joint training of heterogeneous experts and aggregate context information from different perspectives. Each perspective is modeled by a simple aggregation module. The outputs of multiple aggregation modules are fed into a one-timestep pointer network to get the final answer. At the same time, to tackle the problem of insufficient labeled data, we propose an efficient sampling mechanism to automatically generate more training examples by matching the distribution of candidates between labeled and unlabeled data. We conduct our experiments on a recently released cloze-test dataset CLOTH (Xie et al., 2017), which consists of nearly 100k questions designed by professional teachers. Results show that our method achieves new state-of-the-art performance over previous strong baselines.
Liang Wang 0046, Sujian Li, Kewei Shen, Ruoyu Jia, Jingming Liu
COLING2
2018 Joint Learning for Targeted Sentiment Analysis
abstract
Targeted sentiment analysis (TSA) aims at extracting targets and classifying their sentiment classes.Previous works only exploit word embeddings as features and do not explore more potentials of neural networks when jointly learning the two tasks.In this paper, we carefully design the hierarchical multi-layer bidirectional gated recurrent units (HMBi-GRU) model to learn abstract features for both tasks, and we propose a HMBi-GRU based joint model which allows the target label of word to have influence on its sentiment label.Experimental results on two datasets show that our joint learning model can outperform other baselines and demonstrate the effectiveness of HMBi-GRU in learning abstract features.
Dehong Ma, Sujian Li, Houfeng Wang
EMNLP2
2018 Auto-Dialabel: Labeling Dialogue Data with Unsupervised Learning
abstract
The lack of labeled data is one of the main challenges when building a task-oriented dialogue system.Existing dialogue datasets usually rely on human labeling, which is expensive, limited in size, and in low coverage.In this paper, we instead propose our framework auto-dialabel to automatically cluster the dialogue intents and slots.In this framework, we collect a set of context features, leverage an autoencoder for feature assembly, and adapt a dynamic hierarchical clustering method for intent and slot labeling.Experimental results show that our framework can promote human labeling cost to a great extent, achieve good intent clustering accuracy (84.1%), and provide reasonable and instructive slot labeling results.
Qi Chen 0009, Lei Sha, Sujian Li, Xu Sun 0001, Houfeng Wang
EMNLP4
2018 Toward Fast and Accurate Neural Discourse Segmentation
abstract
Discourse segmentation, which segments texts into Elementary Discourse Units, is a fundamental step in discourse analysis.Previous discourse segmenters rely on complicated hand-crafted features and are not practical in actual use.In this paper, we propose an endto-end neural segmenter based on BiLSTM-CRF framework.To improve its accuracy, we address the problem of data insufficiency by transferring a word representation model that is trained on a large corpus.We also propose a restricted self-attention mechanism in order to capture useful information within a neighborhood.Experiments on the RST-DT corpus show that our model is significantly faster than previous methods, while achieving new stateof-the-art performance.1
Yizhong Wang, Sujian Li, Jingfeng Yang 0001
EMNLP2
2018 Query and Output: Generating Words by Querying Distributed Word Representations for Paraphrase Generation
abstract
Shuming Ma, Xu Sun, Wei Li, Sujian Li, Wenjie Li, Xuancheng Ren. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Shuming Ma, Xu Sun 0001, Wei Li 0101, Sujian Li, Wenjie Li 0002, Xuancheng Ren
NAACL-HLT4
2018 Target Extraction via Feature-Enriched Neural Networks Model
Dehong Ma, Sujian Li, Houfeng Wang
NLPCC (1)2
2018 Abstractive Summarization Improved by WordNet-Based Extractive Sentences
Niantao Xie, Sujian Li, Huiling Ren, Qibin Zhai
NLPCC (1)2
2018 Cross-Domain and Semisupervised Named Entity Recognition in Chinese Social Media: A Unified Model
abstract
Named entity recognition (NER) in Chinese social media is an important, but challenging task because Chinese social media language is informal and noisy. Most previous methods on NER focus on in-domain supervised learning, which is limited by scarce annotated data in social media. In this paper, we present that sufficient corpora in formal domains and massive unannotated text can be combined to improve the NER performance in social media. We propose a unified model which can learn from out-of-domain corpora and in-domain unannotated text. The unified model is composed of two parts. One is for cross-domain learning and the other is for semisupervised learning. Cross-domain learning can learn out-of-domain information based on domain similarity. Semisupervised learning can learn in-domain unannotated information by self-training. Experimental results show that our unified model yields a 9.57% improvement over strong baselines and achieves the state-of-the-art performance.
Jingjing Xu 0001, Hangfeng He 0001, Xu Sun 0001, Xuancheng Ren, Sujian Li
IEEE ACM Trans. Audio Speech Lang. Process.5
2017 Joint Copying and Restricted Generation for Paraphrase
abstract
Many natural language generation tasks, such as abstractive summarization and text simplification, are paraphrase-orientated. In these tasks, copying and rewriting are two main writing modes. Most previous sequence-to-sequence (Seq2Seq) models use a single decoder and neglect this fact. In this paper, we develop a novel Seq2Seq model to fuse a copying decoder and a restricted generative decoder. The copying decoder finds the position to be copied based on a typical attention model. The generative decoder produces words limited in the source-specific vocabulary. To combine the two decoders and determine the final output, we develop a predictor to predict the mode of copying or rewriting. This predictor can be guided by the actual writing mode in the training data. We conduct extensive experiments on two different paraphrase datasets. The result shows that our model outperforms the state-of-the-art approaches in terms of both informativeness and language quality.
Ziqiang Cao, Chuwei Luo, Wenjie Li 0002, Sujian Li
AAAI4
2017 Improving Multi-Document Summarization via Text Classification
abstract
Developed so far, multi-document summarization has reached its bottleneck due to the lack of sufficient training data and diverse categories of documents. Text classification just makes up for these deficiencies. In this paper, we propose a novel summarization system called TCSum, which leverages plentiful text classification data to improve the performance of multi-document summarization. TCSum projects documents onto distributed representations which act as a bridge between text classification and summarization. It also utilizes the classification results to produce summaries of different styles. Extensive experiments on DUC generic multi-document summarization datasets show that, TCSum can achieve the state-of-the-art performance without using any hand-crafted features and has the capability to catch the variations of summary styles with respect to different text categories.
Ziqiang Cao, Wenjie Li 0002, Sujian Li, Furu Wei
AAAI3
2017 Attentive Interactive Neural Networks for Answer Selection in Community Question Answering
abstract
Answer selection plays a key role in community question answering (CQA). Previous research on answer selection usually ignores the problems of redundancy and noise prevalent in CQA. In this paper, we propose to treat different text segments differently and design a novel attentive interactive neural network (AI-NN) to focus on those text segments useful to answer selection. The representations of question and answer are first learned by convolutional neural networks (CNNs) or other neural network architectures. Then AI-NN learns interactions of each paired segments of two texts. Row-wise and column-wise pooling are used afterwards to collect the interactions. We adopt attention mechanism to measure the importance of each segment and combine the interactions to obtain fixed-length representations for question and answer. Experimental results on CQA dataset in SemEval-2016 demonstrate that AI-NN outperforms state-of-the-art method.
Xiaodong Zhang 0022, Sujian Li, Lei Sha, Houfeng Wang
AAAI2
2017 Learning to Rank Semantic Coherence for Topic Segmentation
abstract
Topic segmentation plays an important role for discourse parsing and information retrieval.Due to the absence of training data, previous work mainly adopts unsupervised methods to rank semantic coherence between paragraphs for topic segmentation.In this paper, we present an intuitive and simple idea to automatically create a "quasi" training dataset, which includes a large amount of text pairs from the same or different documents with different semantic coherence.With the training corpus, we design a symmetric CNN neural network to model text pairs and rank the semantic coherence within the learning to rank framework.Experiments show that our algorithm is able to achieve competitive performance over strong baselines on several real-world datasets.
Liang Wang 0046, Sujian Li, Yajuan Lü, Houfeng Wang
EMNLP2
2017 Interactive Attention Networks for Aspect-Level Sentiment Classification
abstract
Aspect-level sentiment classification aims at identifying the sentiment polarity of specific target in its context. Previous approaches have realized the importance of targets in sentiment classification and developed various methods with the goal of precisely modeling thier contexts via generating target-specific representations. However, these studies always ignore the separate modeling of targets. In this paper, we argue that both targets and contexts deserve special treatment and need to be learned their own representations via interactive learning. Then, we propose the interactive attention networks (IAN) to interactively learn attentions in the contexts and targets, and generate the representations for targets and contexts separately. With this design, the IAN model can well represent a target and its collocative context, which is helpful to sentiment classification. Experimental results on SemEval 2014 Datasets demonstrate the effectiveness of our model.
Dehong Ma, Sujian Li, Xiaodong Zhang 0022, Houfeng Wang
IJCAI2
2017 Cascading Multiway Attentions for Document-level Sentiment Classification
abstract
Document-level sentiment classification aims to assign the user reviews a sentiment polarity. Previous methods either just utilized the document content without consideration of user and product information, or did not comprehensively consider what roles the three kinds of information play in text modeling. In this paper, to reasonably use all the information, we present the idea that user, product and their combination can all influence the generation of attentions to words and sentences, when judging the sentiment of a document. With this idea, we propose a cascading multiway attention (CMA) model, where multiple ways of using user and product information are cascaded to influence the generation of attentions on the word and sentence layers. Then, sentences and documents are well modeled by multiple representation vectors, which provide rich information for sentiment classification. Experiments on IMDB and Yelp datasets demonstrate the effectiveness of our model.
Dehong Ma, Sujian Li, Xiaodong Zhang 0022, Houfeng Wang, Xu Sun 0001
IJCNLP(1)2
2017 Tag-Enhanced Tree-Structured Neural Networks for Implicit Discourse Relation Classification
abstract
Identifying implicit discourse relations between text spans is a challenging task because it requires understanding the meaning of the text. To tackle this task, recent studies have tried several deep learning methods but few of them exploited the syntactic information. In this work, we explore the idea of incorporating syntactic parse tree into neural networks. Specifically, we employ the Tree-LSTM model and Tree-GRU model, which is based on the tree structure, to encode the arguments in a relation. And we further leverage the constituent tags to control the semantic composition process in these tree-structured neural networks. Experimental results show that our method achieves state-of-the-art performance on PDTB corpus.
Yizhong Wang, Sujian Li, Jingfeng Yang 0001, Xu Sun 0001, Houfeng Wang
IJCNLP(1)2
2016 TGSum: Build Tweet Guided Multi-Document Summarization Dataset
abstract
The development of summarization research has been significantly hampered by the costly acquisition of reference summaries. This paper proposes an effective way to automatically collect large scales of news-related multi-document summaries with reference to social media's reactions. We utilize two types of social labels in tweets, i.e., hashtags and hyper-links. Hashtags are used to cluster documents into different topic sets. Also, a tweet with a hyper-link often highlights certain key points of the corresponding document. We synthesize a linked document cluster to form a reference summary which can cover most key points. To this aim, we adopt the ROUGE metrics to measure the coverage ratio, and develop an Integer Linear Programming solution to discover the sentence set reaching the upper bound of ROUGE. Since we allow summary sentences to be selected from both documents and high-quality tweets, the generated reference summaries could be abstractive. Both informativeness and readability of the collected summaries are verified by manual judgment. In addition, we train a Support Vector Regression summarizer on DUC generic multi-document summarization benchmarks. With the collected data as extra training resource, the performance of the summarizer improves a lot on all the test sets. We release this dataset for further research.
Ziqiang Cao, Chengyao Chen, Wenjie Li 0002, Sujian Li, Furu Wei, Ming Zhou 0001
AAAI4
2016 Implicit Discourse Relation Classification via Multi-Task Neural Networks
abstract
Without discourse connectives, classifying implicit discourse relations is a challenging task and a bottleneck for building a practical discourse parser. Previous research usually makes use of one kind of discourse framework such as PDTB or RST to improve the classification performance on discourse relations. Actually, under different discourse annotation frameworks, there exist multiple corpora which have internal connections. To exploit the combination of different discourse corpora, we design related discourse classification tasks specific to a corpus, and propose a novel Convolutional Neural Network embedded multi-task learning system to synthesize these tasks by learning both unique and shared representations for each task. The experimental results on the PDTB implicit discourse relation classification task demonstrate that our model achieves significant gains over baseline systems.
Yang Liu 0124, Sujian Li, Xiaodong Zhang 0022, Zhifang Sui
AAAI2
2016 RBPB: Regularization-Based Pattern Balancing Method for Event Extraction
abstract
Event extraction is a particularly challenging information extraction task, which intends to identify and classify event triggers and arguments from raw text.In recent works, when determining event types (trigger classification), most of the works are either pattern-only or feature-only.However, although patterns cannot cover all representations of an event, it is still a very important feature.In addition, when identifying and classifying arguments, previous works consider each candidate argument separately while ignoring the relationship between arguments.This paper proposes a Regularization-Based Pattern Balancing Method (RBPB).Inspired by the progress in representation learning, we use trigger embedding, sentence-level embedding and pattern features together as our features for trigger classification so that the effect of patterns and other useful features can be balanced.In addition, RBPB uses a regularization method to take advantage of the relationship between arguments.Experiments show that we achieve results better than current state-of-art equivalents.
Lei Sha, Jing Liu 0022, Chin-Yew Lin, Sujian Li, Baobao Chang, Zhifang Sui
ACL (1)4
2016 AttSum: Joint Learning of Focusing and Summarization with Neural Attention
abstract
Query relevance ranking and sentence saliency ranking are the two main tasks in extractive query-focused summarization. Previous supervised summarization systems often perform the two tasks in isolation. However, since reference summaries are the trade-off between relevance and saliency, using them as supervision, neither of the two rankers could be trained well. This paper proposes a novel summarization system called AttSum, which tackles the two tasks jointly. It automatically learns distributed representations for sentences as well as the document cluster. Meanwhile, it applies the attention mechanism to simulate the attentive reading of human behavior when a query is given. Extensive experiments are conducted on DUC query-focused summarization benchmark datasets. Without using any hand-crafted features, AttSum achieves competitive performance. We also observe that the sentences recognized to focus on the query indeed meet the query need.
Ziqiang Cao, Wenjie Li 0002, Sujian Li, Furu Wei, Yanran Li
COLING3
2016 Towards Time-Aware Knowledge Graph Completion
abstract
Knowledge graph (KG) completion adds new facts to a KG by making inferences from existing facts. Most existing methods ignore the time information and only learn from time-unknown fact triples. In dynamic environments that evolve over time, it is important and challenging for knowledge graph completion models to take into account the temporal aspects of facts. In this paper, we present a novel time-aware knowledge graph completion model that is able to predict links in a KG using both the existing facts and the temporal information of the facts. To incorporate the happening time of facts, we propose a time-aware KG embedding model using temporal order information among facts. To incorporate the valid time of facts, we propose a joint time-aware inference model based on Integer Linear Programming (ILP) using temporal consistencyinformationasconstraints. Wefurtherintegratetwomodelstomakefulluseofglobal temporal information. We empirically evaluate our models on time-aware KG completion task. Experimental results show that our time-aware models achieve the state-of-the-art on temporal facts consistently.
Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Baobao Chang, Sujian Li, Zhifang Sui
COLING6
2016 Reading and Thinking: Re-read LSTM Unit for Textual Entailment Recognition
abstract
Recognizing Textual Entailment (RTE) is a fundamentally important task in natural language processing that has many applications. The recently released Stanford Natural Language Inference (SNLI) corpus has made it possible to develop and evaluate deep neural network methods for the RTE task. Previous neural network based methods usually try to encode the two sentences (premise and hypothesis) and send them together into a multi-layer perceptron to get their entailment type, or use LSTM-RNN to link two sentences together while using attention mechanic to enhance the model’s ability. In this paper, we propose to use the re-read mechanic, which means to read the premise again and again while reading the hypothesis. After read the premise again, the model can get a better understanding of the premise, which can also affect the understanding of the hypothesis. On the contrary, a better understanding of the hypothesis can also affect the understanding of the premise. With the alternative re-read process, the model can “think” of a better decision of entailment type. We designed a new LSTM unit called re-read LSTM (rLSTM) to implement this “thinking” process. Experiments show that we achieve results better than current state-of-the-art equivalents.
Lei Sha, Baobao Chang, Zhifang Sui, Sujian Li
COLING4
2016 News Stream Summarization using Burst Information Networks
abstract
This paper studies summarizing key information from news streams. We propose simple yet effective models to solve the problem based on a novel and promising representation of text streams – Burst Information Networks (BINets). A BINet can be aware of redundant information, allows global analysis of a text stream, and can be efficiently built and dynamically updated, which perfectly fits the demands of text stream summarization. Extensive experiments show that the BINet-based approaches are not only efficient and can be used in a real-time online summarization setting, but also can generate high-quality summaries, outperforming the state-of-the-art approach.
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Sujian Li, Ming Zhou 0001, Zhifang Sui
EMNLP4
2016 Encoding Temporal Information for Time-Aware Link Prediction
abstract
Most existing knowledge base (KB) embedding methods solely learn from time-unknown fact triples but neglect the temporal information in the knowledge base.In this paper, we propose a novel time-aware KB embedding approach taking advantage of the happening time of facts.Specifically, we use temporal order constraints to model transformation between time-sensitive relations and enforce the embeddings to be temporally consistent and more accurate.We empirically evaluate our approach in two tasks of link prediction and triple classification.Experimental results show that our method outperforms other baselines on the two tasks consistently.
Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui
EMNLP5
2016 Recognizing Implicit Discourse Relations via Repeated Reading: Neural Networks with Multi-Level Attention
abstract
Recognizing implicit discourse relations is a challenging but important task in the field of Natural Language Processing.For such a complex text processing task, different from previous studies, we argue that it is necessary to repeatedly read the arguments and dynamically exploit the efficient features useful for recognizing discourse relations.To mimic the repeated reading strategy, we propose the neural networks with multi-level attention (NNMA), combining the attention mechanism and external memories to gradually fix the attention on some specific words helpful to judging the discourse relations.Experiments on the PDTB dataset show that our proposed method achieves the state-ofart results.The visualization of the attention weights also illustrates the progress that our model observes the arguments on each level and progressively locates the important words.
Yang Liu 0124, Sujian Li
EMNLP2
2016 Capturing Argument Relationship for Chinese Semantic Role Labeling
abstract
In this paper, we capture the argument relationships for Chinese semantic role labeling task, and improve the task's performance with the help of argument relationships.We split the relationship between two candidate arguments into two categories: (1) Compatible arguments: if one candidate argument belongs to a given predicate, then the other is more likely to belong to the same predicate; (2) Incompatible arguments: if one candidate argument belongs to a given predicate, then the other is less likely to belong to the same predicate.However, previous works did not explicitly model argument relationships.We use a simple maximum entropy classifier to capture the two categories of argument relationships and test its performance on the Chinese Proposition Bank (CPB).The experiments show that argument relationships is effective in Chinese semantic role labeling task.
Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui, Tingsong Jiang
EMNLP2
2016 A New Approach to Modeling and Analyzing Timed Compatibility of Service Composition under Temporal Constraints
abstract
In enterprises, timed compatibility (i. e., temporal constraint satisfiability) have become an important factor to guarantee the timely completion of service composition so that meet the requirements of customers and ensure the success execution of enterprises. However, existing researches do not fully investigate the mixture situations where distributions of service durations are complex and diverse, and consider both the uncertainty of queue time and operation time of services. In this paper, we propose a new approach to modeling and analyzing timed compatibility of service composition where time durations are mixture distributed. First, we calculate the duration of activities in service with considering both the queue time and operation time. Second, we obtain the time information of interactive activities in service. Finally, we get the time information of paths in services and check the temporal constraints. Furthermore, a real-life case is used to illustrate the efficiency and effectiveness of our approach.
Yanxue Xing, Sujian Li, Yanhua Du, Helan Liang
ISPDC2
2016 Joint Learning Templates and Slots for Event Schema Induction
abstract
Automatic event schema induction (AESI) means to extract meta-event from raw text, in other words, to find out what types (templates) of event may exist in the raw text and what roles (slots) may exist in each event type.In this paper, we propose a joint entity-driven model to learn templates and slots simultaneously based on the constraints of templates and slots in the same sentence.In addition, the entities' semantic information is also considered for the inner connectivity of the entities.We borrow the normalized cut criteria in image segmentation to divide the entities into more accurate template clusters and slot clusters.The experiment shows that our model gains a relatively higher result than previous work.
Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui
HLT-NAACL2
2016 Relation Classification Via Modeling Augmented Dependency Paths
abstract
Previous research on relation classification has verified the effectiveness of using dependency shortest paths or dependency subtrees. How to efficiently unify these two kinds of dependency information in relation classification is still an open problem. In this paper, we propose a novel structure, termed augmented dependency path (ADP), which is composed of the shortest dependency path between two entities and the subtrees attached to the shortest path. To exploit the semantic representation behind the ADP structure, we develop the dependency-based neural networks (DepNN) model which combines the advantages of the recursive neural network (RNN) and the convolutional neural network (CNN). In DepNN, RNN is designed to model the dependency subtrees since it is good at capturing the hierarchical structures. Then, the semantic representation in subtrees is passed to the nodes on the shortest path and CNN is used to get the most important features on the ADP. Experiments on the SemEval-2010 dataset show that the ADP structure including both the shortest dependency path and the attached subtrees is helpful to classify the semantic relations between two entities and our proposed method can achieve the state-of-the-art performance.
Yang Liu 0124, Sujian Li, Furu Wei, Heng Ji 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 A Novel Neural Topic Model and Its Supervised Extension
abstract
Topic modeling techniques have the benefits of modeling words and documents uniformly under a probabilistic framework. However, they also suffer from the limitations of sensitivity to initialization and unigram topic distribution, which can be remedied by deep learning techniques. To explore the combination of topic modeling and deep learning techniques, we first explain the standard topic modelfrom the perspective of a neural network. Based on this, we propose a novel neural topic model (NTM) where the representation of words and documents are efficiently and naturally combined into a uniform framework. Extending from NTM, we can easily add a label layer and propose the supervised neural topic model (sNTM) to tackle supervised tasks. Experiments show that our models are competitive in both topic discovery and classification/regression tasks.
Ziqiang Cao, Sujian Li, Yang Liu 0124, Wenjie Li 0002, Heng Ji 0001
AAAI2
2015 Ranking with Recursive Neural Networks and Its Application to Multi-Document Summarization
abstract
We develop a Ranking framework upon Recursive Neural Networks (R2N2) to rank sentences for multi-document summarization. It formulates the sentence ranking task as a hierarchical regression process, which simultaneously measures the salience of a sentence and its constituents (e.g., phrases) in the parsing tree. This enables us to draw on word-level to sentence-level supervisions derived from reference summaries.In addition, recursive neural networks are used to automatically learn ranking features over the tree, with hand-crafted feature vectors of words as inputs. Hierarchical regressions are then conducted with learned features concatenating raw features.Ranking scores of sentences and words are utilized to effectively select informative and non-redundant sentences to generate summaries.Experiments on the DUC 2001, 2002 and 2004 multi-document summarization datasets show that R2N2 outperforms state-of-the-art extractive summarization approaches.
Ziqiang Cao, Furu Wei, Li Dong 0004, Sujian Li, Ming Zhou 0001
AAAI4
2015 Bring you to the past: Automatic Generation of Topically Relevant Event Chronicles
abstract
Tao Ge, Wenzhe Pei, Heng Ji, Sujian Li, Baobao Chang, Zhifang Sui. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Tao Ge 0001, Wenzhe Pei, Heng Ji 0001, Sujian Li, Baobao Chang, Zhifang Sui
ACL (1)4
2015 Context-aware Entity Morph Decoding
abstract
Boliang Zhang, Hongzhao Huang, Xiaoman Pan, Sujian Li, Chin-Yew Lin, Heng Ji, Kevin Knight, Zhen Wen, Yizhou Sun, Jiawei Han, Bulent Yener. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Boliang Zhang, Hongzhao Huang, Xiaoman Pan, Sujian Li, Chin-Yew Lin, Heng Ji 0001, Kevin Knight, Yizhou Sun, Jiawei Han 0001, Bülent Yener
ACL (1)4
2015 Component-Enhanced Chinese Character Embeddings
abstract
Distributed word representations are very useful for capturing semantic information and have been successfully applied in a variety of NLP tasks, especially on En-glish. In this work, we innovatively de-velop two component-enhanced Chinese character embedding models and their bi-gram extensions. Distinguished from En-glish word embeddings, our models ex-plore the compositions of Chinese char-acters, which often serve as semantic in-dictors inherently. The evaluations on both word similarity and text classification demonstrate the effectiveness of our mod-els. 1
Yanran Li, Wenjie Li 0002, Fei Sun 0001, Sujian Li
EMNLP4
2015 Recognizing Textual Entailment Using Probabilistic Inference
abstract
Recognizing Text Entailment (RTE) plays an important role in NLP applications including question answering, information retrieval, etc.In recent work, some research explore "deep" expressions such as discourse commitments or strict logic for representing the text.However, these expressions suffer from the limitation of inference inconvenience or translation loss.To overcome the limitations, in this paper, we propose to use the predicate-argument structures to represent the discourse commitments extracted from text.At the same time, with the help of the YAGO knowledge, we borrow the distant supervision technique to mine the implicit facts from the text.We also construct a probabilistic network for all the facts and conduct inference to judge the confidence of each fact for RTE.The experimental results show that our proposed method achieves a competitive result compared to the previous work.
Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui, Tingsong Jiang
EMNLP2
2015 Why Read if You Can Scan? Trigger Scoping Strategy for Biographical Fact Extraction
abstract
The rapid growth of information sources brings a unique challenge to biographical information extraction: how to find specific facts without having to read all the words. An effective solution is to follow the human scanning strategy which keeps a specific keyword in mind and searches within a specific scope. In this paper, we mimic a scanning process to extract biographical facts. We use event and relation triggers as
Dian Yu 0001, Heng Ji 0001, Sujian Li, Chin-Yew Lin
HLT-NAACL3
2014 Text-level Discourse Dependency Parsing
abstract
Previous researches on Text-level discourse parsing mainly made use of constituency structure to parse the whole document into one discourse tree. In this paper, we present the limitations of constituency based dis-course parsing and first propose to use de-pendency structure to directly represent the relations between elementary discourse units (EDUs). The state-of-the-art depend-ency parsing techniques, the Eisner algo-rithm and maximum spanning tree (MST) algorithm, are adopted to parse an optimal discourse dependency tree based on the arc-factored model and the large-margin learn-ing techniques. Experiments show that our discourse dependency parsers achieve a competitive performance on text-level dis-course parsing. 1
Sujian Li, Liang Wang 0046, Ziqiang Cao, Wenjie Li 0002
ACL (1)1
2014 Query-focused Multi-Document Summarization: Combining a Topic Model with Graph-based Semi-supervised Learning
Yanran Li, Sujian Li
COLING2
2014 Joint Learning of Chinese Words, Terms and Keywords
abstract
Previous work often used a pipelined framework where Chinese word segmentation is followed by term extraction and keyword extraction.Such framework suffers from error propagation and is unable to leverage information in later modules for prior components.In this paper, we propose a four-level Dirichlet Process based model (DP-4) to jointly learn the word distributions from the corpus, domain and document levels simultaneously.Based on the DP-4 model, a sentence-wise Gibbs sampler is adopted to obtain proper segmentation results.Meanwhile, terms and keywords are acquired in the sampling process.Experimental results have shown the effectiveness of our method.
Ziqiang Cao, Sujian Li, Heng Ji 0001
EMNLP2
2014 Constructing Information Networks Using One Single Model
abstract
In this paper, we propose a new frame-work that unifies the output of three infor-mation extraction (IE) tasks- entity men-tions, relations and events as an informa-tion network representation, and extracts all of them using one single joint model based on structured prediction. This novel formulation allows different parts of the information network fully interact with each other. For example, many rela-tions can now be considered as the re-sultant states of events. Our approach achieves substantial improvements over traditional pipelined approaches, and sig-nificantly advances state-of-the-art end-to-end event argument extraction. 1
Qi Li 0014, Heng Ji 0001, Sujian Li
EMNLP4
2013 Event-Based Time Label Propagation for Automatic Dating of News Articles
abstract
Since many applications such as timeline summaries and temporal IR involving temporal analysis rely on document timestamps, the task of automatic dating of documents has been increasingly important.Instead of using feature-based methods as conventional models, our method attempts to date documents in a year level by exploiting relative temporal relations between documents and events, which are very effective for dating documents.Based on this intuition, we proposed an eventbased time label propagation model called confidence boosting in which time label information can be propagated between documents and events on a bipartite graph.The experiments show that our event-based propagation model can predict document timestamps in high accuracy and the model combined with a MaxEnt classifier outperforms the state-ofthe-art method for this task especially when the size of the training set is small.
Tao Ge 0001, Baobao Chang, Sujian Li, Zhifang Sui
EMNLP3
2013 A novel topic model for automatic term extraction
abstract
Automatic term extraction (ATE) aims at extracting domain-specific terms from a corpus of a certain domain. Termhood is one essential measure for judging whether a phrase is a term. Previous researches on termhood mainly depend on the word frequency information. In this paper, we propose to compute termhood based on semantic representation of words. A novel topic model, namely i-SWB, is developed to map the domain corpus into a latent semantic space, which is composed of some general topics, a background topic and a documents-specific topic. Experiments on four domains demonstrate that our approach outperforms the state-of-the-art ATE approaches.
Sujian Li, Wenjie Li 0002, Baobao Chang
SIGIR1
2013 Real Time Event Detection in Twitter
Feida Zhu 0001, Jing Jiang 0001, Sujian Li
WAIM4
2013 An Improved Genetic Algorithm for Service Selection under Temporal Constraints in Cloud Computing
Helan Liang, Yanhua Du, Sujian Li
WISE (2)3
2013 A progressive sentence selection strategy for document summarization
Ouyang You, Wenjie Li 0002, Renxian Zhang, Sujian Li, Qin Lu 0001
Inf. Process. Manag.4
2013 Exploring hypergraph-based semi-supervised ranking for query-oriented summarization
Wei Wang 0013, Sujian Li, Wenjie Li 0002, Furu Wei
Inf. Sci.2
2013 A Novel Feature-based Bayesian Model for Query Focused Multi-document Summarization
abstract
Supervised learning methods and LDA based topic model have been successfully applied in the field of multi-document summarization. In this paper, we propose a novel supervised approach that can incorporate rich sentence features into Bayesian topic models in a principled way, thus taking advantages of both topic model and feature based supervised learning methods. Experimental results on DUC2007, TAC2008 and TAC2009 demonstrate the effectiveness of our approach.
Sujian Li
Trans. Assoc. Comput. Linguistics2
2012 Exploring simultaneous keyword and key sentence extraction: improve graph-based ranking using wikipedia
abstract
Summarization and Keyword Selection are two important tasks in NLP community. Although both aim to summarize the source articles, they are usually treated separately by using sentences or words. In this paper, we propose a two-level graph based ranking algorithm to generate summarization and extract keywords at the same time. Previous works have reached a consensus that important sentence is composed by important keywords. In this paper, we further study the mutual impact between them through context analysis. We use Wikipedia to build a two-level concept-based graph, instead of traditional term-based graph, to express their homogenous relationship and heterogeneous relationship. We run PageRank and HITS rank on the graph to adjust both homogenous and heterogeneous relationships. A more reasonable relatedness value will be got for key sentence selection and keyword selection. We evaluate our algorithm on TAC 2011 data set. Traditional term-based approach achieves a score of 0.255 in ROUGE-1 and a score of 0.037 and ROUGE-2 and our approach can improve them to 0.323 and 0.048 separately.
Sujian Li
CIKM4
2012 Update Summarization using a Multi-level Hierarchical Dirichlet Process Model
Sujian Li, Baobao Chang
COLING2
2012 Implicit Discourse Relation Recognition by Selecting Typical Training Examples
Sujian Li, Wenjie Li 0002
COLING2
2012 Constructing Chinese Abbreviation Dictionary: A Stacked Approach
Longkai Zhang, Sujian Li, Houfeng Wang, Ni Sun, Xinfan Meng
COLING2
2012 Joint Learning for Coreference Resolution with Markov Logic
Yang Song 0021, Jing Jiang 0001, Wayne Xin Zhao, Sujian Li, Houfeng Wang
EMNLP-CoNLL4
2012 Entity-centric topic-oriented opinion summarization in twitter
abstract
Microblogging services, such as Twitter, have become popular channels for people to express their opinions towards a broad range of topics. Twitter generates a huge volume of instant messages (i.e. tweets) carrying users' sentiments and attitudes every minute, which both necessitates automatic opinion summarization and poses great challenges to the summarization system. In this paper, we study the problem of opinion summarization for entities, such as celebrities and brands, in Twitter. We propose an entity-centric topic-based opinion summarization framework, which aims to produce opinion summaries in accordance with topics and remarkably emphasizing the insight behind the opinions. To this end, we first mine topics from #hashtags, the human-annotated semantic tags in tweets. We integrate the #hashtags as weakly supervised information into topic modeling algorithms to obtain better interpretation and representation for calculating the similarity among them, and adopt Affinity Propagation algorithm to group #hashtags into coherent topics. Subsequently, we use templates generalized from paraphrasing to identify tweets with deep insights, which reveal reasons, express demands or reflect viewpoints. Afterwards, we develop a target (i.e. entity) dependent sentiment classification approach to identifying the opinion towards a given target (i.e. entity) of tweets. Finally, the opinion summary is generated through integrating information from dimensions of topic, opinion and insight, as well as other factors (e.g. topic relevancy, redundancy and language styles) in an unified optimization framework. We conduct extensive experiments on a real-life data set to evaluate the performance of individual opinion summarization modules as well as the quality of the produced summary. The promising experiment results show the effectiveness of the proposed framework and algorithms.
Xinfan Meng, Furu Wei, Ming Zhou 0001, Sujian Li, Houfeng Wang
KDD5
2011 CoRankBayes: bayesian learning to rank under the co-training framework and its application in keyphrase extraction
abstract
Recently, learning to rank algorithms have become a popular and effective tool for ordering objects (e.g. terms) according to their degrees of importance. The contribution of this paper is that we propose a simple and fast learning to rank model RankBayes and embed it in the co-training framework. The detailed proof is given that Naïve Bayes algorithm can be used to implement a learning to rank model. To solve the problem of two-model inconsistency, an ingenious approach is put forward to rank all the phrases by making use of the labeled results of two RankBayes models. Experimental results show that the proposed approach is promising in solving ranking problems.
Chen Wang 0036, Sujian Li
CIKM2
2011 Applying regression models to query-focused multi-document summarization
Ouyang You, Wenjie Li 0002, Sujian Li, Qin Lu 0001
Inf. Process. Manag.3
2010 Intertopic information mining for query-based summarization
abstract
Abstract In this article, the authors address the problem of sentence ranking in summarization. Although most existing summarization approaches are concerned with the information embodied in a particular topic (including a set of documents and an associated query) for sentence ranking, they propose a novel ranking approach that incorporates intertopic information mining. Intertopic information, in contrast to intratopic information, is able to reveal pairwise topic relationships and thus can be considered as the bridge across different topics. In this article, the intertopic information is used for transferring word importance learned from known topics to unknown topics under a learning‐based summarization framework. To mine this information, the authors model the topic relationship by clustering all the words in both known and unknown topics according to various kinds of word conceptual labels, which indicate the roles of the words in the topic. Based on the mined relationships, we develop a probabilistic model using manually generated summaries provided for known topics to predict ranking scores for sentences in unknown topics. A series of experiments have been conducted on the Document Understanding Conference (DUC) 2006 data set. The evaluation results show that intertopic information is indeed effective for sentence ranking and the resultant summarization system performs comparably well to the best‐performing DUC participating systems on the same data set.
Ouyang You, Wenjie Li 0002, Sujian Li, Qin Lu 0001
J. Assoc. Inf. Sci. Technol.3
2009 HyperSum: hypergraph based semi-supervised sentence ranking for query-oriented summarization
abstract
Graph based sentence ranking algorithms such as PageRank and HITS have been successfully used in query-oriented summarization. With these algorithms, the documents to be summarized are often modeled as a text graph where nodes represent sentences and edges represent pairwise similarity relationships between two sentences. A deficiency of conventional graph modeling is its incapability of naturally and effectively representing complex group relationships shared among multiple objects. Simply squeezing complex relationships into pairwise ones will inevitably lead to loss of information which can be useful for ranking and learning. In this paper, we propose to take advantage of hypergraph, i.e. a generalization of graph, to remedy this defect. In a text hypergraph, nodes still represent sentences, yet hyperedges are allowed to connect more than two sentences. With a text hypergraph, we are thus able to integrate both group relationships formulated among multiple sentences and pairwise relationships formulated between two sentences in a unified framework. As essential work, it is first addressed in the paper that how a text hypergraph can be built for summarization by applying clustering techniques. Then, a hypergraph based semi-supervised sentence ranking algorithm is developed for query-oriented extractive summarization, where the influence of query is propagated to sentences through the structure of the constructed text hypergraph. When evaluated on DUC data sets, performance of the proposed approach is remarkable.
Wei Wang 0013, Furu Wei, Wenjie Li 0002, Sujian Li
CIKM4
2007 Developing learning strategies for topic-based summarization
abstract
Most up-to-date well-behaved topic-based summarization systems are built upon the extractive framework. They score the sentences based on the associated features by manually assigning or experimentally tuning the weights of the features. In this paper, we discuss how to develop learning strategies in order to obtain the optimal feature weights automatically, which can be used for assigning a sound score to a sentence characterized with a set of features. The two fundamental issues are about training data and learning models. To save the costly manual annotation time and effort, we construct the training data by labeling the sentence with a "true" score calculated according to human summaries. The Support Vector Regression (SVR) model is then used to learn how to relate the "true" score of the sentence to its features. Once the relations have been mathematically modeled, SVR is able to predict the "estimated" score for any given sentence. The evaluations by ROUGE-2 criterion on DUC 2006 and DUC 2005 document sets demonstrate the competitiveness and the adaptability of the proposed approaches.
Ouyang You, Sujian Li, Wenjie Li 0002
CIKM2
2006 Interaction between Lexical Base and Ontology with Formal Concept Analysis
Sujian Li, Qin Lu 0001, Wenjie Li 0002, Ruifeng Xu 0001
LREC1
2006 The Design and Construction of A Chinese Collocation Bank
Ruifeng Xu 0001, Qin Lu 0001, Sujian Li
LREC3
2004 A Combining Approach to Automatic Keyphrases Indexing for Chinese News Documents
Houfeng Wang, Sujian Li, Shiwen Yu, Byeong Kwu Kang
CICLing2
2003 News-Oriented Keyword Indexing with Maximum Entropy Principle
Sujian Li, Houfeng Wang, Shiwen Yu, Chengsheng Xin
PACLIC1
2002 Semantic Computation in a Chinese Question-Answering System
Sujian Li, Jian Zhang 0003, Huang Xiong, Shuo Bai, Qun Liu 0001
J. Comput. Sci. Technol.1
2001 Automatic extraction of lexical relations from Chinese machine readable dictionary
abstract
Lexical relations are very important for NLP. Most previous work to get them is done by hand. In this paper, we describe an automated strategy which exploits a machine readable dictionary (MRD) to construct a richly-structured network of lexical relations. In our system lexical relations include five basic semantic relations, two phonetic relations and one orthographic relation. These relations constitute the basic framework of our lexical network. Then we present an approach to use heuristic functions to extract semantic relations while we conduct syntactic parsing. Experimental results demonstrate that our method is effective.
Sujian Li, Qun Liu 0001, Shuo Bai, Xueqi Cheng 0001
SMC1