Baobao Chang

dblp:91/6051 · DBLP profile ↗
← Back
106ranked-venue papers
3as first author
47since 2021 · last 2026
0000-0003-2824-6750ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 102 · 3 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021
YearPublicationVenuePosition
2026 Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning
abstract
Teaching large language models (LLMs) to be faithful in the provided context is crucial for building reliable information-seeking systems. Therefore, we propose a systematic framework, CANOE, to reduce faithfulness hallucinations of LLMs across different downstream tasks without human annotations. Specifically, we first synthesize short-form question-answering (QA) data with four diverse tasks to construct high-quality and easily verifiable training data without human annotation. Also, we propose Dual-GRPO, a rule-based reinforcement learning method that includes three tailored rule-based rewards derived from synthesized short-form QA data, while simultaneously optimizing both short-form and long-form response generation. Notably, Dual-GRPO eliminates the need to manually label preference data to train reward models and avoids over-optimizing short-form generation when relying only on the synthesized short-form QA data. Experimental results show that CANOE greatly improves the faithfulness of LLMs across 11 different tasks, even outperforming the most advanced LLMs, e.g., GPT-4o and OpenAI o1.
Shuzheng Si, Haozhe Zhao, Yuzhuo Bai, Zhitong Wang, Bofei Gao, Kangyang Luo, Wenhao Li 0003, Yufei Huang 0008, Gang Chen 0039, Fanchao Qi, Minjia Zhang, Baobao Chang, Maosong Sun 0001
AAAI13
2026 A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Task
abstract
Shuzheng Si, Haozhe Zhao, Kangyang Luo, Gang Chen, Fanchao Qi, Minjia Zhang, Baobao Chang, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shuzheng Si, Haozhe Zhao, Kangyang Luo, Gang Chen 0039, Fanchao Qi, Minjia Zhang, Baobao Chang, Maosong Sun 0001
ACL (1)7
2025 Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
abstract
Shuzheng Si, Haozhe Zhao, Gang Chen, Cheng Gao, Yuzhuo Bai, Zhitong Wang, Kaikai An, Kangyang Luo, Chen Qian, Fanchao Qi, Baobao Chang, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shuzheng Si, Haozhe Zhao, Gang Chen 0039, Yuzhuo Bai, Zhitong Wang, Kaikai An, Kangyang Luo, Fanchao Qi, Baobao Chang, Maosong Sun 0001
ACL (1)11
2025 CCAgent: Coordinating Collaborative Data Scaling for Operating System Agents via Web3
Liang Chen 0024, Haozhe Zhao, Yinzhen Huang, Tsekai Lin, Weichu Xie, Peiyi Wang, Runxin Xu, Ming Wu 0007, Baobao Chang
CIKM11
2025 UltraIF: Advancing Instruction Following from the Wild
abstract
Instruction-following made modern large language models (LLMs) helpful assistants.However, the key to taming LLMs on complex instructions remains mysterious, for that there are huge gaps between models trained by opensource community and those trained by leading companies.To bridge the gap, we propose a simple and scalable approach ULTRAIF for building LLMs that can follow complex instructions with open-source data.ULTRAIF first decomposes real-world user prompts into simpler queries, constraints, and corresponding evaluation questions for the constraints.Then, we train an UltraComposer to compose constraintassociated prompts with evaluation questions.This prompt composer allows us to synthesize complicated instructions as well as filter responses with evaluation questions.In our experiment, for the first time, we successfully align LLaMA-3.1-8B-Base to catch up with its instruct version on 5 instruction-following benchmarks without any benchmark information, using only 8B model as response generator and evaluator.The aligned model also achieved competitive scores on other benchmarks.Moreover, we also show that ULTRAIF could further improve LLaMA-3.1-8B-Instruct through self-alignment, motivating broader use cases for the method.Our code is available at https://github.com/kkk-an/UltraIF.
Kaikai An, Ganqu Cui, Shuzheng Si, Baobao Chang
EMNLP7
2025 Thread: A Logic-Based Data Organization Paradigm for How-To Question Answering with Retrieval Augmented Generation
abstract
Kaikai An, Fangkai Yang, Liqun Li, Junting Lu, Sitao Cheng, Shuzheng Si, Lu Wang, Pu Zhao, Lele Cao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Baobao Chang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Kaikai An, Fangkai Yang, Liqun Li, Junting Lu, Sitao Cheng, Shuzheng Si, Lu Wang 0029, Pu Zhao 0004, Le-le Cao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Baobao Chang
EMNLP13
2025 GATEAU: Selecting Influential Samples for Long Context Alignment
abstract
Shuzheng Si, Haozhe Zhao, Gang Chen, Yunshui Li, Kangyang Luo, Chuancheng Lv, Kaikai An, Fanchao Qi, Baobao Chang, Maosong Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Shuzheng Si, Haozhe Zhao, Gang Chen 0039, Yunshui Li, Kangyang Luo, Chuancheng Lv, Kaikai An, Fanchao Qi, Baobao Chang, Maosong Sun 0001
EMNLP9
2025 Looking Beyond Text: Reducing Language Bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
abstract
Large vision-language models (LVLMs) have achieved impressive results in vision-language tasks.However, LVLMs suffer from hallucinations caused by language bias, which neglects images while over-relying on text.We identify two reasons for the bias: 1).Different training scales between the LLM pretraining and LVLM alignment stage.2).The learned inference bias due to short-term dependency of text data.Therefore, we propose LACING, designed to address such bias with MuLtimodal DuAlattention MeChanIsm (MDA) aNd Soft-Image Guidance (SIG).Specifically, MDA adopts a parallel dual-attention mechanism that constructs separate attention for visual and text inputs to enhance integration of visual inputs across model.SIG uses a learnable soft visual prompt during training and inference to replace visual inputs, designed to compel LVLMs to prioritize text inputs during inference.Experiments across different model architectures and scales demonstrate that LACING effectively debiases LVLMs from their language bias, enhancing visual comprehension and reducing hallucinations without additional resources.
Haozhe Zhao, Shuzheng Si, Liang Chen 0024, Yichi Zhang 0010, Maosong Sun 0001, Baobao Chang, Minjia Zhang
EMNLP6
2025 A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
abstract
This work tackles the information loss bottleneck of vector-quantization (VQ) autoregressive image generation by introducing a novel model architecture called the 2-Dimensional Autoregression (DnD) Transformer. The DnD-Transformer predicts more codes for an image by introducing a new direction, **model depth**, along with the sequence length. Compared to 1D autoregression and previous work using similar 2D image decomposition such as RQ-Transformer, the DnD-Transformer is an end-to-end model that can generate higher quality images with the same backbone model size and sequence length, opening a new optimization perspective for autoregressive image generation. Furthermore, our experiments reveal that the DnD-Transformer's potential extends beyond generating natural images. It can even generate images with rich text and graphical elements in a self-supervised manner, demonstrating an understanding of these combined modalities. This has not been previously demonstrated for popular vision generative models such as diffusion models, showing a spark of vision-language intelligence when trained solely on images. Code, datasets and models are open at https://github.com/chenllliang/DnD-Transformer.
Liang Chen 0024, Sinan Tan, Zefan Cai, Weichu Xie, Haozhe Zhao, Yichi Zhang 0010, Junyang Lin, Jinze Bai, Tianyu Liu 0001, Baobao Chang
ICLR10
2025 Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language Models
abstract
Recent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54% and 52.55% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning.
Bofei Gao, Feifan Song 0001, Zhe Yang 0013, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li 0039, Chenghao Ma, Liang Chen 0024, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang 0009, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu 0001, Baobao Chang
ICLR20
2025 MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
abstract
Jinsheng Huang, Liang Chen, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan, Haozhe Zhao, Zhihui Guo, Yichi Zhang, Jingyang Yuan, Wei Ju, Luchen Liu, Tianyu Liu, Baobao Chang, Ming Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jinsheng Huang, Liang Chen 0024, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan 0016, Haozhe Zhao, Zhihui Guo, Yichi Zhang 0010, Jingyang Yuan, Wei Ju 0001, Luchen Liu, Tianyu Liu 0001, Baobao Chang, Ming Zhang 0004
NAACL (Long Papers)15
2024 Improving Event Definition Following For Zero-Shot Event Detection
abstract
Zefan Cai, Po-Nien Kung, Ashima Suvarna, Mingyu Ma, Hritik Bansal, Baobao Chang, P. Jeffrey Brantingham, Wei Wang, Nanyun Peng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zefan Cai, Po-Nien Kung, Ashima Suvarna, Mingyu Derek Ma, Hritik Bansal, Baobao Chang, P. Jeffrey Brantingham, Wei Wang 0010, Nanyun Peng 0001
ACL (1)6
2024 An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Liang Chen 0024, Haozhe Zhao, Tianyu Liu 0001, Shuai Bai, Junyang Lin, Chang Zhou 0005, Baobao Chang
ECCV (81)7
2024 A Survey on In-context Learning
abstract
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, Lei Li, Zhifang Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Qingxiu Dong, Lei Li 0039, Damai Dai, Jingyuan Ma, Rui Li 0094, Heming Xia, Jingjing Xu 0001, Zhiyong Wu 0011, Baobao Chang, Xu Sun 0001, Lei Li 0005, Zhifang Sui
EMNLP10
2024 MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
abstract
Since the resurgence of deep learning, vision-language models (VLMs) enhanced by large language models (LLMs) have grown exponentially in popularity. However, while LLMs can utilize extensive background knowledge and task information with in-context learning, most VLMs still struggle with understanding complex multi-modal prompts with multiple images, making VLMs less effective in downstream vision-language tasks. In this paper, we address the limitation above by 1) introducing vision-language Model with **M**ulti-**M**odal **I**n-**C**ontext **L**earning(MMICL), a new approach to allow the VLM to deal with multi-modal inputs efficiently; 2) proposing a novel context scheme to augment the in-context learning ability of the VLM; 3) constructing the Multi-modal In-Context Learning (MIC) dataset, designed to enhance the VLM's ability to understand complex multi-modal prompts. Our experiments confirm that MMICL achieves new state-of-the-art zero-shot performance on a wide range of general vision-language tasks, especially for complex benchmarks, including MME and MMBench. Our analysis demonstrates that MMICL effectively tackles the challenge of complex multi-modal prompt understanding and emerges the impressive ICL ability. Furthermore, we observe that MMICL successfully alleviates language bias in VLMs, a common issue for VLMs that often leads to hallucination when faced with extensive textual context. Our code, dataset, dataset tool, and model are available at https://github.com/PKUnlp-icler/MIC.
Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma 0001, Kaikai An, Liang Chen 0024, Zixuan Liu 0001, Sheng Wang 0012, Wenjuan Han, Baobao Chang
ICLR10
2024 VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness
abstract
Finetuning a pretrained vision model (PVM) is a common technique for learning downstream vision tasks. The conventional finetuning process with the randomly sampled data points results in diminished training efficiency. To address this drawback, we propose a novel approach, Vision- languag e C ollaborative A ctive F inetuning (VeCAF). VeCAF optimizes a parametric data selection model by incorporating the training objective of the model being tuned. Effectively, this guides the PVM towards the performance goal with improved data and computational efficiency.With the ever-growing feasibility of acquiring labels and natural language annotations of image data through web-scale crawling, we exploit the inherent semantic richness of the text embedding space and utilize text embeddings of image annotations to augment PVM image features for better data selection and finetuning. Furthermore, the flexibility of text-domain augmentation gives VeCAF the unique ability to handle out-of-distribution scenarios without external augmented data. Extensive experiments show the leading performance and high efficiency of VeCAF that is superior to baselines in both in-distribution and out-of-distribution image classification tasks. On ImageNet, VeCAF needs up to 3.3× less training batches to reach the target performance compared to full fine-tuning and achieves an accuracy improvement of 2.8% over active SOTA fine-tuning methods with the same number of batches. Our code is now available at https://github.com/RoyZry98/VeCAF-Pytorch.
Rongyu Zhang, Zefan Cai, Huanrui Yang, Denis A. Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, Baobao Chang, Yuan Du, Shanghang Zhang
ACM Multimedia9
2024 DialogVCS: Robust Natural Language Understanding in Dialogue System Upgrade
abstract
Zefan Cai, Xin Zheng, Tianyu Liu, Haoran Meng, Jiaqi Han, Gang Yuan, Binghuai Lin, Baobao Chang, Yunbo Cao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zefan Cai, Tianyu Liu 0001, Haoran Meng, Binghuai Lin, Baobao Chang, Yunbo Cao
NAACL-HLT8
2024 Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation
abstract
Haozhe Zhao, Zefan Cai, Shuzheng Si, Liang Chen, Yufeng He, Kaikai An, Baobao Chang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Haozhe Zhao, Zefan Cai, Shuzheng Si, Liang Chen 0024, Yufeng He, Kaikai An, Baobao Chang
NAACL-HLT7
2024 Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
abstract
Fine-tuning is a crucial process for adapting large language models (LLMs) to diverse applications. In certain scenarios, such as multi-tenant serving, deploying multiple LLMs becomes necessary to meet complex demands. Recent studies suggest decomposing a fine-tuned LLM into a base model and corresponding delta weights, which are then compressed using low-rank or low-bit approaches to reduce costs. In this work, we observe that existing low-rank and low-bit compression methods can significantly harm the model performance for task-specific fine-tuned LLMs (e.g., WizardMath for math problems). Motivated by the long-tail distribution of singular values in the delta weights, we propose a delta quantization approach using mixed-precision. This method employs higher-bit representation for singular vectors corresponding to larger singular values. We evaluate our approach on various fine-tuned LLMs, including math LLMs, code LLMs, chat LLMs, and even VLMs. Experimental results demonstrate that our approach performs comparably to full fine-tuned LLMs, surpassing both low-rank and low-bit baselines by a considerable margin. Additionally, we show that our method is compatible with various backbone LLMs, such as Llama-2, Llama-3, and Mistral, highlighting its generalizability.
Bowen Ping, Shuo Wang 0013, Hanqing Wang 0003, Xu Han 0007, Yuzhuang Xu, Yukun Yan, Yun Chen 0007, Baobao Chang, Zhiyuan Liu 0001, Maosong Sun 0001
NeurIPS8
2024 UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
abstract
This paper presents UltraEdit, a large-scale (~ 4M editing samples), automatically generated dataset for instruction-based image editing. Our key idea is to address the drawbacks in existing image editing datasets like InstructPix2Pix and MagicBrush, and provide a systematic approach to producing massive and high-quality image editing samples: 1) UltraEdit includes more diverse editing instructions by combining LLM creativity and in-context editing examples by human raters; 2) UltraEdit is anchored on real images (photographs or artworks), which offers more diversity and less biases than those purely synthesized by text-to-image models; 3) UltraEdit supports region-based editing with high-quality, automatically produced region annotations. Our experiments show that canonical diffusion-based editing baselines trained on UltraEdit set new records on challenging MagicBrush and Emu-Edit benchmarks, respectively. Our analysis further confirms the crucial role of real image anchors and region-based editing data. The dataset, code, and models will be made public.
Haozhe Zhao, Xiaojian Ma 0001, Liang Chen 0024, Shuzheng Si, Rujie Wu, Kaikai An, Peiyu Yu, Minjia Zhang, Qing Li 0003, Baobao Chang
NeurIPS10
2023 Query Your Model with Definitions in FrameNet: An Effective Method for Frame Semantic Role Labeling
abstract
Frame Semantic Role Labeling (FSRL) identifies arguments and labels them with frame semantic roles defined in FrameNet. Previous researches tend to divide FSRL into argument identification and role classification. Such methods usually model role classification as naive multi-class classification and treat arguments individually, which neglects label semantics and interactions between arguments and thus hindering performance and generalization of models. In this paper, we propose a query-based framework named ArGument Extractor with Definitions in FrameNet (AGED) to mitigate these problems. Definitions of frames and frame elements (FEs) in FrameNet can be used to query arguments in text. Encoding text-definition pairs can guide models in learning label semantics and strengthening argument interactions. Experiments show that AGED outperforms previous state-of-the-art by up to 1.3 F1-score in two FrameNet datasets and the generalization power of AGED in zero-shot and fewshot scenarios. Our code and technical appendix is available at https://github.com/PKUnlp-icler/AGED.
Baobao Chang
AAAI3
2023 Can We Edit Factual Knowledge by In-Context Learning?
abstract
Previous studies have shown that large language models (LLMs) like GPTs store massive factual knowledge in their parameters.However, the stored knowledge could be false or outdated.Traditional knowledge editing methods refine LLMs via fine-tuning on texts containing specific knowledge.However, with the increasing scales of LLMs, these gradient-based approaches bring large computation costs.The trend of model-as-a-service also makes it impossible to modify knowledge in black-box LLMs.Inspired by in-context learning (ICL), a new paradigm based on demonstration contexts without parameter updating, we explore whether ICL can edit factual knowledge.To answer this question, we give a comprehensive empirical study of ICL strategies.Experiments show that in-context knowledge editing (IKE), without any gradient and parameter updating, achieves a competitive success rate compared to gradient-based methods on GPT-J (6B) but with much fewer side effects, including less over-editing on similar but unrelated facts and less knowledge forgetting on previously stored knowledge.We also apply the method to larger LMs with tens or hundreds of parameters like OPT-175B, which shows the scalability of our method.The code is available at https://github.com/pkunlp-icler/IKE.
Lei Li 0039, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu 0011, Jingjing Xu 0001, Baobao Chang
EMNLP7
2023 TABLEIE: Capturing the Interactions Among Sub-Tasks in Information Extraction via Double Tables
abstract
Information Extraction mainly consists of three sub-tasks, Named Entity Recognition, Relation Extraction and Event Extraction. Although these sub-tasks are highly correlated with each other, most previous works simply focus on part of them and ignore the interactions among different sub-tasks. Recently, some graph-based models are proposed to cover all the interactions among different IE sub-tasks. However, the use of Graph Neural Network brings heavy computation burden, damaging the model efficiency. In this paper, we propose a double-table framework, TableIE, to capture the interactions among IE sub-tasks as well as improve the model efficiency. Specifically, TableIE has an entity-relation table and an event table, based on which we propose both within-table and cross-table interaction through a novel table integration technique. Such technique makes use of an information-aware mask to extract more essential information in the table during the integration, which we call discriminative interaction. Our extensive experiments demonstrate that TableIE outperforms the previous state-of-the-art up to 1.4 on the ACE05 dataset. Besides, since TableIE does not involve the time-consuming graph operation, it is also more efficient than the previous graph-based models, with 13x speed-up in the inference stage. Our code is available at https://github.com/PKUnlp-icler/TableIE
Jiaxing Lin, Runxin Xu, Baobao Chang
ICASSP3
2023 On the Pareto Front of Multilingual Neural Machine Translation
abstract
In this work, we study how the performance of a given direction changes with its sampling ratio in Multilingual Neural Machine Translation (MNMT). By training over 200 multilingual models with various model sizes, data sizes, and language directions, we find it interesting that the performance of certain translation direction does not always improve with the increase of its weight in the multi-task optimization objective. Accordingly, scalarization method leads to a multitask trade-off front that deviates from the traditional Pareto front when there exists data imbalance in the training corpus, which poses a great challenge to improve the overall performance of all directions. Based on our observations, we propose the Double Power Law to predict the unique performance trade-off front in MNMT, which is robust across various languages, data adequacy, and the number of tasks. Finally, we formulate the sample ratio selection problem in MNMT as an optimization problem based on the Double Power Law. Extensive experiments show that it achieves better performance than temperature searching and gradient manipulation methods with only 1/5 to 1/2 of the total training budget. We release the code at https://github.com/pkunlp-icler/ParetoMNMT for reproduction.
Liang Chen 0024, Shuming Ma, Dongdong Zhang 0001, Furu Wei, Baobao Chang
NeurIPS5
2023 Mixture-of-Experts for Biomedical Question Answering
Damai Dai, Wenbin Jiang 0002, Yajuan Lyu, Zhifang Sui, Baobao Chang
NLPCC (1)6
2023 Coarse-to-Fine Entity Representations for Document-Level Relation Extraction
Damai Dai, Shuang Zeng, Baobao Chang, Zhifang Sui
NLPCC (2)4
2023 A Span-based Target-aware Relation Model for Frame-semantic Parsing
abstract
Frame-semantic Parsing (FSP) is a challenging and critical task in Natural Language Processing (NLP). Most of the existing studies decompose the FSP task into frame identification (FI) and frame semantic role labeling (FSRL) subtasks, and adopt a pipeline model architecture that clearly causes error propagation problem. However, recent jointly learning models aim to address the above problem and generally treat FSP as a span-level structured prediction task, which, unfortunately, leads to cascading error propagation problem between roles and less-efficient solutions due to huge search space of roles. To address these problems, we reformulate the FSRL task into a target-aware relation classification task and propose a novel and lightweight jointly learning framework that simultaneously processes three subtasks of FSP, including frame identification, argument identification, and role classification. The novel task formulation and jointly learning with interaction mechanisms among subtasks can help improve the overall system performance and reduce the search space and time complexity, compared with existing methods. Extensive experimental results demonstrate that our proposed model significantly outperforms 10 state-of-the-art models in terms of F1 score across two benchmark datasets.
Xuefeng Su, Ru Li 0001, Xiaoli Li 0001, Baobao Chang, Zhiwei Hu, Xiaoqi Han, Zhichao Yan 0002
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2022 From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model Compression
abstract
Pre-trained Language Models (PLMs) have achieved great success in various Natural Language Processing (NLP) tasks under the pre-training and fine-tuning paradigm. With large quantities of parameters, PLMs are computation-intensive and resource-hungry. Hence, model pruning has been introduced to compress large-scale PLMs. However, most prior approaches only consider task-specific knowledge towards downstream tasks, but ignore the essential task-agnostic knowledge during pruning, which may cause catastrophic forgetting problem and lead to poor generalization ability. To maintain both task-agnostic and task-specific knowledge in our pruned model, we propose ContrAstive Pruning (CAP) under the paradigm of pre-training and fine-tuning. It is designed as a general framework, compatible with both structured and unstructured pruning. Unified in contrastive learn- ing, CAP enables the pruned model to learn from the pre-trained model for task-agnostic knowledge, and fine-tuned model for task-specific knowledge. Besides, to better retain the performance of the pruned model, the snapshots (i.e., the intermediate models at each pruning iteration) also serve as effective supervisions for pruning. Our extensive experiments show that adopting CAP consistently yields significant improvements, especially in extremely high sparsity scenarios. With only 3% model parameters reserved (i.e., 97% sparsity), CAP successfully achieves 99.2% and 96.3% of the original BERT performance in QQP and MNLI tasks. In addition, our probing experiments demonstrate that the model pruned by CAP tends to achieve better generalization ability.
Runxin Xu, Fuli Luo, Chengyu Wang 0001, Baobao Chang, Jun Huang 0007, Songfang Huang, Fei Huang 0002
AAAI4
2022 StableMoE: Stable Routing Strategy for Mixture of Experts
abstract
The Mixture-of-Experts (MoE) technique can scale up the model size of Transformers with an affordable computational overhead.We point out that existing learning-to-route MoE methods suffer from the routing fluctuation issue, i.e., the target expert of the same input may change along with training, but only one expert will be activated for the input during inference.The routing fluctuation tends to harm sample efficiency because the same input updates different experts but only one is finally used.In this paper, we propose STABLEMOE with two training stages to address the routing fluctuation problem.In the first training stage, we learn a balanced and cohesive routing strategy and distill it into a lightweight router decoupled from the backbone model.In the second training stage, we utilize the distilled router to determine the token-to-expert assignment and freeze it for a stable routing strategy.We validate our method on language modeling and multilingual machine translation.The results show that STABLEMOE outperforms existing MoE methods in terms of both convergence speed and performance.
Damai Dai, Li Dong 0004, Shuming Ma, Bo Zheng 0010, Zhifang Sui, Baobao Chang, Furu Wei
ACL (1)6
2022 Knowledge Neurons in Pretrained Transformers
abstract
Large-scale pretrained language models are surprisingly good at recalling factual knowledge presented in the training corpus (Petroni et al., 2019; Jiang et al., 2020b).In this paper, we present preliminary studies on how factual knowledge is stored in pretrained Transformers by introducing the concept of knowledge neurons.Specifically, we examine the fill-in-the-blank cloze task for BERT.Given a relational fact, we propose a knowledge attribution method to identify the neurons that express the fact.We find that the activation of such knowledge neurons is positively correlated to the expression of their corresponding facts.In our case studies, we attempt to leverage knowledge neurons to edit (such as update, and erase) specific factual knowledge without fine-tuning.Our results shed light on understanding the storage of knowledge within pretrained Transformers.The code is available at https://github.com/ Hunter-DDM/knowledge-neurons.
Damai Dai, Li Dong 0004, Yaru Hao, Zhifang Sui, Baobao Chang, Furu Wei
ACL (1)5
2022 Premise-based Multimodal Reasoning: Conditional Inference on Joint Textual and Visual Clues
abstract
Qingxiu Dong, Ziwei Qin, Heming Xia, Tian Feng, Shoujie Tong, Haoran Meng, Lin Xu, Zhongyu Wei, Weidong Zhan, Baobao Chang, Sujian Li, Tianyu Liu, Zhifang Sui. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Qingxiu Dong, Ziwei Qin, Heming Xia, Shoujie Tong, Haoran Meng, Zhongyu Wei, Weidong Zhan, Baobao Chang, Sujian Li, Tianyu Liu 0001, Zhifang Sui
ACL (1)10
2022 CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
abstract
Ningyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li, Xin Shang, Kangping Yin, Chuanqi Tan, Jian Xu, Fei Huang, Luo Si, Yuan Ni, Guotong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan, Linfeng Li, Jun Yan, Hongying Zan, Kunli Zhang, Buzhou Tang, Qingcai Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Ningyu Zhang 0001, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li 0040, Xin Shang, Kangping Yin, Chuanqi Tan, Fei Huang 0002, Luo Si, Yuan Ni, Guo Tong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan 0002, Jun Yan 0010, Hongying Zan, Kunli Zhang, Buzhou Tang, Qingcai Chen
ACL (1)15
2022 DISK: Domain-constrained Instance Sketch for Math Word Problem Generation
abstract
A math word problem (MWP) is a coherent narrative which reflects the underlying logic of math equations. Successful MWP generation can automate the writing of mathematics questions. Previous methods mainly generate MWP text based on inflexible pre-defined templates. In this paper, we propose a neural model for generating MWP text from math equations. Firstly, we incorporate a matching model conditioned on the domain knowledge to retrieve a MWP instance which is most consistent with the ground-truth, where the domain is a latent variable extracted with a domain summarizer. Secondly, by constructing a Quantity Cell Graph (QCG) from the retrieved MWP instance and reasoning over it, we improve the model’s comprehension of real-world scenarios and derive a domain-constrained instance sketch to guide the generation. Besides, the QCG also interacts with the equation encoder to enhance the alignment between math tokens (e.g., quantities and variables) and MWP text. Experiments and empirical analysis on educational MWP set show that our model achieves impressive performance in both automatic evaluation metrics and human evaluation metrics.
Tianyang Cao, Shuang Zeng, Xiaodan Xu, Mairgup Mansur, Baobao Chang
COLING5
2022 SCL-RAI: Span-based Contrastive Learning with Retrieval Augmented Inference for Unlabeled Entity Problem in NER
abstract
Unlabeled Entity Problem (UEP) in Named Entity Recognition (NER) datasets seriously hinders the improvement of NER performance. This paper proposes SCL-RAI to cope with this problem. Firstly, we decrease the distance of span representations with the same label while increasing it for different ones via span-based contrastive learning, which relieves the ambiguity among entities and improves the robustness of the model over unlabeled entities. Then we propose retrieval augmented inference to mitigate the decision boundary shifting problem. Our method significantly outperforms the previous SOTA method by 4.21% and 8.64% F1-score on two real-world datasets.
Shuzheng Si, Shuang Zeng, Jiaxing Lin, Baobao Chang
COLING4
2022 Robust Fine-tuning via Perturbation and Interpolation from In-batch Instances
abstract
Fine-tuning pretrained language models (PLMs) on downstream tasks has become common practice in natural language processing. However, most of the PLMs are vulnerable, e.g., they are brittle under adversarial attacks or imbalanced data, which hinders the application of the PLMs on some downstream tasks, especially in safe-critical scenarios. In this paper, we propose a simple yet effective fine-tuning method called Match-Tuning to force the PLMs to be more robust. For each instance in a batch, we involve other instances in the same batch to interact with it. To be specific, regarding the instances with other labels as a perturbation, Match-Tuning makes the model more robust to noise at the beginning of training. While nearing the end, Match-Tuning focuses more on performing an interpolation among the instances with the same label for better generalization. Extensive experiments on various tasks in GLUE benchmark show that Match-Tuning consistently outperforms the vanilla fine-tuning by 1.64 scores. Moreover, Match-Tuning exhibits remarkable robustness to adversarial attacks and data imbalance.
Shoujie Tong, Qingxiu Dong, Damai Dai, Yifan Song 0002, Tianyu Liu 0001, Baobao Chang, Zhifang Sui
IJCAI6
2022 CLINER: Clinical Interrogation Named Entity Recognition
Tianyang Cao, Yifan Yang 0008, Yunyan Zhang, Xi Chen 0003, Baobao Chang, Zhifang Sui, Ruihui Zhao, Yefeng Zheng 0001, Bang Liu 0003
KSEM (2)7
2022 Mining Clues from Incomplete Utterance: A Query-enhanced Network for Incomplete Utterance Rewriting
abstract
Incomplete utterance rewriting has recently raised wide attention.However, previous works do not consider the semantic structural information between incomplete utterance and rewritten utterance or model the semantic structure implicitly and insufficiently.To address this problem, we propose a QUEry-Enhanced Network (QUEEN).Firstly, our proposed query template explicitly brings guided semantic structural knowledge between the incomplete utterance and the rewritten utterance making model perceive where to refer back to or recover omitted tokens.Then, we adopt a fast and effective edit operation scoring network to model the relation between two tokens.Benefiting from extra information and the well-designed network, QUEEN achieves state-of-the-art performance on several public datasets.
Shuzheng Si, Shuang Zeng, Baobao Chang
NAACL-HLT3
2022 An Enhanced Span-based Decomposition Method for Few-Shot Sequence Labeling
abstract
Peiyi Wang, Runxin Xu, Tianyu Liu, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui
NAACL-HLT6
2022 A Two-Stream AMR-enhanced Model for Document-level Event Argument Extraction
abstract
Runxin Xu, Peiyi Wang, Tianyu Liu, Shuang Zeng, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Runxin Xu, Peiyi Wang, Tianyu Liu 0001, Shuang Zeng, Baobao Chang, Zhifang Sui
NAACL-HLT5
2022 A Double-Graph Based Framework for Frame Semantic Parsing
abstract
Frame semantic parsing is a fundamental NLP task, which consists of three subtasks: frame identification, argument identification and role classification.Most previous studies tend to neglect relations between different subtasks and arguments and pay little attention to ontological frame knowledge defined in FrameNet.In this paper, we propose a Knowledge-guided Incremental semantic parser with Double-graph (KID).We first introduce Frame Knowledge Graph (FKG), a heterogeneous graph containing both frames and FEs (Frame Elements) built on the frame knowledge so that we can derive knowledgeenhanced representations for frames and FEs.Besides, we propose Frame Semantic Graph (FSG) to represent frame semantic structures extracted from the text with graph structures.In this way, we can transform frame semantic parsing into an incremental graph construction problem to strengthen interactions between subtasks and relations between arguments.Our experiments show that KID outperforms the previous state-of-the-art method by up to 1.7 F1-score on two FrameNet datasets.Our code is availavle at https://github. com/PKUnlp-icler/KID.
Runxin Xu, Baobao Chang
NAACL-HLT4
2022 Plug-and-Play Module for Commonsense Reasoning in Machine Reading Comprehension
Damai Dai, Zhifang Sui, Baobao Chang
NLPCC (2)4
2021 Towards Faithfulness in Open Domain Table-to-text Generation from an Entity-centric View
abstract
In open domain table-to-text generation, we notice the unfaithful generation usually contains hallucinated entities which can not be aligned to any input table record. We thus try to evaluate the generation faithfulness with two entity-centric metrics: table record coverage and the ratio of hallucinated entities in text, both of which are shown to have strong agreement with human judgements. Then based on these metrics, we quantitatively analyze the correlation between training data quality and generation fidelity which indicates the potential usage of entity information in faithful generation. Motivated by these findings, we propose two methods for faithful generation: 1) augmented training by incorporating the auxiliary entity information, including both an augmented plan-based model and an unsupervised model and 2) training instance selection based on faithfulness ranking. We show these approaches improve generation fidelity in both full dataset setting and few shot setting by both automatic and human evaluations.
Tianyu Liu 0001, Baobao Chang, Zhifang Sui
AAAI3
2021 Document-level Event Extraction via Heterogeneous Graph-based Interaction Model with a Tracker
abstract
Runxin Xu, Tianyu Liu, Lei Li, Baobao Chang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Runxin Xu, Tianyu Liu 0001, Lei Li 0005, Baobao Chang
ACL/IJCNLP (1)4
2021 Behind the Scenes: An Exploration of Trigger Biases Problem in Few-Shot Event Classification
abstract
Few-Shot Event Classification (FSEC) aims at developing a model for event prediction, which can generalize to new event types with a limited number of annotated data. Existing FSEC studies have achieved high accuracy on different benchmarks. However, we find they suffer from trigger biases that signify the statistical homogeneity between some trigger words and target event types, which we summarize as trigger overlapping and trigger separability. The biases can result in context-bypassing problem, i.e., correct classifications can be gained by looking at only the trigger words while ignoring the entire context. Therefore, existing models can be weak in generalizing to unseen data in real scenarios. To further uncover the trigger biases and assess the generalization ability of the models, we propose two new sampling methods, Trigger-Uniform Sampling (TUS) and COnfusion Sampling (COS), for the meta tasks construction during evaluation. Besides, to cope with the context-bypassing problem in FSEC models, we introduce adversarial training and trigger reconstruction techniques. Experiments show these techniques help not only improve the performance, but also enhance the generalization ability of models.
Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Damai Dai, Baobao Chang, Zhifang Sui
CIKM5
2021 Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning
abstract
Recent pretrained language models extend from millions to billions of parameters.Thus the need to fine-tune an extremely large pretrained model with a limited training corpus arises in various downstream tasks.In this paper, we propose a straightforward yet effective fine-tuning technique, CHILD-TUNING, which updates a subset of parameters (called child network) of large pretrained models via strategically masking out the gradients of the non-child network during the backward process.Experiments on various downstream tasks in GLUE benchmark show that CHILD-TUNING consistently outperforms the vanilla fine-tuning by 1.5 ∼ 8.6 average score among four different pretrained models, and surpasses the prior fine-tuning techniques by 0.6 ∼ 1.3 points.Furthermore, empirical results on domain transfer and task transfer show that CHILD-TUNING can obtain better generalization performance by large margins.
Runxin Xu, Fuli Luo, Chuanqi Tan, Baobao Chang, Songfang Huang, Fei Huang 0002
EMNLP (1)5
2021 Generating Math Word Problems from Equations with Topic Consistency Maintaining and Commonsense Enforcement
Tianyang Cao, Shuang Zeng, Songge Zhao, Mairgup Mansur, Baobao Chang
ICANN (3)5
2021 Decompose, Fuse and Generate: A Formation-Informed Method for Chinese Definition Generation
abstract
Hua Zheng, Damai Dai, Lei Li, Tianyu Liu, Zhifang Sui, Baobao Chang, Yang Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Damai Dai, Lei Li 0039, Tianyu Liu 0001, Zhifang Sui, Baobao Chang, Yang Liu 0124
NAACL-HLT6
2020 Interactive Neural Network: Leveraging Part-of-Speech Window for Aspect Term Extraction (Student Abstract)
abstract
Aspect term extraction is a fundamental task for aspect-level sentiment analysis. Previous methods tend to extract noun aspect terms due to the large quantities of them, and perform badly on extracting aspect terms containing words with other POS tags, according to experimental results. In addition, few works focus on the POS tags of adjacent words which are critical to aspect term extraction. We propose a novel model which combines POS and word features in an interactive way, and makes full use of the POS tags of adjacent words by POS window. We conduct experiments on two datasets, and prove the effectiveness of our model.
Da Yin, Xiuyu Wu, Baobao Chang
AAAI3
2020 Soap: Soaking Capacity Optimization for Multi-Document Summarization
abstract
Multi-document summarization (MDS) aims at giving a brief summary for a cluster of related documents. In this paper, we consider the MDS task as an optimization problem with a novel measure named soaking capacity being the objective function. The origin of our method is the classic hypothesis: the summary components are the sinks of information diffusion. We point out that the hypothesis only gives the role of summary but does not cover how well a summary acts as this role. To fill in the gap, soaking capacity is formally defined to quantify the ability of summary to soak up information. We explicitly demonstrate its fitness as an indicator for both the saliency and the diversity goal of MDS. For solving the optimization problem, we propose a greedy algorithm named Soap by adopting a surrogate of soaking capacity to accelerate the computation. Experiments on MDS datasets across various domains show the great potential of Soap as compared with the state-of-the-art MDS systems.
Kexiang Wang, Baobao Chang, Zhifang Sui
CIKM2
2020 An Anchor-Based Automatic Evaluation Metric for Document Summarization
abstract
The widespread adoption of reference-based automatic evaluation metrics such as ROUGE has promoted the development of document summarization.In this paper, we consider a new protocol for designing reference-based metrics that require the endorsement of source document(s).Following protocol, we propose an anchored ROUGE metric fixing each summary particle on source document, which bases the computation on more solid ground.Empirical results on benchmark datasets validate that source document helps to induce a higher correlation with human judgments for ROUGE metric.Being self-explanatory and easy-to-implement, the protocol can naturally foster various effective designs of reference-based metrics besides the anchored ROUGE introduced here.
Kexiang Wang, Tianyu Liu 0001, Baobao Chang, Zhifang Sui
COLING3
2020 An Empirical Study on Model-agnostic Debiasing Strategies for Robust Natural Language Inference
abstract
The prior work on natural language inference (NLI) debiasing mainly targets at one or few known biases while not necessarily making the models more robust.In this paper, we focus on the model-agnostic debiasing strategies and explore how to (or is it possible to) make the NLI models robust to multiple distinct adversarial attacks while keeping or even strengthening the models' generalization power.We firstly benchmark prevailing neural NLI models including pretrained ones on various adversarial datasets.We then try to combat distinct known biases by modifying a mixture of experts (MoE) ensemble method (Clark et al., 2019) and show that it's nontrivial to mitigate multiple NLI biases at the same time, and that model-level ensemble method outperforms MoE ensemble method.We also perform data augmentation including text swap, word substitution and paraphrase and prove its efficiency in combating various (though not all) adversarial attacks at the same time.Finally, we investigate several methods to merge heterogeneous training data (1.35M) and perform model ensembling, which are straightforward but effective to strengthen NLI models.
Tianyu Liu 0001, Xiaoan Ding, Baobao Chang, Zhifang Sui
CoNLL4
2020 Discriminatively-Tuned Generative Classifiers for Robust Natural Language Inference
abstract
While discriminative neural network classifiers are generally preferred, recent work has shown advantages of generative classifiers in term of data efficiency and robustness.In this paper, we focus on natural language inference (NLI).We propose GenNLI, a generative classifier for NLI tasks, and empirically characterize its performance by comparing it to five baselines, including discriminative models and large-scale pretrained language representation models like BERT.We explore training objectives for discriminative fine-tuning of our generative classifiers, showing improvements over log loss fine-tuning from prior work (Lewis and Fan, 2019).In particular, we find strong results with a simple unbounded modification to log loss, which we call the "infinilog loss".Our experiments show that GenNLI outperforms both discriminative and pretrained baselines across several challenging NLI experimental settings, including small training sets, imbalanced label distributions, and label noise.
Xiaoan Ding, Tianyu Liu 0001, Baobao Chang, Zhifang Sui, Kevin Gimpel
EMNLP (1)3
2020 A Spectral Method for Unsupervised Multi-Document Summarization
abstract
Multi-document summarization (MDS) aims at producing a good-quality summary for several related documents.In this paper, we propose a spectral-based hypothesis, which states that the goodness of summary candidate is closely linked to its so-called spectral impact.Here spectral impact considers the perturbation to the dominant eigenvalue of affinity matrix when dropping the summary candidate from the document cluster.The hypothesis is validated by three theoretical perspectives: semantic scaling, propagation dynamics and matrix perturbation.According to the hypothesis, we formulate the MDS task as the combinatorial optimization of spectral impact and propose an accelerated greedy solution based on a surrogate of spectral impact.The evaluation results on various datasets demonstrate:(1) The performance of the summary candidate is positively correlated with its spectral impact, which accords with our hypothesis; (2) Our spectral-based method has a competitive result as compared to state-of-the-art MDS systems.
Kexiang Wang, Baobao Chang, Zhifang Sui
EMNLP (1)2
2020 Double Graph Based Reasoning for Document-level Relation Extraction
abstract
Document-level relation extraction aims to extract relations among entities within a document.Different from sentence-level relation extraction, it requires reasoning over multiple sentences across paragraphs.In this paper, we propose Graph Aggregation-and-Inference Network (GAIN), a method to recognize such relations for long paragraphs.GAIN constructs two graphs, a heterogeneous mentionlevel graph (MG) and an entity-level graph (EG).The former captures complex interaction among different mentions and the latter aggregates mentions underlying for the same entities.Based on the graphs we propose a novel path reasoning mechanism to infer relations between entities.Experiments on the public dataset, DocRED, show GAIN achieves a significant performance improvement (2.85 on F1) over the previous state-of-the-art.Our code is available at https://github.com/ PKUnlp-icler/GAIN.
Shuang Zeng, Runxin Xu, Baobao Chang, Lei Li 0005
EMNLP (1)3
2020 HypoNLI: Exploring the Artificial Patterns of Hypothesis-only Bias in Natural Language Inference
abstract
Many recent studies have shown that for models trained on datasets for natural language inference (NLI), it is possible to make correct predictions by merely looking at the hypothesis while completely ignoring the premise. In this work, we manage to derive adversarial examples in terms of the hypothesis-only bias and explore eligible ways to mitigate such bias. Specifically, we extract various phrases from the hypotheses (artificial patterns) in the training sets, and show that they have been strong indicators to the specific labels. We then figure out ‘hard’ and ‘easy’ instances from the original test sets whose labels are opposite to or consistent with those indications. We also set up baselines including both pretrained models (BERT, RoBerta, XLNet) and competitive non-pretrained models (InferSent, DAM, ESIM). Apart from the benchmark and baselines, we also investigate two debiasing approaches which exploit the artificial pattern modeling to mitigate such hypothesis-only bias: down-sampling and adversarial training. We believe those methods can be treated as competitive baselines in NLI debiasing tasks.
Tianyu Liu 0001, Baobao Chang, Zhifang Sui
LREC3
2019 Hierarchical Encoder with Auxiliary Supervision for Neural Table-to-Text Generation: Learning Better Representation for Tables
abstract
Generating natural language descriptions for the structured tables which consist of multiple attribute-value tuples is a convenient way to help people to understand the tables. Most neural table-to-text models are based on the encoder-decoder framework. However, it is hard for a vanilla encoder to learn the accurate semantic representation of a complex table. The challenges are two-fold: firstly, the table-to-text datasets often contain large number of attributes across different domains, thus it is hard for the encoder to incorporate these heterogeneous resources. Secondly, the single encoder also has difficulties in modeling the complex attribute-value structure of the tables. To this end, we first propose a two-level hierarchical encoder with coarse-to-fine attention to handle the attribute-value structure of the tables. Furthermore, to capture the accurate semantic representations of the tables, we propose 3 joint tasks apart from the prime encoder-decoder learning, namely auxiliary sequence labeling task, text autoencoder and multi-labeling classification, as the auxiliary supervisions for the table encoder. We test our models on the widely used dataset WIKIBIO which contains Wikipedia infoboxes and related descriptions. The dataset contains complex tables as well as large number of attributes across different domains. We achieve the state-of-the-art performance on both automatic and human evaluation metrics.
Tianyu Liu 0001, Fuli Luo, Qiaolin Xia, Shuming Ma, Baobao Chang, Zhifang Sui
AAAI5
2019 Towards Comprehensive Description Generation from Factual Attribute-value Tables
abstract
The comprehensive descriptions for factual attribute-value tables, which should be accurate, informative and loyal, can be very helpful for end users to understand the structured data in this form.However previous neural generators might suffer from key attributes missing, less informative and groundless information problems, which impede the generation of high-quality comprehensive descriptions for tables.To relieve these problems, we first propose force attention (FA) method to encourage the generator to pay more attention to the uncovered attributes to avoid potential key attributes missing.Furthermore, we propose reinforcement learning for information richness to generate more informative as well as more loyal descriptions for tables.In our experiments, we utilize the widely used WIKIBIO dataset as a benchmark.Additionally we create WB-filter based on WIKIBIO to test our model in the simulated user-oriented scenarios, in which the generated descriptions should accord with particular user interests.Experimental results show that our model outperforms the state-of-the-art baselines on both automatic and human evaluation.
Tianyu Liu 0001, Fuli Luo, Wei Wu 0044, Baobao Chang, Zhifang Sui
ACL (1)5
2019 Learning to Control the Fine-grained Sentiment for Story Ending Generation
abstract
Automatic story ending generation is an interesting and challenging task in natural language generation.Previous studies are mainly limited to generate coherent, reasonable and diversified story endings, and few works focus on controlling the sentiment of story endings.This paper focuses on generating a story ending which meets the given fine-grained sentiment intensity.There are two major challenges to this task.First is the lack of story corpus which has fine-grained sentiment labels.Second is the difficulty of explicitly controlling sentiment intensity when generating endings.Therefore, we propose a generic and novel framework which consists of a sentiment analyzer and a sentimental generator, respectively addressing the two challenges.The sentiment analyzer adopts a series of methods to acquire sentiment intensities of the story dataset.The sentimental generator introduces the sentiment intensity into decoder via a Gaussian Kernel Layer to control the sentiment of the output.To the best of our knowledge, this is the first endeavor to control the fine-grained sentiment for story ending generation without manually annotating sentiment labels.Experiments show that our proposed framework can generate story endings which are not only more coherent and fluent but also able to meet the given sentiment intensity better. 1
Fuli Luo, Damai Dai, Tianyu Liu 0001, Baobao Chang, Zhifang Sui, Xu Sun 0001
ACL (1)5
2019 Towards Fine-grained Text Sentiment Transfer
abstract
In this paper, we focus on the task of finegrained text sentiment transfer (FGST).This task aims to revise an input sequence to satisfy a given sentiment intensity, while preserving the original semantic content.Different from conventional sentiment transfer task that only reverses the sentiment polarity (positive/negative) of text, the FTST task requires more nuanced and fine-grained control of sentiment.To remedy this, we propose a novel Seq2SentiSeq model.Specifically, the numeric sentiment intensity value is incorporated into the decoder via a Gaussian kernel layer to finely control the sentiment intensity of the output.Moreover, to tackle the problem of lacking parallel data, we propose a cycle reinforcement learning algorithm to guide the model training.In this framework, the elaborately designed rewards can balance both sentiment transformation and content preservation, while not requiring any ground truth output.Experimental results show that our approach can outperform existing methods by a large margin in both automatic evaluation and human evaluation.Our code and data, including outputs of all baselines and our model are available at https://github.com/luofuli/ Fine-grained-Sentiment-Transfer. 1
Fuli Luo, Peng Li 0030, Jie Zhou 0016, Yutong Tan, Baobao Chang, Zhifang Sui, Xu Sun 0001
ACL (1)6
2019 Pun-GAN: Generative Adversarial Network for Pun Generation
abstract
Fuli Luo, Shunyao Li, Pengcheng Yang, Lei Li, Baobao Chang, Zhifang Sui, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Fuli Luo, Shunyao Li, Lei Li 0039, Baobao Chang, Zhifang Sui, Xu Sun 0001
EMNLP/IJCNLP (1)5
2019 A Dual Reinforcement Learning Framework for Unsupervised Text Style Transfer
abstract
Unsupervised text style transfer aims to transfer the underlying style of text but keep its main content unchanged without parallel data. Most existing methods typically follow two steps: first separating the content from the original style, and then fusing the content with the desired style. However, the separation in the first step is challenging because the content and style interact in subtle ways in natural language. Therefore, in this paper, we propose a dual reinforcement learning framework to directly transfer the style of the text via a one-step mapping model, without any separation of content and style. Specifically, we consider the learning of the source-to-target and target-to-source mappings as a dual task, and two rewards are designed based on such a dual structure to reflect the style accuracy and content preservation, respectively. In this way, the two one-step mapping models can be trained via reinforcement learning, without any use of parallel data. Automatic evaluations show that our model outperforms the state-of-the-art systems by a large margin, especially with more than 10 BLEU points improvement averaged on two benchmark datasets. Human evaluations also validate the effectiveness of our model in terms of style accuracy, content preservation and fluency. Our code and data, including outputs of all baselines and our model are available at https://github.com/luofuli/DualRL.
Fuli Luo, Peng Li 0030, Jie Zhou 0016, Baobao Chang, Xu Sun 0001, Zhifang Sui
IJCAI5
2018 Table-to-Text Generation by Structure-Aware Seq2seq Learning
abstract
Table-to-text generation aims to generate a description for a factual table which can be viewed as a set of field-value records. To encode both the content and the structure of a table, we propose a novel structure-aware seq2seq architecture which consists of field-gating encoder and description generator with dual attention. In the encoding phase, we update the cell memory of the LSTM unit by a field gate and its corresponding field value in order to incorporate field information into table representation. In the decoding phase, dual attention mechanism which contains word level attention and field level attention is proposed to model the semantic relevance between the generated description and the table. We conduct experiments on the WIKIBIO dataset which contains over 700k biographies and corresponding infoboxes from Wikipedia. The attention visualizations and case studies show that our model is capable of generating coherent and informative descriptions based on the comprehensive understanding of both the content and the structure of a table. Automatic evaluations also show our model outperforms the baselines by a great margin. Code for this work is available on https://github.com/tyliupku/wiki2bio.
Tianyu Liu 0001, Kexiang Wang, Lei Sha, Baobao Chang, Zhifang Sui
AAAI4
2018 Order-Planning Neural Text Generation From Structured Data
abstract
Generating texts from structured data (e.g., a table) is important for various natural language processing tasks such as question answering and dialog systems. In recent studies, researchers use neural language models and encoder-decoder frameworks for table-to-text generation. However, these neural network-based approaches typically do not model the order of content during text generation. When a human writes a summary based on a given table, he or she would probably consider the content order before wording. In this paper, we propose an order-planning text generation model, where order information is explicitly captured by link-based attention. Then a self-adaptive gate combines the link-based attention with traditional content-based attention. We conducted experiments on the WikiBio dataset and achieve higher performance than previous methods in terms of BLEU, ROUGE, and NIST scores; we also performed ablation tests to analyze each component of our model.
Lei Sha, Lili Mou, Tianyu Liu 0001, Pascal Poupart, Sujian Li, Baobao Chang, Zhifang Sui
AAAI6
2018 Jointly Extracting Event Triggers and Arguments by Dependency-Bridge RNN and Tensor-Based Argument Interaction
abstract
Event extraction plays an important role in natural language processing (NLP) applications including question answering and information retrieval. Traditional event extraction relies heavily on lexical and syntactic features, which require intensive human engineering and may not generalize to different datasets. Deep neural networks, on the other hand, are able to automatically learn underlying features, but existing networks do not make full use of syntactic relations. In this paper, we propose a novel dependency bridge recurrent neural network (dbRNN) for event extraction. We build our model upon a recurrent neural network, but enhance it with dependency bridges, which carry syntactically related information when modeling each word.We illustrates that simultaneously applying tree structure and sequence structure in RNN brings much better performance than only uses sequential RNN. In addition, we use a tensor layer to simultaneously capture the various types of latent interaction between candidate arguments as well as identify/classify all arguments of an event. Experiments show that our approach achieves competitive results compared with previous work.
Lei Sha, Baobao Chang, Zhifang Sui
AAAI3
2018 A Multi-View Fusion Neural Network for Answer Selection
abstract
Community question answering aims at choosing the most appropriate answer for a given question, which is important in many NLP applications. Previous neural network-based methods consider several different aspects of information through calculating attentions. These different kinds of attentions are always simply summed up and can be seen as a ``single view", causing severe information loss. To overcome this problem, we propose a Multi-View Fusion Neural Network, where each attention component generates a ``view'' of the QA pair and a fusion RNN integrates the generated views to form a more holistic representation. In this fusion RNN method, a filter gate collects important information of input and directly adds it to the output, which borrows the idea of residual networks. Experimental results on the WikiQA and SemEval-2016 CQA datasets demonstrate that our proposed model outperforms the state-of-the-art methods.
Lei Sha, Xiaodong Zhang 0022, Baobao Chang, Zhifang Sui
AAAI4
2018 Incorporating Glosses into Neural Word Sense Disambiguation
abstract
Word Sense Disambiguation (WSD) aims to identify the correct meaning of polysemous words in the particular context.Lexical resources like WordNet which are proved to be of great help for WSD in the knowledge-based methods.However, previous neural networks for WSD always rely on massive labeled data (context), ignoring lexical resources like glosses (sense definitions).In this paper, we integrate the context and glosses of the target word into a unified framework in order to make full use of both labeled data and lexical knowledge.Therefore, we propose GAS: a gloss-augmented WSD neural network which jointly encodes the context and glosses of the target word.GAS models the semantic relationship between the context and the gloss in an improved memory network framework, which breaks the barriers of the previous supervised methods and knowledge-based methods.We further extend the original gloss of word sense via its semantic relations in WordNet to enrich the gloss information.The experimental results show that our model outperforms the state-of-theart systems on several English all-words WSD datasets.
Fuli Luo, Tianyu Liu 0001, Qiaolin Xia, Baobao Chang, Zhifang Sui
ACL (1)4
2018 Fine-grained Coordinated Cross-lingual Text Stream Alignment for Endless Language Knowledge Acquisition
abstract
This paper proposes to study fine-grained coordinated cross-lingual text stream alignment through a novel information network decipherment paradigm.We use Burst Information Networks as media to represent text streams and present a simple yet effective network decipherment algorithm with diverse clues to decipher the networks for accurate text stream alignment.Experiments on Chinese-English news streams show our approach not only outperforms previous approaches on bilingual lexicon extraction from coordinated text streams but also can harvest high-quality alignments from large amounts of streaming data for endless language knowledge mining, which makes it promising to be a new paradigm for automatic language knowledge acquisition.
Tao Ge 0001, Qing Dou, Heng Ji 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001
EMNLP5
2018 Leveraging Gloss Knowledge in Neural Word Sense Disambiguation by Hierarchical Co-Attention
abstract
The goal of Word Sense Disambiguation (WSD) is to identify the correct meaning of a word in the particular context.Traditional supervised methods only use labeled data (context), while missing rich lexical knowledge such as the gloss which defines the meaning of a word sense.Recent studies have shown that incorporating glosses into neural networks for WSD has made significant improvement.However, the previous models usually build the context representation and gloss representation separately.In this paper, we find that the learning for the context and gloss representation can benefit from each other.Gloss can help to highlight the important words in the context, thus building a better context representation.Context can also help to locate the key words in the gloss of the correct word sense.Therefore, we introduce a co-attention mechanism to generate co-dependent representations for the context and gloss.Furthermore, in order to capture both word-level and sentence-level information, we extend the attention mechanism in a hierarchical fashion.Experimental results show that our model achieves the state-of-the-art results on several standard English all-words WSD test datasets.
Fuli Luo, Tianyu Liu 0001, Zexue He, Qiaolin Xia, Zhifang Sui, Baobao Chang
EMNLP6
2018 Improved Dependency Parsing using Implicit Word Connections Learned from Unlabeled Data
abstract
Pre-trained word embeddings and language model have been shown useful in a lot of tasks.However, both of them cannot directly capture word connections in a sentence, which is important for dependency parsing given its goal is to establish dependency relations between words.In this paper, we propose to implicitly capture word connections from unlabeled data by a word ordering model with selfattention mechanism.Experiments show that these implicit word connections do improve our parsing model.Furthermore, by combining with a pre-trained language model, our model gets state-of-the-art performance on the English PTB dataset, achieving 96.35% UAS and 95.25% LAS.
Wenhui Wang 0003, Baobao Chang, Mairgup Mansur
EMNLP2
2018 EventWiki: A Knowledge Base of Major Events
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001
LREC3
2018 SeRI: A Dataset for Sub-event Relation Inference from an Encyclopedia
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001
NLPCC (2)3
2017 Gated Self-Matching Networks for Reading Comprehension and Question Answering
abstract
In this paper, we present the gated selfmatching networks for reading comprehension style question answering, which aims to answer questions from a given passage.We first match the question and passage with gated attention-based recurrent networks to obtain the question-aware passage representation.Then we propose a self-matching attention mechanism to refine the representation by matching the passage against itself, which effectively encodes information from the whole passage.We finally employ the pointer networks to locate the positions of answers from the passages.We conduct extensive experiments on the SQuAD dataset.The single model achieves 71.3% on the evaluation metrics of exact match on the hidden test set, while the ensemble model further boosts the results to 75.9%.At the time of submission of the paper, our model holds the first place on the SQuAD leaderboard for both single and ensemble model.
Wenhui Wang 0003, Nan Yang 0002, Furu Wei, Baobao Chang, Ming Zhou 0001
ACL (1)4
2017 A Progressive Learning Approach to Chinese SRL Using Heterogeneous Data
abstract
Previous studies on Chinese semantic role labeling (SRL) have concentrated on a single semantically annotated corpus.But the training data of single corpus is often limited.Whereas the other existing semantically annotated corpora for Chinese SRL are scattered across different annotation frameworks.But still, Data sparsity remains a bottleneck.This situation calls for larger training datasets, or effective approaches which can take advantage of highly heterogeneous data.In this paper, we focus mainly on the latter, that is, to improve Chinese SRL by using heterogeneous corpora together.We propose a novel progressive learning model which augments the Progressive Neural Network with Gated Recurrent Adapters.The model can accommodate heterogeneous inputs and effectively transfer knowledge between them.We also release a new corpus, Chinese Sem-Bank, for Chinese SRL 1 .Experiments on CPB 1.0 show that our model outperforms state-of-the-art methods.
Qiaolin Xia, Lei Sha, Baobao Chang, Zhifang Sui
ACL (1)3
2017 A Soft-label Method for Noise-tolerant Distantly Supervised Relation Extraction
abstract
Distant-supervised relation extraction inevitably suffers from wrong labeling problems because it heuristically labels relational facts with knowledge bases.Previous sentence level denoise models don't achieve satisfying performances because they use hard labels which are determined by distant supervision and immutable during training.To this end, we introduce an entity-pair level denoise method which exploits semantic information from correctly labeled entity pairs to correct wrong labels dynamically during training.We propose a joint score function which combines the relational scores based on the entity-pair representation and the confidence of the hard label to obtain a new label, namely a soft label, for certain entity pair.During training, soft labels instead of hard labels serve as gold labels.Experiments on the benchmark dataset show that our method dramatically reduces noisy instances and outperforms the state-of-the-art systems.
Tianyu Liu 0001, Kexiang Wang, Baobao Chang, Zhifang Sui
EMNLP3
2017 Affinity-Preserving Random Walk for Multi-Document Summarization
abstract
Multi-document summarization provides users with a short text that summarizes the information in a set of related documents.This paper introduces affinitypreserving random walk to the summarization task, which preserves the affinity relations of sentences by an absorbing random walk model.Meanwhile, we put forward adjustable affinity-preserving random walk to enforce the diversity constraint of summarization in the random walk process.The ROUGE evaluations on DUC 2003 topic-focused summarization task and DUC 2004 generic summarization task show the good performance of our method, which has the best ROUGE-2 recall among the graph-based ranking methods.
Kexiang Wang, Tianyu Liu 0001, Zhifang Sui, Baobao Chang
EMNLP4
2017 Large-Scale Simple Question Generation by Template-Based Seq2seq Learning
Tianyu Liu 0001, Bingzhen Wei, Baobao Chang, Zhifang Sui
NLPCC3
2016 RBPB: Regularization-Based Pattern Balancing Method for Event Extraction
abstract
Event extraction is a particularly challenging information extraction task, which intends to identify and classify event triggers and arguments from raw text.In recent works, when determining event types (trigger classification), most of the works are either pattern-only or feature-only.However, although patterns cannot cover all representations of an event, it is still a very important feature.In addition, when identifying and classifying arguments, previous works consider each candidate argument separately while ignoring the relationship between arguments.This paper proposes a Regularization-Based Pattern Balancing Method (RBPB).Inspired by the progress in representation learning, we use trigger embedding, sentence-level embedding and pattern features together as our features for trigger classification so that the effect of patterns and other useful features can be balanced.In addition, RBPB uses a regularization method to take advantage of the relationship between arguments.Experiments show that we achieve results better than current state-of-art equivalents.
Lei Sha, Jing Liu 0022, Chin-Yew Lin, Sujian Li, Baobao Chang, Zhifang Sui
ACL (1)5
2016 Graph-based Dependency Parsing with Bidirectional LSTM
abstract
In this paper, we propose a neural network model for graph-based dependency parsing which utilizes Bidirectional LSTM (BLSTM) to capture richer contextual information instead of using high-order factorization, and enable our model to use much fewer features than previous work.In addition, we propose an effective way to learn sentence segment embedding on sentence-level based on an extra forward LSTM network.Although our model uses only first-order factorization, experiments on English Peen Treebank and Chinese Penn Treebank show that our model could be competitive with previous higher-order graph-based dependency parsing models and state-of-the-art models.
Wenhui Wang 0003, Baobao Chang
ACL (1)2
2016 Event Detection with Burst Information Networks
abstract
Retrospective event detection is an important task for discovering previously unidentified events in a text stream. In this paper, we propose two fast centroid-aware event detection models based on a novel text stream representation – Burst Information Networks (BINets) for addressing the challenge. The BINets are time-aware, efficient and can be easily analyzed for identifying key information (centroids). These advantages allow the BINet-based approaches to achieve the state-of-the-art performance on multiple datasets, demonstrating the efficacy of BINets for the task of event detection.
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Ming Zhou 0001
COLING3
2016 Towards Time-Aware Knowledge Graph Completion
abstract
Knowledge graph (KG) completion adds new facts to a KG by making inferences from existing facts. Most existing methods ignore the time information and only learn from time-unknown fact triples. In dynamic environments that evolve over time, it is important and challenging for knowledge graph completion models to take into account the temporal aspects of facts. In this paper, we present a novel time-aware knowledge graph completion model that is able to predict links in a KG using both the existing facts and the temporal information of the facts. To incorporate the happening time of facts, we propose a time-aware KG embedding model using temporal order information among facts. To incorporate the valid time of facts, we propose a joint time-aware inference model based on Integer Linear Programming (ILP) using temporal consistencyinformationasconstraints. Wefurtherintegratetwomodelstomakefulluseofglobal temporal information. We empirically evaluate our models on time-aware KG completion task. Experimental results show that our time-aware models achieve the state-of-the-art on temporal facts consistently.
Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Baobao Chang, Sujian Li, Zhifang Sui
COLING5
2016 Reading and Thinking: Re-read LSTM Unit for Textual Entailment Recognition
abstract
Recognizing Textual Entailment (RTE) is a fundamentally important task in natural language processing that has many applications. The recently released Stanford Natural Language Inference (SNLI) corpus has made it possible to develop and evaluate deep neural network methods for the RTE task. Previous neural network based methods usually try to encode the two sentences (premise and hypothesis) and send them together into a multi-layer perceptron to get their entailment type, or use LSTM-RNN to link two sentences together while using attention mechanic to enhance the model’s ability. In this paper, we propose to use the re-read mechanic, which means to read the premise again and again while reading the hypothesis. After read the premise again, the model can get a better understanding of the premise, which can also affect the understanding of the hypothesis. On the contrary, a better understanding of the hypothesis can also affect the understanding of the premise. With the alternative re-read process, the model can “think” of a better decision of entailment type. We designed a new LSTM unit called re-read LSTM (rLSTM) to implement this “thinking” process. Experiments show that we achieve results better than current state-of-the-art equivalents.
Lei Sha, Baobao Chang, Zhifang Sui, Sujian Li
COLING2
2016 News Stream Summarization using Burst Information Networks
abstract
This paper studies summarizing key information from news streams. We propose simple yet effective models to solve the problem based on a novel and promising representation of text streams – Burst Information Networks (BINets). A BINet can be aware of redundant information, allows global analysis of a text stream, and can be efficiently built and dynamically updated, which perfectly fits the demands of text stream summarization. Extensive experiments show that the BINet-based approaches are not only efficient and can be used in a real-time online summarization setting, but also can generate high-quality summaries, outperforming the state-of-the-art approach.
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Sujian Li, Ming Zhou 0001, Zhifang Sui
EMNLP3
2016 Encoding Temporal Information for Time-Aware Link Prediction
abstract
Most existing knowledge base (KB) embedding methods solely learn from time-unknown fact triples but neglect the temporal information in the knowledge base.In this paper, we propose a novel time-aware KB embedding approach taking advantage of the happening time of facts.Specifically, we use temporal order constraints to model transformation between time-sensitive relations and enforce the embeddings to be temporally consistent and more accurate.We empirically evaluate our approach in two tasks of link prediction and triple classification.Experimental results show that our method outperforms other baselines on the two tasks consistently.
Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui
EMNLP6
2016 Discourse Parsing with Attention-based Hierarchical Neural Networks
Tianshi Li 0002, Baobao Chang
EMNLP3
2016 Capturing Argument Relationship for Chinese Semantic Role Labeling
abstract
In this paper, we capture the argument relationships for Chinese semantic role labeling task, and improve the task's performance with the help of argument relationships.We split the relationship between two candidate arguments into two categories: (1) Compatible arguments: if one candidate argument belongs to a given predicate, then the other is more likely to belong to the same predicate; (2) Incompatible arguments: if one candidate argument belongs to a given predicate, then the other is less likely to belong to the same predicate.However, previous works did not explicitly model argument relationships.We use a simple maximum entropy classifier to capture the two categories of argument relationships and test its performance on the Chinese Proposition Bank (CPB).The experiments show that argument relationships is effective in Chinese semantic role labeling task.
Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui, Tingsong Jiang
EMNLP3
2016 Joint Learning Templates and Slots for Event Schema Induction
abstract
Automatic event schema induction (AESI) means to extract meta-event from raw text, in other words, to find out what types (templates) of event may exist in the raw text and what roles (slots) may exist in each event type.In this paper, we propose a joint entity-driven model to learn templates and slots simultaneously based on the constraints of templates and slots in the same sentence.In addition, the entities' semantic information is also considered for the inner connectivity of the entities.We borrow the normalized cut criteria in image segmentation to divide the entities into more accurate template clusters and slot clusters.The experiment shows that our model gains a relatively higher result than previous work.
Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui
HLT-NAACL3
2015 Bring you to the past: Automatic Generation of Topically Relevant Event Chronicles
abstract
Tao Ge, Wenzhe Pei, Heng Ji, Sujian Li, Baobao Chang, Zhifang Sui. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Tao Ge 0001, Wenzhe Pei, Heng Ji 0001, Sujian Li, Baobao Chang, Zhifang Sui
ACL (1)5
2015 An Effective Neural Network Model for Graph-based Dependency Parsing
abstract
Wenzhe Pei, Tao Ge, Baobao Chang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Wenzhe Pei, Tao Ge 0001, Baobao Chang
ACL (1)3
2015 Distinguishing Specific and Daily Topics
Tao Ge 0001, Wenzhe Pei, Baobao Chang, Zhifang Sui
APWeb3
2015 Multi-label Text Categorization with Joint Learning Predictions-as-Features Method
abstract
Multi-label text categorization is a type of text categorization, where each document is assigned to one or more categories.Recently, a series of methods have been developed, which train a classifier for each label, organize the classifiers in a partially ordered structure and take predictions produced by the former classifiers as the latter classifiers' features.These predictions-asfeatures style methods model high order label dependencies and obtain high performance.Nevertheless, the predictionsas-features methods suffer a drawback.When training a classifier for one label, the predictions-as-features methods can model dependencies between former labels and the current label, but they can't model dependencies between the current label and the latter labels.To address this problem, we propose a novel joint learning algorithm that allows the feedbacks to be propagated from the classifiers for latter labels to the classifier for the current label.We conduct experiments using real-world textual data sets, and these experiments illustrate the predictions-as-features models trained by our algorithm outperform the original models.
Houfeng Wang, Xu Sun 0001, Baobao Chang, Shi Zhao, Lei Sha
EMNLP4
2015 Recognizing Textual Entailment Using Probabilistic Inference
abstract
Recognizing Text Entailment (RTE) plays an important role in NLP applications including question answering, information retrieval, etc.In recent work, some research explore "deep" expressions such as discourse commitments or strict logic for representing the text.However, these expressions suffer from the limitation of inference inconvenience or translation loss.To overcome the limitations, in this paper, we propose to use the predicate-argument structures to represent the discourse commitments extracted from text.At the same time, with the help of the YAGO knowledge, we borrow the distant supervision technique to mine the implicit facts from the text.We also construct a probabilistic network for all the facts and conduct inference to judge the confidence of each fact for RTE.The experimental results show that our proposed method achieves a competitive result compared to the previous work.
Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui, Tingsong Jiang
EMNLP3
2015 Chinese Semantic Role Labeling with Bidirectional Recurrent Neural Networks
abstract
Traditional approaches to Chinese Seman-tic Role Labeling (SRL) almost heavily re-ly on feature engineering. Even worse, the long-range dependencies in a sentence can hardly be modeled by these method-s. In this paper, we introduce bidirection-al recurrent neural network (RNN) with long-short-term memory (LSTM) to cap-ture bidirectional and long-range depen-dencies in a sentence with minimal fea-ture engineering. Experimental results on Chinese Proposition Bank (CPB) show a significant improvement over the state-of-the-art methods. Moreover, our model makes it convenient to introduce hetero-geneous resource, which makes a further improvement on our experimental perfor-mance. 1
Tingsong Jiang, Baobao Chang, Zhifang Sui
EMNLP3
2015 ERSOM: A Structural Ontology Matching Approach Using Automatically Learned Entity Representation
abstract
As a key representation model of knowledge, ontology has been widely used in a lot of NLP related tasks, such as semantic parsing, information extraction and text mining etc.In this paper, we study the task of ontology matching, which concentrates on finding semantically related entities between different ontologies that describe the same domain, to solve the semantic heterogeneity problem.Previous works exploit different kinds of descriptions of an entity in ontology directly and separately to find the correspondences without considering the higher level correlations between the descriptions.Besides, the structural information of ontology haven't been utilized adequately for ontology matching.We propose in this paper an ontology matching approach, named ERSOM, which mainly includes an unsupervised representation learning method based on the deep neural networks to learn the general representation of the entities and an iterative similarity propagation method that takes advantage of more abundant structure information of the ontology to discover more mappings.The experimental results on the datasets from Ontology Alignment Evaluation Initiative (OAEI 1 ) show that ER-SOM achieves a competitive performance compared to the state-of-the-art ontology matching systems.
Chuncheng Xiang, Tingsong Jiang, Baobao Chang, Zhifang Sui
EMNLP3
2015 An Ontology Matching Approach Based on Affinity-Preserving Random Walks
Chuncheng Xiang, Baobao Chang, Zhifang Sui
IJCAI2
2014 Max-Margin Tensor Neural Network for Chinese Word Segmentation
abstract
Recently, neural network models for natural language processing tasks have been increasingly focused on for their ability to alleviate the burden of manual feature engineering.In this paper, we propose a novel neural network model for Chinese word segmentation called Max-Margin Tensor Neural Network (MMTNN).By exploiting tag embeddings and tensorbased transformation, MMTNN has the ability to model complicated interactions between tags and context characters.Furthermore, a new tensor factorization approach is proposed to speed up the model and avoid overfitting.Experiments on the benchmark dataset show that our model achieves better performances than previous neural network models and that our model can achieve a competitive performance with minimal feature engineering.Despite Chinese word segmentation being a specific case, MMTNN can be easily generalized and applied to other sequence labeling tasks.
Wenzhe Pei, Tao Ge 0001, Baobao Chang
ACL (1)3
2014 Inducing Word Sense with Automatically Learned Hidden Concepts
Baobao Chang, Wenzhe Pei, Miaohong Chen
COLING1
2014 A Joint Model for Unsupervised Chinese Word Segmentation
abstract
In this paper, we propose a joint model for unsupervised Chinese word segmentation (CWS). Inspired by the “products of ex-perts ” idea, our joint model firstly com-bines two generative models, which are word-based hierarchical Dirichlet process model and character-based hidden Markov model, by simply multiplying their proba-bilities together. Gibbs sampling is used for model inference. In order to further combine the strength of goodness-based model, we then integrated nVBE into our joint model by using it to initializing the Gibbs sampler. We conduct our experi-ments on PKU and MSRA datasets pro-vided by the second SIGHAN bakeoff. Test results on these two datasets show that the joint model achieves much bet-ter results than all of its component mod-els. Statistical significance tests also show that it is significantly better than state-of-the-art systems, achieving the highest F-scores. Finally, analysis indicates that compared with nVBE and HDP, the joint model has a stronger ability to solve both combinational and overlapping ambigui-ties in Chinese word segmentation. 1
Miaohong Chen, Baobao Chang, Wenzhe Pei
EMNLP2
2013 Exploiting collaborative filtering techniques for automatic assessment of student free-text responses
abstract
The automatic assessment of free-text responses of students is a relatively newer task in both computational linguistics and educational technology. The goal of the task is to produce an assessment of student answers to explanation and definition questions typically asked in problems seen in practice exercises or tests. Unlike some conventional methods which assess the student responses based on only information about their corresponding questions, this paper exploits idea of collaborative filtering to analyze student responses and used an effective collaborative filtering model -- feature-based matrix factorization model to deal with this challenge. The experimental results show that our feature-based matrix factorization model outperforms the baseline models and the model with a re-ranking phase can achieve a better and competitive performance -- 63.6% overall accuracy on the Beetle dataset.
Tao Ge 0001, Zhifang Sui, Baobao Chang
CIKM3
2013 Event-Based Time Label Propagation for Automatic Dating of News Articles
abstract
Since many applications such as timeline summaries and temporal IR involving temporal analysis rely on document timestamps, the task of automatic dating of documents has been increasingly important.Instead of using feature-based methods as conventional models, our method attempts to date documents in a year level by exploiting relative temporal relations between documents and events, which are very effective for dating documents.Based on this intuition, we proposed an eventbased time label propagation model called confidence boosting in which time label information can be propagated between documents and events on a bipartite graph.The experiments show that our event-based propagation model can predict document timestamps in high accuracy and the model combined with a MaxEnt classifier outperforms the state-ofthe-art method for this task especially when the size of the training set is small.
Tao Ge 0001, Baobao Chang, Sujian Li, Zhifang Sui
EMNLP2
2013 Feature-based Neural Language Model and Chinese Word Segmentation
Mairgup Mansur, Wenzhe Pei, Baobao Chang
IJCNLP3
2013 A novel topic model for automatic term extraction
abstract
Automatic term extraction (ATE) aims at extracting domain-specific terms from a corpus of a certain domain. Termhood is one essential measure for judging whether a phrase is a term. Previous researches on termhood mainly depend on the word frequency information. In this paper, we propose to compute termhood based on semantic representation of words. A novel topic model, namely i-SWB, is developed to map the domain corpus into a latent semantic space, which is composed of some general topics, a background topic and a documents-specific topic. Experiments on four domains demonstrate that our approach outperforms the state-of-the-art ATE approaches.
Sujian Li, Wenjie Li 0002, Baobao Chang
SIGIR5
2012 Update Summarization using a Multi-level Hierarchical Dirichlet Process Model
Sujian Li, Baobao Chang
COLING5
2010 Enhancing Domain Portability of Chinese Segmentation Model Using Chi-Square Statistics and Bootstrapping
Baobao Chang, Dongxu Han
EMNLP1
2008 Improving Chinese Semantic Role Classification with Hierarchical Feature Selection Strategy
Weiwei Ding, Baobao Chang
EMNLP2
2005 Extracting Terminologically Relevant Collocations in the Translation of Chinese Monograph
Byeong Kwu Kang, Baobao Chang, Yi-Rong Chen, Shiwen Yu
IJCNLP2
2004 Chinese-English Parallel Corpus Construction and its Application
Baobao Chang
PACLIC1