VLDB 2026 Research / reviewers in the wild / expert
Zhifang Sui
dblp:22/5834
· DBLP profile ↗
95ranked-venue papers
1as first author
39since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 94 · 1 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time ScalingabstractRecent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, including training methodologies, scalability, and generalization capabilities. We investigate the interplay between pre-training and reward model training FLOPs to assess their influence on PRM efficiency and accuracy in complex reasoning tasks. Our analysis reveals a pattern of diminishing returns in performance with increasing PRM scale, highlighting the importance of balancing model size and computational cost. Furthermore, the diversity of training datasets significantly impacts PRM performance, emphasizing the importance of diverse data to enhance both accuracy and efficiency. We further examine test-time scaling strategies, identifying Monte Carlo Tree Search as the most effective method when computational resources are abundant, while Best-of-N Sampling serves as a practical alternative under resource-limited conditions. Notably, our findings indicate that PRMs trained on mathematical datasets exhibit performance comparable to those tailored for code generation, suggesting robust cross-domain generalization. Employing a gradient-based metric, we observe that PRMs exhibit a preference for selecting responses with similar underlying patterns, further informing their optimization. Zhengyu Chen 0001, Teng Xiao, Ruochen Zhou, Xuesheng Yang, Zhifang Sui, Jingang Wang |
AAAI | 7 |
| 2026 | Large Language Models Struggle with Unreasonability in Math ProblemsabstractLarge Language Models (LLMs) have shown remarkable success on a wide range of math and reasoning benchmarks. However, we observe that they often struggle when faced with unreasonable math problems. Instead of recognizing these issues, models frequently proceed as if the problem is well-posed, producing incorrect answers or falling into overthinking and verbose self-correction. To systematically investigate this overlooked vulnerability, we propose the Unreasonable Math Problems (UMP) benchmark, designed to evaluate LLMs' ability to detect and respond to unreasonable math problem statements. Based on extensive experiments covering 19 LLMs, we find that even state-of-the-art general models like GPT-4o struggle on UMP. While reasoning models such as DeepSeek-R1 demonstrate a higher sensitivity to unreasonable inputs, this often comes at the cost of generating overly long and meaningless responses that fail to converge. We further find that prompting and fine-tuning enhance the detection of unreasonable inputs, with minor and acceptable trade-offs, making them practical solutions in this challenging setting. Jingyuan Ma, Damai Dai, Zihang Yuan, Rui Li 0094, Weilin Luo, Lei Sha, Zhifang Sui |
AAAI | 9 |
| 2026 | RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data SelectionabstractData selection for instruction tuning is crucial for improving the performance of large language models (LLMs) while reducing training costs. In this paper, we propose Refined Contribution Measurement with In-Context Learning (RICo), a novel gradient-free method that quantifies the fine-grained contribution of individual samples to both task-level and global-level model performance. RICo enables more accurate identification of high-contribution data, leading to better instruction tuning. We also introduce a lightweight selection paradigm trained on RICo scores, enabling scalable data selection with strictly linear inference complexity. Extensive experiments on 3 LLMs across 12 benchmarks and 5 pairwise evaluation sets demonstrate the effectiveness of RICo. Remarkably, on LLaMA3.1-8B, models trained in 15% of RICo-selected data outperform full datasets by 5.42 percentage points and exceed the best performance of widely used selection methods by 1.48 percentage points. We further analyze high-contribution samples selected by RICo, which show both diverse tasks and appropriate difficulty levels, rather than merely the most difficult cases. Qingxiu Dong, Linli Yao, Fangwei Zhu, Weilin Luo, Zhifang Sui |
AAAI | 7 |
| 2026 | HistLens: Mapping Idea Change across Concepts and CorporaabstractLanguage change both reflects and shapes social processes, and the semantic evolution of foundational concepts provides a measurable trace of historical and social transformation.Despite recent advances in diachronic semantics and discourse analysis, existing computational approaches often (i) concentrate on a single concept or a single corpus, making findings difficult to compare across heterogeneous sources, and (ii) remain confined to surface lexical evidence, offering insufficient computational and interpretive granularity when concepts are expressed implicitly.We propose HistLens, a unified, SAE-based framework for multi-concept, multi-corpus conceptualhistory analysis.The framework decomposes concept representations into interpretable features and tracks their activation dynamics over time and across sources, yielding comparable conceptual trajectories within a shared coordinate system.Experiments on long-span press corpora show that HistLens supports crossconcept, cross-corpus computation of patterns of idea evolution and enables implicit concept computation.By bridging conceptual modeling with interpretive needs, HistLens broadens the analytical perspectives and methodological repertoire available to social science and the humanities for diachronic text analysis. Yi Jing, Weiyun Qiu, Yihang Peng, Zhifang Sui |
ACL (1) | 4 |
| 2026 | Towards Stable and Effective Reinforcement Learning for Mixture-of-ExpertsabstractDi Zhang, Xun Wu, Shaohan Huang, Lingjie Jiang, Yaru Hao, Li Dong, Zewen Chi, Zhifang Sui, Furu Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaohan Huang, Lingjie Jiang, Yaru Hao, Li Dong 0004, Zewen Chi, Zhifang Sui, Furu Wei |
ACL (1) | 8 |
| 2025 | Exploring Activation Patterns of Parameters in Language ModelsabstractMost work treats large language models as black boxes without an in-depth understanding of their internal working mechanism. To explain the internal representations of LLMs, we utilize a gradient-based metric to assess the activation level of model parameters. Based on this metric, we obtain three preliminary findings. (1) When the inputs are in the same domain, parameters in the shallow layers will be activated densely, which means a larger portion of parameters will have great impacts on the outputs. In contrast, parameters in the deep layers are activated sparsely. (2) When the inputs are across different domains, parameters in shallow layers exhibit higher similarity in the activation behavior than in deep layers. (3) In deep layers, the similarity of the distributions of activated parameters is positively correlated to the empirical data relevance. Further, we develop three validation experiments to solidify these findings. (1) Firstly, starting from the first finding, we attempt to configure different sparsities for different layers and find this method can benefit model pruning. (2) Secondly, we find that a pruned model based on one calibration set can better handle tasks related to the calibration task than those not related, which validates the second finding. (3) Thirdly, Based on the STS-B and SICK benchmarks, we find that two sentences with consistent semantics tend to share similar parameter activation patterns in deep layers, which aligns with our third finding. Our work sheds light on the behavior of parameter activation in LLMs, and we hope these findings will have the potential to inspire more practical applications. Yudong Wang 0005, Damai Dai, Zhe Yang 0013, Jingyuan Ma, Zhifang Sui |
AAAI | 5 |
| 2025 | Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMsabstractLarge Language Models (LLMs) can correct their self-generated responses, but a decline in accuracy after self-correction is also witnessed.To have a deeper understanding of selfcorrection, we endeavor to decompose, evaluate, and analyze the self-correction behaviors of LLMs.By enumerating and analyzing answer correctness before and after self-correction, we decompose the self-correction capability into confidence (being confident to correct answers) and critique (turning wrong answers to correct) capabilities, and propose two metrics from a probabilistic perspective to measure these 2 capabilities, along with another metric for overall self-correction capability evaluation.Based on our decomposition and evaluation metrics, we conduct extensive experiments and draw some empirical conclusions.For example, we find different models can exhibit distinct behaviors: some models are confident while others are more critical.We also find the trade-off between the two capabilities (i.e.improving one can lead to a decline in the other) when manipulating model self-correction behavior by prompts or in-context learning.Further, we find a simple yet efficient strategy to improve self-correction capability by transforming Supervision Fine-Tuning (SFT) data format, and our strategy outperforms vanilla SFT in both capabilities and achieves much higher accuracy after self-correction.Our code is publicly available on GitHub. Zhe Yang 0013, Yichang Zhang, Yudong Wang 0005, Ziyao Xu 0001, Junyang Lin, Zhifang Sui |
ACL (1) | 6 |
| 2025 | Towards Harmonized Uncertainty Estimation for Large Language ModelsabstractTo facilitate robust and trustworthy deployment of large language models (LLMs), it is essential to quantify the reliability of their generations through uncertainty estimation.While recent efforts have made significant advancements by leveraging the internal logic and linguistic features of LLMs to estimate uncertainty scores, our empirical analysis highlights the pitfalls of these methods to strike a harmonized estimation between indication, balance, and calibration, which hinders their broader capability for accurate uncertainty estimation.To address this challenge, we propose CUE (Corrector for Uncertainty Estimation): A straightforward yet effective method that employs a lightweight model trained on data aligned with the target LLM's performance to adjust uncertainty scores.Comprehensive experiments across diverse models and tasks demonstrate its effectiveness, which achieves consistent improvements of up to 60% over existing methods.Resources are available at https://github. com/O-L1RU1/Corrector4UE. Rui Li 0094, Jing Long, Muge Qi, Heming Xia, Lei Sha, Peiyi Wang, Zhifang Sui |
ACL (1) | 7 |
| 2025 | Language Models Encode the Value of Numbers LinearlyabstractLarge language models (LLMs) have exhibited impressive competence in various tasks, but their internal mechanisms on mathematical problems are still under-explored. In this paper, we study a fundamental question: how language models encode the value of numbers, a basic element in math. To study the question, we construct a synthetic dataset comprising addition problems and utilize linear probes to read out input numbers from the hidden states. Experimental results support the existence of encoded number values in LLMs on different layers, and these values can be extracted via linear probes. Further experiments show that LLMs store their calculation results in a similar manner, and we can intervene the output via simple vector additions, proving the causal connection between encoded numbers and language model outputs. Our research provides evidence that LLMs encode the value of numbers linearly, offering insights for better exploring, designing, and utilizing numeric information in LLMs. Fangwei Zhu, Damai Dai, Zhifang Sui |
COLING | 3 |
| 2025 | AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient OptimizationabstractRecently, model merging methods have demonstrated powerful strengths in combining abilities on various tasks from multiple Large Language Models (LLMs). While previous model merging methods mainly focus on merging homogeneous models with identical architecture, they meet challenges when dealing with Multimodal Large Language Models (MLLMs) with inherent heterogeneous property, including differences in model architecture and the asymmetry in the parameter space. In this work, we propose AdaMMS1, a novel model merging method tailored for heterogeneous MLLMs. Our method tackles the challenges in three steps: mapping, merging and searching. Specifically, we first design mapping function between models to apply model merging on MLLMs with different architecture. Then we apply linear interpolation on model weights to actively adapt the asymmetry in the heterogeneous MLLMs. Finally in the hyper-parameter searching step, we propose an unsupervised hyper-parameter selection method for model merging. As the first model merging method capable of merging heterogeneous MLLMs without labeled data, extensive experiments on various model combinations demonstrated that AdaMMS outperforms previous model merging methods on various vision-language benchmarks.2 Yiyang Du, Xiaochen Wang 0002, Chi Chen 0005, Jiabo Ye, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Zhifang Sui, Maosong Sun 0001, Yang Liu 0005 |
CVPR | 10 |
| 2025 | A Probabilistic Inference Scaling Theory for LLM Self-CorrectionabstractLarge Language Models (LLMs) have demonstrated the capability to refine their generated answers through self-correction, enabling continuous performance improvement over multiple rounds. However, the mechanisms underlying how and why accuracy evolves during this iterative process remain unexplored. To fill this gap, we propose a probabilistic theory to model the dynamics of accuracy change and explain the performance improvements observed in multi-round self-correction. Through mathematical derivation, we establish that the accuracy after the t^{th} round of self-correction is given by: Acc_t = Upp - \alpha^t(Upp - Acc_0),where Acc_0 denotes the initial accuracy, Upp represents the upper bound of accuracy convergence, and \alpha determines the rate of convergence. Based on our theory, these parameters can be calculated and the predicted accuracy curve then can be obtained through only a single round of self-correction. Extensive experiments across diverse models and datasets demonstrate that our theoretical predictions align closely with empirical accuracy curves, validating the effectiveness of the theory. Our work provides a theoretical foundation for understanding LLM self-correction, thus paving the way for further explorations. Yichang Zhang, Junyang Lin, Zhifang Sui |
EMNLP | 6 |
| 2025 | Self-Boosting Large Language Models with Synthetic Preference DataabstractThrough alignment with human preferences, Large Language Models (LLMs) have advanced significantly in generating honest, harmless, and helpful responses. However, collecting high-quality preference data is a resource-intensive and creativity-demanding process, especially for the continual improvement of LLMs. We introduce SynPO, a self-boosting paradigm that leverages synthetic preference data for model alignment. SynPO employs an iterative mechanism wherein a self-prompt generator creates diverse prompts, and a response improver refines model responses progressively. This approach trains LLMs to autonomously learn the generative rewards for their own outputs and eliminates the need for large-scale annotation of prompts and human preferences. After four SynPO iterations, Llama3-8B and Mistral-7B show significant enhancements in instruction-following abilities, achieving over 22.1% win rate improvements on AlpacaEval 2.0 and ArenaHard. Simultaneously, SynPO improves the general performance of LLMs on various tasks, validated by a 3.2 to 5.0 average score increase on the well-recognized Open LLM leaderboard. Qingxiu Dong, Li Dong 0004, Xingxing Zhang 0002, Zhifang Sui, Furu Wei |
ICLR | 4 |
| 2024 | DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsabstractDamai Dai, Chengqi Deng, Chenggang Zhao, R.x. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y.k. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, Wenfeng Liang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, Wenfeng Liang |
ACL (1) | 16 |
| 2024 | Large Language Models are not Fair EvaluatorsabstractPeiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Peiyi Wang, Lei Li 0039, Liang Chen 0024, Zefan Cai, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu 0049, Tianyu Liu 0001, Zhifang Sui |
ACL (1) | 11 |
| 2024 | Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human AnnotationsabstractPeiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Peiyi Wang, Lei Li 0039, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li 0005, Deli Chen, Zhifang Sui |
ACL (1) | 9 |
| 2024 | FaGANet: An Evidence-Based Fact-Checking Model with Integrated Encoder Leveraging Contextual InformationabstractIn the face of the rapidly growing spread of false and misleading information in the real world, manual evidence-based fact-checking efforts become increasingly challenging and time-consuming. In order to tackle this issue, we propose FaGANet, an automated and accurate fact-checking model that leverages the power of sentence-level attention and graph attention network to enhance performance. This model adeptly integrates encoder-only models with graph attention network, effectively fusing claims and evidence information for accurate identification of even well-disguised data. Experiment results showcase the significant improvement in accuracy achieved by our FaGANet model, as well as its state-of-the-art performance in the evidence-based fact-checking task. We release our code and data in https://github.com/WeiyaoLuo/FaGANet. Weiyao Luo, Junfeng Ran, Zailong Tian, Sujian Li, Zhifang Sui |
LREC/COLING | 5 |
| 2024 | A Survey on In-context LearningabstractQingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, Lei Li, Zhifang Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Qingxiu Dong, Lei Li 0039, Damai Dai, Jingyuan Ma, Rui Li 0094, Heming Xia, Jingjing Xu 0001, Zhiyong Wu 0011, Baobao Chang, Xu Sun 0001, Lei Li 0005, Zhifang Sui |
EMNLP | 13 |
| 2024 | Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?abstractLarge language models (LLMs) have demonstrated impressive capabilities, but still suffer from inconsistency issues (e.g.LLMs can react differently to disturbances like rephrasing or inconsequential order change).In addition to these inconsistencies, we also observe that LLMs, while capable of solving hard problems, can paradoxically fail at easier ones.To evaluate this hard-to-easy inconsistency, we develop the ConsisEval benchmark, where each entry comprises a pair of questions with a strict order of difficulty.Furthermore, we introduce the concept of consistency score to quantitatively measure this inconsistency and analyze the potential for improvement in consistency by relative consistency score.Based on comprehensive experiments across a variety of existing models, we find: (1) GPT-4 achieves the highest consistency score of 92.2% but is still inconsistent to specific questions due to distraction by redundant information, misinterpretation of questions, etc.; (2) models with stronger capabilities typically exhibit higher consistency, but exceptions also exist; (3) hard data enhances consistency for both fine-tuning and in-context learning.Our data and code will be publicly available on GitHub. 1 Zhe Yang 0013, Yichang Zhang, Tianyu Liu 0001, Jian Yang 0003, Junyang Lin, Chang Zhou 0005, Zhifang Sui |
EMNLP | 7 |
| 2023 | Denoising Bottleneck with Mutual Information Maximization for Video Multimodal FusionabstractShaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu, Binghuai Lin, Yunbo Cao, Zhifang Sui. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Zhifang Sui |
ACL (1) | 7 |
| 2023 | Statistical Knowledge Assessment for Large Language ModelsabstractGiven varying prompts regarding a factoid question, can a large language model (LLM) reliably generate factually correct answers? Existing LLMs may generate distinct responses for different prompts. In this paper, we study the problem of quantifying knowledge contained in an LLM regarding a given set of facts. We propose KaRR, a statistical approach to assess factual knowledge for LLMs. The main idea is to estimate the ratio of LLM generating text corresponding to the answer entity given diverse prompts of the subject and the querying relation, versus it generating by random chances. Our assessment suite contains a comprehensive set of 994,123 entities and 600 relations, with 1,395,905 text aliases. We use our method to evaluate 20 LLMs of various sizes, including LLaMA, Alpaca, OPT, etc. Experiments show that our results have a strong correlation (0.43 Kendall's $\tau$) with the results of human assessment on LLMs. Our results reveal that the knowledge in LLMs with the same backbone architecture adheres to the scaling law, while tuning on instruction-following data sometimes compromises the model's capability to generate factually correct text reliably. Qingxiu Dong, Jingjing Xu 0001, Lingpeng Kong, Zhifang Sui, Lei Li 0005 |
NeurIPS | 4 |
| 2023 | Neural Knowledge Bank for Pretrained Transformers
Damai Dai, Wenbin Jiang 0002, Qingxiu Dong, Yajuan Lyu, Zhifang Sui |
NLPCC (2) | 5 |
| 2023 | Mixture-of-Experts for Biomedical Question Answering
Damai Dai, Wenbin Jiang 0002, Yajuan Lyu, Zhifang Sui, Baobao Chang |
NLPCC (1) | 5 |
| 2023 | Coarse-to-Fine Entity Representations for Document-Level Relation Extraction
Damai Dai, Shuang Zeng, Baobao Chang, Zhifang Sui |
NLPCC (2) | 5 |
| 2022 | StableMoE: Stable Routing Strategy for Mixture of ExpertsabstractThe Mixture-of-Experts (MoE) technique can scale up the model size of Transformers with an affordable computational overhead.We point out that existing learning-to-route MoE methods suffer from the routing fluctuation issue, i.e., the target expert of the same input may change along with training, but only one expert will be activated for the input during inference.The routing fluctuation tends to harm sample efficiency because the same input updates different experts but only one is finally used.In this paper, we propose STABLEMOE with two training stages to address the routing fluctuation problem.In the first training stage, we learn a balanced and cohesive routing strategy and distill it into a lightweight router decoupled from the backbone model.In the second training stage, we utilize the distilled router to determine the token-to-expert assignment and freeze it for a stable routing strategy.We validate our method on language modeling and multilingual machine translation.The results show that STABLEMOE outperforms existing MoE methods in terms of both convergence speed and performance. Damai Dai, Li Dong 0004, Shuming Ma, Bo Zheng 0010, Zhifang Sui, Baobao Chang, Furu Wei |
ACL (1) | 5 |
| 2022 | Knowledge Neurons in Pretrained TransformersabstractLarge-scale pretrained language models are surprisingly good at recalling factual knowledge presented in the training corpus (Petroni et al., 2019; Jiang et al., 2020b).In this paper, we present preliminary studies on how factual knowledge is stored in pretrained Transformers by introducing the concept of knowledge neurons.Specifically, we examine the fill-in-the-blank cloze task for BERT.Given a relational fact, we propose a knowledge attribution method to identify the neurons that express the fact.We find that the activation of such knowledge neurons is positively correlated to the expression of their corresponding facts.In our case studies, we attempt to leverage knowledge neurons to edit (such as update, and erase) specific factual knowledge without fine-tuning.Our results shed light on understanding the storage of knowledge within pretrained Transformers.The code is available at https://github.com/ Hunter-DDM/knowledge-neurons. Damai Dai, Li Dong 0004, Yaru Hao, Zhifang Sui, Baobao Chang, Furu Wei |
ACL (1) | 4 |
| 2022 | Premise-based Multimodal Reasoning: Conditional Inference on Joint Textual and Visual CluesabstractQingxiu Dong, Ziwei Qin, Heming Xia, Tian Feng, Shoujie Tong, Haoran Meng, Lin Xu, Zhongyu Wei, Weidong Zhan, Baobao Chang, Sujian Li, Tianyu Liu, Zhifang Sui. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Qingxiu Dong, Ziwei Qin, Heming Xia, Shoujie Tong, Haoran Meng, Zhongyu Wei, Weidong Zhan, Baobao Chang, Sujian Li, Tianyu Liu 0001, Zhifang Sui |
ACL (1) | 13 |
| 2022 | A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationabstractTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, Bill Dolan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Tianyu Liu 0001, Yizhe Zhang 0002, Chris Brockett, Zhifang Sui, Weizhu Chen, William B. Dolan |
ACL (1) | 5 |
| 2022 | CBLUE: A Chinese Biomedical Language Understanding Evaluation BenchmarkabstractNingyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li, Xin Shang, Kangping Yin, Chuanqi Tan, Jian Xu, Fei Huang, Luo Si, Yuan Ni, Guotong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan, Linfeng Li, Jun Yan, Hongying Zan, Kunli Zhang, Buzhou Tang, Qingcai Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Ningyu Zhang 0001, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li 0040, Xin Shang, Kangping Yin, Chuanqi Tan, Fei Huang 0002, Luo Si, Yuan Ni, Guo Tong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan 0002, Jun Yan 0010, Hongying Zan, Kunli Zhang, Buzhou Tang, Qingcai Chen |
ACL (1) | 14 |
| 2022 | Learning Robust Representations for Continual Relation Extraction via Adversarial Class AugmentationabstractContinual relation extraction (CRE) aims to continually learn new relations from a classincremental data stream.CRE model usually suffers from catastrophic forgetting problem, i.e., the performance of old relations seriously degrades when the model learns new relations.Most previous work attributes catastrophic forgetting to the corruption of the learned representations as new relations come, with an implicit assumption that the CRE models have adequately learned the old relations.In this paper, through empirical studies we argue that this assumption may not hold, and an important reason for catastrophic forgetting is that the learned representations do not have good robustness against the appearance of analogous relations in the subsequent learning process.To address this issue, we encourage the model to learn more precise and robust representations through a simple yet effective adversarial class augmentation mechanism (ACA), which is easy to implement and model-agnostic.Experimental results show that ACA can consistently improve the performance of state-of-theart CRE models on two popular benchmarks. Peiyi Wang, Yifan Song 0002, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Sujian Li, Zhifang Sui |
EMNLP | 7 |
| 2022 | HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text ClassificationabstractHierarchical text classification (HTC) is a challenging subtask of multi-label classification due to its complex label hierarchy.Recently, the pretrained language models (PLM) have been widely adopted in HTC through a finetuning paradigm.However, in this paradigm, there exists a huge gap between the classification tasks with sophisticated label hierarchy and the masked language model (MLM) pretraining tasks of PLMs and thus the potential of PLMs cannot be fully tapped.To bridge the gap, in this paper, we propose HPT, a Hierarchy-aware Prompt Tuning method to handle HTC from a multi-label MLM perspective.Specifically, we construct a dynamic virtual template and label words that take the form of soft prompts to fuse the label hierarchy knowledge and introduce a zero-bounded multi-label cross-entropy loss to harmonize the objectives of HTC and MLM.Extensive experiments show HPT achieves state-of-the-art performances on 3 popular HTC datasets and is adept at handling the imbalance and low resource situations. Peiyi Wang, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Zhifang Sui, Houfeng Wang |
EMNLP | 6 |
| 2022 | Robust Fine-tuning via Perturbation and Interpolation from In-batch InstancesabstractFine-tuning pretrained language models (PLMs) on downstream tasks has become common practice in natural language processing. However, most of the PLMs are vulnerable, e.g., they are brittle under adversarial attacks or imbalanced data, which hinders the application of the PLMs on some downstream tasks, especially in safe-critical scenarios. In this paper, we propose a simple yet effective fine-tuning method called Match-Tuning to force the PLMs to be more robust. For each instance in a batch, we involve other instances in the same batch to interact with it. To be specific, regarding the instances with other labels as a perturbation, Match-Tuning makes the model more robust to noise at the beginning of training. While nearing the end, Match-Tuning focuses more on performing an interpolation among the instances with the same label for better generalization. Extensive experiments on various tasks in GLUE benchmark show that Match-Tuning consistently outperforms the vanilla fine-tuning by 1.64 scores. Moreover, Match-Tuning exhibits remarkable robustness to adversarial attacks and data imbalance. Shoujie Tong, Qingxiu Dong, Damai Dai, Yifan Song 0002, Tianyu Liu 0001, Baobao Chang, Zhifang Sui |
IJCAI | 7 |
| 2022 | CLINER: Clinical Interrogation Named Entity Recognition
Tianyang Cao, Yifan Yang 0008, Yunyan Zhang, Xi Chen 0003, Baobao Chang, Zhifang Sui, Ruihui Zhao, Yefeng Zheng 0001, Bang Liu 0003 |
KSEM (2) | 8 |
| 2022 | An Enhanced Span-based Decomposition Method for Few-Shot Sequence LabelingabstractPeiyi Wang, Runxin Xu, Tianyu Liu, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui |
NAACL-HLT | 7 |
| 2022 | A Two-Stream AMR-enhanced Model for Document-level Event Argument ExtractionabstractRunxin Xu, Peiyi Wang, Tianyu Liu, Shuang Zeng, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Runxin Xu, Peiyi Wang, Tianyu Liu 0001, Shuang Zeng, Baobao Chang, Zhifang Sui |
NAACL-HLT | 6 |
| 2022 | Plug-and-Play Module for Commonsense Reasoning in Machine Reading Comprehension
Damai Dai, Zhifang Sui, Baobao Chang |
NLPCC (2) | 3 |
| 2021 | Towards Faithfulness in Open Domain Table-to-text Generation from an Entity-centric ViewabstractIn open domain table-to-text generation, we notice the unfaithful generation usually contains hallucinated entities which can not be aligned to any input table record. We thus try to evaluate the generation faithfulness with two entity-centric metrics: table record coverage and the ratio of hallucinated entities in text, both of which are shown to have strong agreement with human judgements. Then based on these metrics, we quantitatively analyze the correlation between training data quality and generation fidelity which indicates the potential usage of entity information in faithful generation. Motivated by these findings, we propose two methods for faithful generation: 1) augmented training by incorporating the auxiliary entity information, including both an augmented plan-based model and an unsupervised model and 2) training instance selection based on faithfulness ranking. We show these approaches improve generation fidelity in both full dataset setting and few shot setting by both automatic and human evaluations. Tianyu Liu 0001, Baobao Chang, Zhifang Sui |
AAAI | 4 |
| 2021 | Behind the Scenes: An Exploration of Trigger Biases Problem in Few-Shot Event ClassificationabstractFew-Shot Event Classification (FSEC) aims at developing a model for event prediction, which can generalize to new event types with a limited number of annotated data. Existing FSEC studies have achieved high accuracy on different benchmarks. However, we find they suffer from trigger biases that signify the statistical homogeneity between some trigger words and target event types, which we summarize as trigger overlapping and trigger separability. The biases can result in context-bypassing problem, i.e., correct classifications can be gained by looking at only the trigger words while ignoring the entire context. Therefore, existing models can be weak in generalizing to unseen data in real scenarios. To further uncover the trigger biases and assess the generalization ability of the models, we propose two new sampling methods, Trigger-Uniform Sampling (TUS) and COnfusion Sampling (COS), for the meta tasks construction during evaluation. Besides, to cope with the context-bypassing problem in FSEC models, we introduce adversarial training and trigger reconstruction techniques. Experiments show these techniques help not only improve the performance, but also enhance the generalization ability of models. Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Damai Dai, Baobao Chang, Zhifang Sui |
CIKM | 6 |
| 2021 | Decompose, Fuse and Generate: A Formation-Informed Method for Chinese Definition GenerationabstractHua Zheng, Damai Dai, Lei Li, Tianyu Liu, Zhifang Sui, Baobao Chang, Yang Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Damai Dai, Lei Li 0039, Tianyu Liu 0001, Zhifang Sui, Baobao Chang, Yang Liu 0124 |
NAACL-HLT | 5 |
| 2021 | XGPT: Cross-modal Generative Pre-Training for Image Captioning
Qiaolin Xia, Haoyang Huang, Nan Duan 0001, Dongdong Zhang 0001, Lei Ji 0001, Zhifang Sui, Edward Dong Bo Cui, Taroon Bharti, Ming Zhou 0001 |
NLPCC (1) | 6 |
| 2020 | Soap: Soaking Capacity Optimization for Multi-Document SummarizationabstractMulti-document summarization (MDS) aims at giving a brief summary for a cluster of related documents. In this paper, we consider the MDS task as an optimization problem with a novel measure named soaking capacity being the objective function. The origin of our method is the classic hypothesis: the summary components are the sinks of information diffusion. We point out that the hypothesis only gives the role of summary but does not cover how well a summary acts as this role. To fill in the gap, soaking capacity is formally defined to quantify the ability of summary to soak up information. We explicitly demonstrate its fitness as an indicator for both the saliency and the diversity goal of MDS. For solving the optimization problem, we propose a greedy algorithm named Soap by adopting a surrogate of soaking capacity to accelerate the computation. Experiments on MDS datasets across various domains show the great potential of Soap as compared with the state-of-the-art MDS systems. Kexiang Wang, Baobao Chang, Zhifang Sui |
CIKM | 3 |
| 2020 | An Anchor-Based Automatic Evaluation Metric for Document SummarizationabstractThe widespread adoption of reference-based automatic evaluation metrics such as ROUGE has promoted the development of document summarization.In this paper, we consider a new protocol for designing reference-based metrics that require the endorsement of source document(s).Following protocol, we propose an anchored ROUGE metric fixing each summary particle on source document, which bases the computation on more solid ground.Empirical results on benchmark datasets validate that source document helps to induce a higher correlation with human judgments for ROUGE metric.Being self-explanatory and easy-to-implement, the protocol can naturally foster various effective designs of reference-based metrics besides the anchored ROUGE introduced here. Kexiang Wang, Tianyu Liu 0001, Baobao Chang, Zhifang Sui |
COLING | 4 |
| 2020 | An Empirical Study on Model-agnostic Debiasing Strategies for Robust Natural Language InferenceabstractThe prior work on natural language inference (NLI) debiasing mainly targets at one or few known biases while not necessarily making the models more robust.In this paper, we focus on the model-agnostic debiasing strategies and explore how to (or is it possible to) make the NLI models robust to multiple distinct adversarial attacks while keeping or even strengthening the models' generalization power.We firstly benchmark prevailing neural NLI models including pretrained ones on various adversarial datasets.We then try to combat distinct known biases by modifying a mixture of experts (MoE) ensemble method (Clark et al., 2019) and show that it's nontrivial to mitigate multiple NLI biases at the same time, and that model-level ensemble method outperforms MoE ensemble method.We also perform data augmentation including text swap, word substitution and paraphrase and prove its efficiency in combating various (though not all) adversarial attacks at the same time.Finally, we investigate several methods to merge heterogeneous training data (1.35M) and perform model ensembling, which are straightforward but effective to strengthen NLI models. Tianyu Liu 0001, Xiaoan Ding, Baobao Chang, Zhifang Sui |
CoNLL | 5 |
| 2020 | Discriminatively-Tuned Generative Classifiers for Robust Natural Language InferenceabstractWhile discriminative neural network classifiers are generally preferred, recent work has shown advantages of generative classifiers in term of data efficiency and robustness.In this paper, we focus on natural language inference (NLI).We propose GenNLI, a generative classifier for NLI tasks, and empirically characterize its performance by comparing it to five baselines, including discriminative models and large-scale pretrained language representation models like BERT.We explore training objectives for discriminative fine-tuning of our generative classifiers, showing improvements over log loss fine-tuning from prior work (Lewis and Fan, 2019).In particular, we find strong results with a simple unbounded modification to log loss, which we call the "infinilog loss".Our experiments show that GenNLI outperforms both discriminative and pretrained baselines across several challenging NLI experimental settings, including small training sets, imbalanced label distributions, and label noise. Xiaoan Ding, Tianyu Liu 0001, Baobao Chang, Zhifang Sui, Kevin Gimpel |
EMNLP (1) | 4 |
| 2020 | A Spectral Method for Unsupervised Multi-Document SummarizationabstractMulti-document summarization (MDS) aims at producing a good-quality summary for several related documents.In this paper, we propose a spectral-based hypothesis, which states that the goodness of summary candidate is closely linked to its so-called spectral impact.Here spectral impact considers the perturbation to the dominant eigenvalue of affinity matrix when dropping the summary candidate from the document cluster.The hypothesis is validated by three theoretical perspectives: semantic scaling, propagation dynamics and matrix perturbation.According to the hypothesis, we formulate the MDS task as the combinatorial optimization of spectral impact and propose an accelerated greedy solution based on a surrogate of spectral impact.The evaluation results on various datasets demonstrate:(1) The performance of the summary candidate is positively correlated with its spectral impact, which accords with our hypothesis; (2) Our spectral-based method has a competitive result as compared to state-of-the-art MDS systems. Kexiang Wang, Baobao Chang, Zhifang Sui |
EMNLP (1) | 3 |
| 2020 | HypoNLI: Exploring the Artificial Patterns of Hypothesis-only Bias in Natural Language InferenceabstractMany recent studies have shown that for models trained on datasets for natural language inference (NLI), it is possible to make correct predictions by merely looking at the hypothesis while completely ignoring the premise. In this work, we manage to derive adversarial examples in terms of the hypothesis-only bias and explore eligible ways to mitigate such bias. Specifically, we extract various phrases from the hypotheses (artificial patterns) in the training sets, and show that they have been strong indicators to the specific labels. We then figure out ‘hard’ and ‘easy’ instances from the original test sets whose labels are opposite to or consistent with those indications. We also set up baselines including both pretrained models (BERT, RoBerta, XLNet) and competitive non-pretrained models (InferSent, DAM, ESIM). Apart from the benchmark and baselines, we also investigate two debiasing approaches which exploit the artificial pattern modeling to mitigate such hypothesis-only bias: down-sampling and adversarial training. We believe those methods can be treated as competitive baselines in NLI debiasing tasks. Tianyu Liu 0001, Baobao Chang, Zhifang Sui |
LREC | 4 |
| 2019 | WSD-GAN: Word Sense Disambiguation Using Generative Adversarial NetworksabstractWord Sense Disambiguation (WSD), as a tough task in Natural Language Processing (NLP), aims to identify the correct sense of an ambiguous word in a given context. There are two mainstreams in WSD. Supervised methods mainly utilize labeled context to train a classifier which generates the right probability distribution of word senses. Meanwhile knowledge-based (unsupervised) methods which focus on glosses (word sense definitions) always calculate the similarity of context-gloss pair as score to find out the right word sense. In this paper, we propose a generative adversarial framework WSD-GAN which combines two mainstream methods in WSD. The generative model, based on supervised methods, tries to generate a probability distribution over the word senses. Meanwhile the discriminative model, based on knowledge-based methods, focuses on predicting the relevancy of the context-gloss pairs and identifies the correct pairs over the others. Furthermore, in order to optimize both two models, we leverage policy gradient to enhance the performances of the two models mutually. Our experimental results show that WSD-GAN achieves competitive results on several English all-words WSD datasets. Fuli Luo, Yutong Tan, Wenxin Zeng, Zhifang Sui |
AAAI | 5 |
| 2019 | Hierarchical Encoder with Auxiliary Supervision for Neural Table-to-Text Generation: Learning Better Representation for TablesabstractGenerating natural language descriptions for the structured tables which consist of multiple attribute-value tuples is a convenient way to help people to understand the tables. Most neural table-to-text models are based on the encoder-decoder framework. However, it is hard for a vanilla encoder to learn the accurate semantic representation of a complex table. The challenges are two-fold: firstly, the table-to-text datasets often contain large number of attributes across different domains, thus it is hard for the encoder to incorporate these heterogeneous resources. Secondly, the single encoder also has difficulties in modeling the complex attribute-value structure of the tables. To this end, we first propose a two-level hierarchical encoder with coarse-to-fine attention to handle the attribute-value structure of the tables. Furthermore, to capture the accurate semantic representations of the tables, we propose 3 joint tasks apart from the prime encoder-decoder learning, namely auxiliary sequence labeling task, text autoencoder and multi-labeling classification, as the auxiliary supervisions for the table encoder. We test our models on the widely used dataset WIKIBIO which contains Wikipedia infoboxes and related descriptions. The dataset contains complex tables as well as large number of attributes across different domains. We achieve the state-of-the-art performance on both automatic and human evaluation metrics. Tianyu Liu 0001, Fuli Luo, Qiaolin Xia, Shuming Ma, Baobao Chang, Zhifang Sui |
AAAI | 6 |
| 2019 | Towards Comprehensive Description Generation from Factual Attribute-value TablesabstractThe comprehensive descriptions for factual attribute-value tables, which should be accurate, informative and loyal, can be very helpful for end users to understand the structured data in this form.However previous neural generators might suffer from key attributes missing, less informative and groundless information problems, which impede the generation of high-quality comprehensive descriptions for tables.To relieve these problems, we first propose force attention (FA) method to encourage the generator to pay more attention to the uncovered attributes to avoid potential key attributes missing.Furthermore, we propose reinforcement learning for information richness to generate more informative as well as more loyal descriptions for tables.In our experiments, we utilize the widely used WIKIBIO dataset as a benchmark.Additionally we create WB-filter based on WIKIBIO to test our model in the simulated user-oriented scenarios, in which the generated descriptions should accord with particular user interests.Experimental results show that our model outperforms the state-of-the-art baselines on both automatic and human evaluation. Tianyu Liu 0001, Fuli Luo, Wei Wu 0044, Baobao Chang, Zhifang Sui |
ACL (1) | 6 |
| 2019 | Learning to Control the Fine-grained Sentiment for Story Ending GenerationabstractAutomatic story ending generation is an interesting and challenging task in natural language generation.Previous studies are mainly limited to generate coherent, reasonable and diversified story endings, and few works focus on controlling the sentiment of story endings.This paper focuses on generating a story ending which meets the given fine-grained sentiment intensity.There are two major challenges to this task.First is the lack of story corpus which has fine-grained sentiment labels.Second is the difficulty of explicitly controlling sentiment intensity when generating endings.Therefore, we propose a generic and novel framework which consists of a sentiment analyzer and a sentimental generator, respectively addressing the two challenges.The sentiment analyzer adopts a series of methods to acquire sentiment intensities of the story dataset.The sentimental generator introduces the sentiment intensity into decoder via a Gaussian Kernel Layer to control the sentiment of the output.To the best of our knowledge, this is the first endeavor to control the fine-grained sentiment for story ending generation without manually annotating sentiment labels.Experiments show that our proposed framework can generate story endings which are not only more coherent and fluent but also able to meet the given sentiment intensity better. 1 Fuli Luo, Damai Dai, Tianyu Liu 0001, Baobao Chang, Zhifang Sui, Xu Sun 0001 |
ACL (1) | 6 |
| 2019 | Towards Fine-grained Text Sentiment TransferabstractIn this paper, we focus on the task of finegrained text sentiment transfer (FGST).This task aims to revise an input sequence to satisfy a given sentiment intensity, while preserving the original semantic content.Different from conventional sentiment transfer task that only reverses the sentiment polarity (positive/negative) of text, the FTST task requires more nuanced and fine-grained control of sentiment.To remedy this, we propose a novel Seq2SentiSeq model.Specifically, the numeric sentiment intensity value is incorporated into the decoder via a Gaussian kernel layer to finely control the sentiment intensity of the output.Moreover, to tackle the problem of lacking parallel data, we propose a cycle reinforcement learning algorithm to guide the model training.In this framework, the elaborately designed rewards can balance both sentiment transformation and content preservation, while not requiring any ground truth output.Experimental results show that our approach can outperform existing methods by a large margin in both automatic evaluation and human evaluation.Our code and data, including outputs of all baselines and our model are available at https://github.com/luofuli/ Fine-grained-Sentiment-Transfer. 1 Fuli Luo, Peng Li 0030, Jie Zhou 0016, Yutong Tan, Baobao Chang, Zhifang Sui, Xu Sun 0001 |
ACL (1) | 7 |
| 2019 | Pun-GAN: Generative Adversarial Network for Pun GenerationabstractFuli Luo, Shunyao Li, Pengcheng Yang, Lei Li, Baobao Chang, Zhifang Sui, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Fuli Luo, Shunyao Li, Lei Li 0039, Baobao Chang, Zhifang Sui, Xu Sun 0001 |
EMNLP/IJCNLP (1) | 6 |
| 2019 | A Dual Reinforcement Learning Framework for Unsupervised Text Style TransferabstractUnsupervised text style transfer aims to transfer the underlying style of text but keep its main content unchanged without parallel data. Most existing methods typically follow two steps: first separating the content from the original style, and then fusing the content with the desired style. However, the separation in the first step is challenging because the content and style interact in subtle ways in natural language. Therefore, in this paper, we propose a dual reinforcement learning framework to directly transfer the style of the text via a one-step mapping model, without any separation of content and style. Specifically, we consider the learning of the source-to-target and target-to-source mappings as a dual task, and two rewards are designed based on such a dual structure to reflect the style accuracy and content preservation, respectively. In this way, the two one-step mapping models can be trained via reinforcement learning, without any use of parallel data. Automatic evaluations show that our model outperforms the state-of-the-art systems by a large margin, especially with more than 10 BLEU points improvement averaged on two benchmark datasets. Human evaluations also validate the effectiveness of our model in terms of style accuracy, content preservation and fluency. Our code and data, including outputs of all baselines and our model are available at https://github.com/luofuli/DualRL. Fuli Luo, Peng Li 0030, Jie Zhou 0016, Baobao Chang, Xu Sun 0001, Zhifang Sui |
IJCAI | 7 |
| 2018 | Table-to-Text Generation by Structure-Aware Seq2seq LearningabstractTable-to-text generation aims to generate a description for a factual table which can be viewed as a set of field-value records. To encode both the content and the structure of a table, we propose a novel structure-aware seq2seq architecture which consists of field-gating encoder and description generator with dual attention. In the encoding phase, we update the cell memory of the LSTM unit by a field gate and its corresponding field value in order to incorporate field information into table representation. In the decoding phase, dual attention mechanism which contains word level attention and field level attention is proposed to model the semantic relevance between the generated description and the table. We conduct experiments on the WIKIBIO dataset which contains over 700k biographies and corresponding infoboxes from Wikipedia. The attention visualizations and case studies show that our model is capable of generating coherent and informative descriptions based on the comprehensive understanding of both the content and the structure of a table. Automatic evaluations also show our model outperforms the baselines by a great margin. Code for this work is available on https://github.com/tyliupku/wiki2bio. Tianyu Liu 0001, Kexiang Wang, Lei Sha, Baobao Chang, Zhifang Sui |
AAAI | 5 |
| 2018 | Order-Planning Neural Text Generation From Structured DataabstractGenerating texts from structured data (e.g., a table) is important for various natural language processing tasks such as question answering and dialog systems. In recent studies, researchers use neural language models and encoder-decoder frameworks for table-to-text generation. However, these neural network-based approaches typically do not model the order of content during text generation. When a human writes a summary based on a given table, he or she would probably consider the content order before wording. In this paper, we propose an order-planning text generation model, where order information is explicitly captured by link-based attention. Then a self-adaptive gate combines the link-based attention with traditional content-based attention. We conducted experiments on the WikiBio dataset and achieve higher performance than previous methods in terms of BLEU, ROUGE, and NIST scores; we also performed ablation tests to analyze each component of our model. Lei Sha, Lili Mou, Tianyu Liu 0001, Pascal Poupart, Sujian Li, Baobao Chang, Zhifang Sui |
AAAI | 7 |
| 2018 | Jointly Extracting Event Triggers and Arguments by Dependency-Bridge RNN and Tensor-Based Argument InteractionabstractEvent extraction plays an important role in natural language processing (NLP) applications including question answering and information retrieval. Traditional event extraction relies heavily on lexical and syntactic features, which require intensive human engineering and may not generalize to different datasets. Deep neural networks, on the other hand, are able to automatically learn underlying features, but existing networks do not make full use of syntactic relations. In this paper, we propose a novel dependency bridge recurrent neural network (dbRNN) for event extraction. We build our model upon a recurrent neural network, but enhance it with dependency bridges, which carry syntactically related information when modeling each word.We illustrates that simultaneously applying tree structure and sequence structure in RNN brings much better performance than only uses sequential RNN. In addition, we use a tensor layer to simultaneously capture the various types of latent interaction between candidate arguments as well as identify/classify all arguments of an event. Experiments show that our approach achieves competitive results compared with previous work. Lei Sha, Baobao Chang, Zhifang Sui |
AAAI | 4 |
| 2018 | A Multi-View Fusion Neural Network for Answer SelectionabstractCommunity question answering aims at choosing the most appropriate answer for a given question, which is important in many NLP applications. Previous neural network-based methods consider several different aspects of information through calculating attentions. These different kinds of attentions are always simply summed up and can be seen as a ``single view", causing severe information loss. To overcome this problem, we propose a Multi-View Fusion Neural Network, where each attention component generates a ``view'' of the QA pair and a fusion RNN integrates the generated views to form a more holistic representation. In this fusion RNN method, a filter gate collects important information of input and directly adds it to the output, which borrows the idea of residual networks. Experimental results on the WikiQA and SemEval-2016 CQA datasets demonstrate that our proposed model outperforms the state-of-the-art methods. Lei Sha, Xiaodong Zhang 0022, Baobao Chang, Zhifang Sui |
AAAI | 5 |
| 2018 | Incorporating Glosses into Neural Word Sense DisambiguationabstractWord Sense Disambiguation (WSD) aims to identify the correct meaning of polysemous words in the particular context.Lexical resources like WordNet which are proved to be of great help for WSD in the knowledge-based methods.However, previous neural networks for WSD always rely on massive labeled data (context), ignoring lexical resources like glosses (sense definitions).In this paper, we integrate the context and glosses of the target word into a unified framework in order to make full use of both labeled data and lexical knowledge.Therefore, we propose GAS: a gloss-augmented WSD neural network which jointly encodes the context and glosses of the target word.GAS models the semantic relationship between the context and the gloss in an improved memory network framework, which breaks the barriers of the previous supervised methods and knowledge-based methods.We further extend the original gloss of word sense via its semantic relations in WordNet to enrich the gloss information.The experimental results show that our model outperforms the state-of-theart systems on several English all-words WSD datasets. Fuli Luo, Tianyu Liu 0001, Qiaolin Xia, Baobao Chang, Zhifang Sui |
ACL (1) | 5 |
| 2018 | Modeling Consumer Buying Decision for Recommendation Based on Multi-Task Deep LearningabstractAlthough marketing researchers and sociologists have recognized the importance of buying decision process and its significant influence on consumer's purchasing behaviors, existing recommender systems do not explicitly model the consumer buying decision process or capture the sequential regularities of what happens before and after each purchase. In this paper, we try to bridge the gap and improve recommendation systems by explicitly modeling consumer buying decision process and corresponding stages. In particular, we propose a multi-task learning model with long short-term memory networks (LSTM) to learn consumer buying decision process. It maps items, users, product categories, and the behavior sequences into real valued vectors, with which the probability of purchasing a product can be estimated. In this way, the model can capture user intentions and preferences, predicts the conversion rate of each candidate product, and makes recommendations accordingly. Experiments on real world data demonstrate the effectiveness of the proposed approach. Qiaolin Xia, Peng Jiang 0002, Fei Sun 0001, Yi Zhang 0001, Xiaobo Wang 0002, Zhifang Sui |
CIKM | 6 |
| 2018 | Fine-grained Coordinated Cross-lingual Text Stream Alignment for Endless Language Knowledge AcquisitionabstractThis paper proposes to study fine-grained coordinated cross-lingual text stream alignment through a novel information network decipherment paradigm.We use Burst Information Networks as media to represent text streams and present a simple yet effective network decipherment algorithm with diverse clues to decipher the networks for accurate text stream alignment.Experiments on Chinese-English news streams show our approach not only outperforms previous approaches on bilingual lexicon extraction from coordinated text streams but also can harvest high-quality alignments from large amounts of streaming data for endless language knowledge mining, which makes it promising to be a new paradigm for automatic language knowledge acquisition. Tao Ge 0001, Qing Dou, Heng Ji 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
EMNLP | 6 |
| 2018 | Leveraging Gloss Knowledge in Neural Word Sense Disambiguation by Hierarchical Co-AttentionabstractThe goal of Word Sense Disambiguation (WSD) is to identify the correct meaning of a word in the particular context.Traditional supervised methods only use labeled data (context), while missing rich lexical knowledge such as the gloss which defines the meaning of a word sense.Recent studies have shown that incorporating glosses into neural networks for WSD has made significant improvement.However, the previous models usually build the context representation and gloss representation separately.In this paper, we find that the learning for the context and gloss representation can benefit from each other.Gloss can help to highlight the important words in the context, thus building a better context representation.Context can also help to locate the key words in the gloss of the correct word sense.Therefore, we introduce a co-attention mechanism to generate co-dependent representations for the context and gloss.Furthermore, in order to capture both word-level and sentence-level information, we extend the attention mechanism in a hierarchical fashion.Experimental results show that our model achieves the state-of-the-art results on several standard English all-words WSD test datasets. Fuli Luo, Tianyu Liu 0001, Zexue He, Qiaolin Xia, Zhifang Sui, Baobao Chang |
EMNLP | 5 |
| 2018 | EventWiki: A Knowledge Base of Major Events
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
LREC | 4 |
| 2018 | Revisiting Distant Supervision for Relation Extraction
Tingsong Jiang, Jing Liu 0022, Chin-Yew Lin, Zhifang Sui |
LREC | 4 |
| 2018 | SeRI: A Dataset for Sub-event Relation Inference from an Encyclopedia
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
NLPCC (2) | 4 |
| 2017 | A Progressive Learning Approach to Chinese SRL Using Heterogeneous DataabstractPrevious studies on Chinese semantic role labeling (SRL) have concentrated on a single semantically annotated corpus.But the training data of single corpus is often limited.Whereas the other existing semantically annotated corpora for Chinese SRL are scattered across different annotation frameworks.But still, Data sparsity remains a bottleneck.This situation calls for larger training datasets, or effective approaches which can take advantage of highly heterogeneous data.In this paper, we focus mainly on the latter, that is, to improve Chinese SRL by using heterogeneous corpora together.We propose a novel progressive learning model which augments the Progressive Neural Network with Gated Recurrent Adapters.The model can accommodate heterogeneous inputs and effectively transfer knowledge between them.We also release a new corpus, Chinese Sem-Bank, for Chinese SRL 1 .Experiments on CPB 1.0 show that our model outperforms state-of-the-art methods. Qiaolin Xia, Lei Sha, Baobao Chang, Zhifang Sui |
ACL (1) | 4 |
| 2017 | A Soft-label Method for Noise-tolerant Distantly Supervised Relation ExtractionabstractDistant-supervised relation extraction inevitably suffers from wrong labeling problems because it heuristically labels relational facts with knowledge bases.Previous sentence level denoise models don't achieve satisfying performances because they use hard labels which are determined by distant supervision and immutable during training.To this end, we introduce an entity-pair level denoise method which exploits semantic information from correctly labeled entity pairs to correct wrong labels dynamically during training.We propose a joint score function which combines the relational scores based on the entity-pair representation and the confidence of the hard label to obtain a new label, namely a soft label, for certain entity pair.During training, soft labels instead of hard labels serve as gold labels.Experiments on the benchmark dataset show that our method dramatically reduces noisy instances and outperforms the state-of-the-art systems. Tianyu Liu 0001, Kexiang Wang, Baobao Chang, Zhifang Sui |
EMNLP | 4 |
| 2017 | Affinity-Preserving Random Walk for Multi-Document SummarizationabstractMulti-document summarization provides users with a short text that summarizes the information in a set of related documents.This paper introduces affinitypreserving random walk to the summarization task, which preserves the affinity relations of sentences by an absorbing random walk model.Meanwhile, we put forward adjustable affinity-preserving random walk to enforce the diversity constraint of summarization in the random walk process.The ROUGE evaluations on DUC 2003 topic-focused summarization task and DUC 2004 generic summarization task show the good performance of our method, which has the best ROUGE-2 recall among the graph-based ranking methods. Kexiang Wang, Tianyu Liu 0001, Zhifang Sui, Baobao Chang |
EMNLP | 3 |
| 2017 | Large-Scale Simple Question Generation by Template-Based Seq2seq Learning
Tianyu Liu 0001, Bingzhen Wei, Baobao Chang, Zhifang Sui |
NLPCC | 4 |
| 2017 | Will Repeated Reading Benefit Natural Language Understanding?
Lei Sha, Zhifang Sui |
NLPCC | 3 |
| 2016 | Implicit Discourse Relation Classification via Multi-Task Neural NetworksabstractWithout discourse connectives, classifying implicit discourse relations is a challenging task and a bottleneck for building a practical discourse parser. Previous research usually makes use of one kind of discourse framework such as PDTB or RST to improve the classification performance on discourse relations. Actually, under different discourse annotation frameworks, there exist multiple corpora which have internal connections. To exploit the combination of different discourse corpora, we design related discourse classification tasks specific to a corpus, and propose a novel Convolutional Neural Network embedded multi-task learning system to synthesize these tasks by learning both unique and shared representations for each task. The experimental results on the PDTB implicit discourse relation classification task demonstrate that our model achieves significant gains over baseline systems. Yang Liu 0124, Sujian Li, Xiaodong Zhang 0022, Zhifang Sui |
AAAI | 4 |
| 2016 | RBPB: Regularization-Based Pattern Balancing Method for Event ExtractionabstractEvent extraction is a particularly challenging information extraction task, which intends to identify and classify event triggers and arguments from raw text.In recent works, when determining event types (trigger classification), most of the works are either pattern-only or feature-only.However, although patterns cannot cover all representations of an event, it is still a very important feature.In addition, when identifying and classifying arguments, previous works consider each candidate argument separately while ignoring the relationship between arguments.This paper proposes a Regularization-Based Pattern Balancing Method (RBPB).Inspired by the progress in representation learning, we use trigger embedding, sentence-level embedding and pattern features together as our features for trigger classification so that the effect of patterns and other useful features can be balanced.In addition, RBPB uses a regularization method to take advantage of the relationship between arguments.Experiments show that we achieve results better than current state-of-art equivalents. Lei Sha, Jing Liu 0022, Chin-Yew Lin, Sujian Li, Baobao Chang, Zhifang Sui |
ACL (1) | 6 |
| 2016 | Event Detection with Burst Information NetworksabstractRetrospective event detection is an important task for discovering previously unidentified events in a text stream. In this paper, we propose two fast centroid-aware event detection models based on a novel text stream representation – Burst Information Networks (BINets) for addressing the challenge. The BINets are time-aware, efficient and can be easily analyzed for identifying key information (centroids). These advantages allow the BINet-based approaches to achieve the state-of-the-art performance on multiple datasets, demonstrating the efficacy of BINets for the task of event detection. Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Ming Zhou 0001 |
COLING | 4 |
| 2016 | Towards Time-Aware Knowledge Graph CompletionabstractKnowledge graph (KG) completion adds new facts to a KG by making inferences from existing facts. Most existing methods ignore the time information and only learn from time-unknown fact triples. In dynamic environments that evolve over time, it is important and challenging for knowledge graph completion models to take into account the temporal aspects of facts. In this paper, we present a novel time-aware knowledge graph completion model that is able to predict links in a KG using both the existing facts and the temporal information of the facts. To incorporate the happening time of facts, we propose a time-aware KG embedding model using temporal order information among facts. To incorporate the valid time of facts, we propose a joint time-aware inference model based on Integer Linear Programming (ILP) using temporal consistencyinformationasconstraints. Wefurtherintegratetwomodelstomakefulluseofglobal temporal information. We empirically evaluate our models on time-aware KG completion task. Experimental results show that our time-aware models achieve the state-of-the-art on temporal facts consistently. Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Baobao Chang, Sujian Li, Zhifang Sui |
COLING | 7 |
| 2016 | Reading and Thinking: Re-read LSTM Unit for Textual Entailment RecognitionabstractRecognizing Textual Entailment (RTE) is a fundamentally important task in natural language processing that has many applications. The recently released Stanford Natural Language Inference (SNLI) corpus has made it possible to develop and evaluate deep neural network methods for the RTE task. Previous neural network based methods usually try to encode the two sentences (premise and hypothesis) and send them together into a multi-layer perceptron to get their entailment type, or use LSTM-RNN to link two sentences together while using attention mechanic to enhance the model’s ability. In this paper, we propose to use the re-read mechanic, which means to read the premise again and again while reading the hypothesis. After read the premise again, the model can get a better understanding of the premise, which can also affect the understanding of the hypothesis. On the contrary, a better understanding of the hypothesis can also affect the understanding of the premise. With the alternative re-read process, the model can “think” of a better decision of entailment type. We designed a new LSTM unit called re-read LSTM (rLSTM) to implement this “thinking” process. Experiments show that we achieve results better than current state-of-the-art equivalents. Lei Sha, Baobao Chang, Zhifang Sui, Sujian Li |
COLING | 3 |
| 2016 | News Stream Summarization using Burst Information NetworksabstractThis paper studies summarizing key information from news streams. We propose simple yet effective models to solve the problem based on a novel and promising representation of text streams – Burst Information Networks (BINets). A BINet can be aware of redundant information, allows global analysis of a text stream, and can be efficiently built and dynamically updated, which perfectly fits the demands of text stream summarization. Extensive experiments show that the BINet-based approaches are not only efficient and can be used in a real-time online summarization setting, but also can generate high-quality summaries, outperforming the state-of-the-art approach. Tao Ge 0001, Lei Cui 0001, Baobao Chang, Sujian Li, Ming Zhou 0001, Zhifang Sui |
EMNLP | 6 |
| 2016 | Encoding Temporal Information for Time-Aware Link PredictionabstractMost existing knowledge base (KB) embedding methods solely learn from time-unknown fact triples but neglect the temporal information in the knowledge base.In this paper, we propose a novel time-aware KB embedding approach taking advantage of the happening time of facts.Specifically, we use temporal order constraints to model transformation between time-sensitive relations and enforce the embeddings to be temporally consistent and more accurate.We empirically evaluate our approach in two tasks of link prediction and triple classification.Experimental results show that our method outperforms other baselines on the two tasks consistently. Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui |
EMNLP | 7 |
| 2016 | Capturing Argument Relationship for Chinese Semantic Role LabelingabstractIn this paper, we capture the argument relationships for Chinese semantic role labeling task, and improve the task's performance with the help of argument relationships.We split the relationship between two candidate arguments into two categories: (1) Compatible arguments: if one candidate argument belongs to a given predicate, then the other is more likely to belong to the same predicate; (2) Incompatible arguments: if one candidate argument belongs to a given predicate, then the other is less likely to belong to the same predicate.However, previous works did not explicitly model argument relationships.We use a simple maximum entropy classifier to capture the two categories of argument relationships and test its performance on the Chinese Proposition Bank (CPB).The experiments show that argument relationships is effective in Chinese semantic role labeling task. Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui, Tingsong Jiang |
EMNLP | 4 |
| 2016 | Joint Learning Templates and Slots for Event Schema InductionabstractAutomatic event schema induction (AESI) means to extract meta-event from raw text, in other words, to find out what types (templates) of event may exist in the raw text and what roles (slots) may exist in each event type.In this paper, we propose a joint entity-driven model to learn templates and slots simultaneously based on the constraints of templates and slots in the same sentence.In addition, the entities' semantic information is also considered for the inner connectivity of the entities.We borrow the normalized cut criteria in image segmentation to divide the entities into more accurate template clusters and slot clusters.The experiment shows that our model gains a relatively higher result than previous work. Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui |
HLT-NAACL | 4 |
| 2015 | Bring you to the past: Automatic Generation of Topically Relevant Event ChroniclesabstractTao Ge, Wenzhe Pei, Heng Ji, Sujian Li, Baobao Chang, Zhifang Sui. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Tao Ge 0001, Wenzhe Pei, Heng Ji 0001, Sujian Li, Baobao Chang, Zhifang Sui |
ACL (1) | 6 |
| 2015 | Distinguishing Specific and Daily Topics
Tao Ge 0001, Wenzhe Pei, Baobao Chang, Zhifang Sui |
APWeb | 4 |
| 2015 | Recognizing Textual Entailment Using Probabilistic InferenceabstractRecognizing Text Entailment (RTE) plays an important role in NLP applications including question answering, information retrieval, etc.In recent work, some research explore "deep" expressions such as discourse commitments or strict logic for representing the text.However, these expressions suffer from the limitation of inference inconvenience or translation loss.To overcome the limitations, in this paper, we propose to use the predicate-argument structures to represent the discourse commitments extracted from text.At the same time, with the help of the YAGO knowledge, we borrow the distant supervision technique to mine the implicit facts from the text.We also construct a probabilistic network for all the facts and conduct inference to judge the confidence of each fact for RTE.The experimental results show that our proposed method achieves a competitive result compared to the previous work. Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui, Tingsong Jiang |
EMNLP | 4 |
| 2015 | Chinese Semantic Role Labeling with Bidirectional Recurrent Neural NetworksabstractTraditional approaches to Chinese Seman-tic Role Labeling (SRL) almost heavily re-ly on feature engineering. Even worse, the long-range dependencies in a sentence can hardly be modeled by these method-s. In this paper, we introduce bidirection-al recurrent neural network (RNN) with long-short-term memory (LSTM) to cap-ture bidirectional and long-range depen-dencies in a sentence with minimal fea-ture engineering. Experimental results on Chinese Proposition Bank (CPB) show a significant improvement over the state-of-the-art methods. Moreover, our model makes it convenient to introduce hetero-geneous resource, which makes a further improvement on our experimental perfor-mance. 1 Tingsong Jiang, Baobao Chang, Zhifang Sui |
EMNLP | 4 |
| 2015 | ERSOM: A Structural Ontology Matching Approach Using Automatically Learned Entity RepresentationabstractAs a key representation model of knowledge, ontology has been widely used in a lot of NLP related tasks, such as semantic parsing, information extraction and text mining etc.In this paper, we study the task of ontology matching, which concentrates on finding semantically related entities between different ontologies that describe the same domain, to solve the semantic heterogeneity problem.Previous works exploit different kinds of descriptions of an entity in ontology directly and separately to find the correspondences without considering the higher level correlations between the descriptions.Besides, the structural information of ontology haven't been utilized adequately for ontology matching.We propose in this paper an ontology matching approach, named ERSOM, which mainly includes an unsupervised representation learning method based on the deep neural networks to learn the general representation of the entities and an iterative similarity propagation method that takes advantage of more abundant structure information of the ontology to discover more mappings.The experimental results on the datasets from Ontology Alignment Evaluation Initiative (OAEI 1 ) show that ER-SOM achieves a competitive performance compared to the state-of-the-art ontology matching systems. Chuncheng Xiang, Tingsong Jiang, Baobao Chang, Zhifang Sui |
EMNLP | 4 |
| 2015 | An Ontology Matching Approach Based on Affinity-Preserving Random Walks
Chuncheng Xiang, Baobao Chang, Zhifang Sui |
IJCAI | 3 |
| 2014 | Event Schema Induction Based on Relational Co-occurrence over Multiple Documents
Tingsong Jiang, Lei Sha, Zhifang Sui |
NLPCC | 3 |
| 2013 | Exploiting collaborative filtering techniques for automatic assessment of student free-text responsesabstractThe automatic assessment of free-text responses of students is a relatively newer task in both computational linguistics and educational technology. The goal of the task is to produce an assessment of student answers to explanation and definition questions typically asked in problems seen in practice exercises or tests. Unlike some conventional methods which assess the student responses based on only information about their corresponding questions, this paper exploits idea of collaborative filtering to analyze student responses and used an effective collaborative filtering model -- feature-based matrix factorization model to deal with this challenge. The experimental results show that our feature-based matrix factorization model outperforms the baseline models and the model with a re-ranking phase can achieve a better and competitive performance -- 63.6% overall accuracy on the Beetle dataset. Tao Ge 0001, Zhifang Sui, Baobao Chang |
CIKM | 2 |
| 2013 | Event-Based Time Label Propagation for Automatic Dating of News ArticlesabstractSince many applications such as timeline summaries and temporal IR involving temporal analysis rely on document timestamps, the task of automatic dating of documents has been increasingly important.Instead of using feature-based methods as conventional models, our method attempts to date documents in a year level by exploiting relative temporal relations between documents and events, which are very effective for dating documents.Based on this intuition, we proposed an eventbased time label propagation model called confidence boosting in which time label information can be propagated between documents and events on a bipartite graph.The experiments show that our event-based propagation model can predict document timestamps in high accuracy and the model combined with a MaxEnt classifier outperforms the state-ofthe-art method for this task especially when the size of the training set is small. Tao Ge 0001, Baobao Chang, Sujian Li, Zhifang Sui |
EMNLP | 4 |
| 2009 | Chinese Semantic Role Labeling with Shallow Parsing
Zhifang Sui |
EMNLP | 2 |
| 2009 | Chinese Function Tag Labeling
Zhifang Sui |
PACLIC | 2 |
| 2008 | Prediction of Maximal Projection for Semantic Role Labeling
Weiwei Sun 0007, Zhifang Sui |
COLING | 2 |
| 2008 | The Integration of Dependency Relation Classification and Semantic Role Labeling Using Bilayer Maximum Entropy Markov Models
Hongzhan Li, Zhifang Sui |
CoNLL | 3 |
| 2007 | A Collocation-Based WSD Model: RFR-SUM
Weiguang Qu, Zhifang Sui, Genlin Ji, Shiwen Yu, Junsheng Zhou |
IEA/AIE | 2 |
| 2006 | A Study on Terminology Extraction Based on Classified Corpora
Yi-Rong Chen, Qin Lu 0001, Wenjie Li 0002, Zhifang Sui, Luning Ji |
LREC | 4 |
| 2000 | An Information-Theory-Based Feature Type Analysis for the Modeling of Statistical ParsingabstractThe paper proposes an information-theory-based method for feature types analysis in probabilistic evaluation modelling for statistical parsing. The basic idea is that we use entropy and conditional entropy to measure whether a feature type grasps some of the information for syntactic structure prediction. Our experiment quantitatively analyzes several feature types' power for syntactic structure prediction and draws a series of interesting conclusions. Zhifang Sui, Jun Zhao 0001, Dekai Wu |
ACL | 1 |
| 1999 | An Information-Theoretic Empirical Analysis of Dependency-Based Feature Types for Word Prediction Models
Dekai Wu, Jun Zhao 0001, Zhifang Sui |
EMNLP | 3 |
| 1999 | An information-based method for selecting feature types for word prediction
Dekai Wu, Zhifang Sui, Jun Zhao 0001 |
EUROSPEECH | 2 |