EDBT 2026 Demo / reviewers in the wild / expert
Qun Liu 0001
dblp:75/4402-1
· DBLP profile ↗
231ranked-venue papers
6as first author
89since 2021 · last 2026
0000-0002-7000-1792ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 216 · 4 first-author · 84 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 19 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool learningabstractTool learning, which allows Large Language Models (LLMs) to leverage external tools for solving complex user tasks, has emerged as a promising avenue for extending model capabilities. However, existing approaches primarily focus on data synthesis for fine-tuning LLMs to invoke tools effectively, largely ignoring how to fully stimulate the potential of the model. In this paper, we propose ToolACE-R, a novel framework that includes both model-aware iterative training and adaptive refinement for tool learning. ToolACE-R features a model-aware iterative training procedure that progressively adjust training samples based on the model’s evolving capabilities to maximize its potential. Additionally, it incorporates self-refinement training corpus which emphasizes LLM's ability to iteratively refine their tool calls, optimizing performance without requiring external feedback. Furthermore, we introduce adaptive self-refinement for efficient test-time scaling, where the trained model can autonomously determine when to stop the process based on iterative self-refinement. We conduct extensive experiments across several benchmark datasets, showing that ToolACE-R achieves competitive performance compared to advanced LLMs. The performance can be further improved efficiently through adaptive self-refinement. These results highlight the effectiveness and generalizability of ToolACE-R, offering a promising direction for more efficient and scalable tool learning. Xingshan Zeng, Weiwen Liu, Xu Huang 0008, Zezhong Wang 0004, Lingzhi Wang 0001, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Qun Liu 0001 |
AAAI | 11 |
| 2026 | Analyzing how pre-trained language models capture factual knowledge using attribution methods
Shaobo Li 0004, Chengjie Sun, Bingquan Liu, Lifeng Shang, Zhenhua Dong, Zhenzhou Ji, Xin Jiang 0002, Qun Liu 0001 |
Knowl. Based Syst. | 9 |
| 2025 | A Comprehensive Evaluation on Event Reasoning of Large Language ModelsabstractEvent reasoning is a fundamental ability that underlies many applications. It requires event schema knowledge to perform global reasoning and needs to deal with the diversity of the inter-event relations and the reasoning paradigms. The extent to which LLMs excel in event reasoning across various relations and reasoning paradigms has not been thoroughly investigated. Additionally, it is still unclear whether LLMs utilize event knowledge in the same way humans do. To mitigate this disparity, we comprehensively evaluate the abilities of event reasoning of LLMs on different relations, paradigms, and levels of abstraction. We introduce a novel benchmark EV2 for EValuation of EVent reasoning. EV2 consists of two levels of evaluation on schema and instance and is comprehensive in relations and reasoning paradigms. We conduct extensive experiments on EV2. We find that 1) LLMs have abilities to accomplish event reasoning but their performances are far from satisfactory. 2) There are imbalances of event reasoning abilities on different relations and paradigms. 3) LLMs have event schema knowledge, however, they're not aligned with humans on how to utilize the knowledge. Based on these findings, we guide the LLMs in utilizing the event schema knowledge as memory leading to improvements in event reasoning. Zhengwei Tao, Zhi Jin 0001, Yifan Zhang 0004, Xiancai Chen, Haiyan Zhao 0001, Jia Li 0003, Bin Liang 0004, Chongyang Tao, Qun Liu 0001, Kam-Fai Wong |
AAAI | 9 |
| 2025 | Friends-MMC: A Dataset for Multi-modal Multi-party Conversation UnderstandingabstractMulti-modal multi-party conversation (MMC) is a less studied yet important topic of research due to that it well fits real-world scenarios and thus potentially has more widely-used applications. Compared with the traditional multi-modal conversations, MMC requires stronger character-centered understanding abilities as there are many interlocutors appearing in both the visual and textual context. To facilitate the study of this problem, we present Friends-MMC in this paper, an MMC dataset that contains 24,000+ unique utterances paired with video context. To explore the character-centered understanding of the dialogue, we also annotate the speaker of each utterance, the names and bounding bboxes of faces that appear in the video. Based on this Friends-MMC dataset, we further study two fundamental MMC tasks: conversation speaker identification and conversation response prediction, both of which have the multi-party nature with the video or image as visual context. For conversation speaker identification, we demonstrate the inefficiencies of existing methods such as pre-trained models, and propose a simple yet effective baseline method that leverages an optimization solver to utilize the context of two modalities to achieve better performance. For conversation response prediction, we fine-tune generative dialogue models on Friend-MMC, and analyze the benefits of speaker information. The code and dataset will be publicly available, and thus we call for more attention on modelling speaker information when understanding conversations. Yueqian Wang, Xiaojun Meng, Yuxuan Wang 0004, Jianxin Liang, Qun Liu 0001, Dongyan Zhao 0001 |
AAAI | 5 |
| 2025 | Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal VerificationabstractChain-of-Thought (CoT) prompting has become the de facto method to elicit reasoning capabilities from large language models (LLMs). However, to mitigate hallucinations in CoT that are notoriously difficult to detect, current methods such as process reward models (PRMs) or self-consistency operate as opaque boxes and do not provide checkable evidence for their judgments, possibly limiting their effectiveness. To address this issue, we draw inspiration from the idea that “the gold standard for supporting a mathematical claim is to provide a proof”. We propose a retrospective, step-aware formal verification framework Safe. Rather than assigning arbitrary scores, we strive to articulate mathematical claims in formal mathematical language Lean 4 at each reasoning step and provide formal proofs to identify hallucinations. We evaluate our framework Safe across multiple language models and various mathematical datasets, demonstrating a significant performance improvement while offering interpretable and verifiable evidence. We also propose FormalStep as a benchmark for step correctness theorem proving with 30,809 formal statements. To the best of our knowledge, our work represents the first endeavor to utilize formal mathematical language Lean 4 for verifying content generated by LLMs, aligning with the reason why formal mathematical languages were created in the first place: to provide a robust foundation for hallucination-prone human-written proofs. Chengwu Liu 0001, Ye Yuan 0016, Yichun Yin, Zaoyu Chen, Yasheng Wang, Lifeng Shang, Qun Liu 0001, Ming Zhang 0004 |
ACL (1) | 9 |
| 2025 | Mixture of insighTful Experts (MoTE): The Synergy of Reasoning Chains and Expert Mixtures in Self-AlignmentabstractZhili Liu, Yunhao Gou, Kai Chen, Lanqing Hong, Jiahui Gao, Fei Mi, Yu Zhang, Zhenguo Li, Xin Jiang, Qun Liu, James Kwok. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhili Liu, Yunhao Gou, Kai Chen 0023, Lanqing Hong, Jiahui Gao 0002, Fei Mi, Yu Zhang 0006, Zhenguo Li, Xin Jiang 0002, Qun Liu 0001, James T. Kwok |
ACL (1) | 10 |
| 2025 | Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editingabstractKaishuai Xu, Tiezheng Yu, Wenjun Hou, Yi Cheng, Chak Tou Leong, Liangyou Li, Xin Jiang, Lifeng Shang, Qun Liu, Wenjie Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Kaishuai Xu, Tiezheng Yu, Chak Tou Leong, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001, Wenjie Li 0002 |
ACL (1) | 9 |
| 2025 | Hierarchical Memory Organization for Wikipedia GenerationabstractEugene J. Yu, Dawei Zhu, Yifan Song, Xiangyu Wong, Jiebin Zhang, Wenxuan Shi, Xiaoguang Li, Qun Liu, Sujian Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Eugene J. Yu, Yifan Song 0002, Xiangyu Wong, Jiebin Zhang, Qun Liu 0001, Sujian Li |
ACL (1) | 8 |
| 2025 | WIKIGENBENCH: Exploring Full-length Wikipedia Generation under Real-World ScenarioabstractIt presents significant challenges to generate comprehensive and accurate Wikipedia articles for newly emerging events under real-world scenario. Existing attempts fall short either by focusing only on short snippets or by using metrics that are insufficient to evaluate real-world scenarios. In this paper, we construct WIKIGENBENCH, a new benchmark consisting of 1,320 entries, designed to align with real-world scenarios in both generation and evaluation. For generation, we explore a real-world scenario where structured, full-length Wikipedia articles with citations are generated for new events using input documents from web sources. For evaluation, we integrate systematic metrics and LLM-based metrics to assess the verifiability, organization, and other aspects aligned with real-world scenarios. Based on this benchmark, we conduct extensive experiments using various models within three commonly used frameworks: direct RAG, hierarchical structure-based RAG, and RAG with fine-tuned generation model. Experimental results show that hierarchical-based methods can generate more comprehensive content, while fine-tuned methods achieve better verifiability. However, even the best methods still show a significant gap compared to existing Wikipedia content, indicating that further research is necessary. Jiebin Zhang, Eugene J. Yu, Qinyu Chen, Chenhao Xiong, Han Qian, Mingbo Song, Weimin Xiong, Qun Liu 0001, Sujian Li |
COLING | 10 |
| 2025 | EMOVA: Empowering Language Models to See, Hear and Speak with Vivid EmotionsabstractGPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging for the open-source community. Existing vision-language models rely on external tools for speech processing, while speech-language models still suffer from limited or totally without vision-understanding capabilities. To address this gap, we propose the EMOVA (EMotionally Omni-present Voice Assistant), to enable Large Language Models with end-to-end speech abilities while maintaining the leading vision-language performance. With a semantic-acoustic disentangled speech tokenizer, we surprisingly notice that omni-modal alignment can further enhance vision-language and speech abilities compared with the bi-modal aligned counterparts. Moreover, a lightweight style module is introduced for the flexible speech style controls including emotions and pitches. For the first time, EMOVA achieves state-of-the-art performance on both the vision-language and speech benchmarks, and meanwhile, supporting omni-modal spoken dialogue with vivid emotions. Yunhao Gou, Runhui Huang, Zhili Liu, Daxin Tan, Chunwei Wang, Yihan Zeng, Dingdong Wang, Kun Xiang, Haoli Bai, Jianhua Han, Weike Jin, Nian Xie, James T. Kwok, Hengshuang Zhao, Xiaodan Liang, Dit-Yan Yeung, Zhenguo Li, Qun Liu 0001, Lanqing Hong, Lu Hou 0002 |
CVPR | 27 |
| 2025 | Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction TuningabstractYunhao Gou, Hansi Yang, Zhili Liu, Kai Chen, Yihan Zeng, Lanqing Hong, Zhenguo Li, Qun Liu, Bo Han, James Kwok, Yu Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yunhao Gou, Hansi Yang, Zhili Liu, Kai Chen 0023, Yihan Zeng, Lanqing Hong, Zhenguo Li, Qun Liu 0001, Bo Han 0003, James T. Kwok, Yu Zhang 0006 |
EMNLP | 8 |
| 2025 | Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' ReasoningabstractZezhong Wang, Xingshan Zeng, Weiwen Liu, Yufei Wang, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Yufei Wang 0005, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong |
EMNLP | 9 |
| 2025 | ToolACE: Winning the Points of LLM Function CallingabstractFunction calling significantly extends the application boundary of large language models (LLMs), where high-quality and diverse training data is critical for unlocking this capability. However, collecting and annotating real function-calling data is challenging, while synthetic data from existing pipelines often lack coverage and accuracy. In this paper, we present ToolACE, an automatic agentic pipeline designed to generate accurate, complex, and diverse tool-learning data, specifically tailored to the capabilities of LLMs. ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs. Dialogs are further generated through the interplay among multiple agents, under the guidance of a complexity evaluator. To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks. We demonstrate that models trained on our synthesized data---even with only 8B parameters---achieve state-of-the-art performance, comparable to the latest GPT-4 models. Our model and a subset of the data are publicly available at https://huggingface.co/Team-ACE. Weiwen Liu, Xu Huang 0008, Xingshan Zeng, Xinlong Hao, Dexun Li, Shuai Wang 0020, Weinan Gan, Zhengying Liu, Yuanqing Yu, Zezhong Wang 0004, Yuxian Wang, Wu Ning, Yutai Hou, Bin Wang 0004, Chuhan Wu, Yong Liu 0020, Yasheng Wang, Duyu Tang, Dandan Tu, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Defu Lian, Qun Liu 0001, Enhong Chen |
ICLR | 26 |
| 2025 | ReAttention: Training-Free Infinite Context with Finite Attention ScopeabstractThe long-context capability of the Large Language Models (LLM) has made significant breakthroughs, but \textit{the maximum supported context length in length extrapolation} remains a critical bottleneck limiting their practical applications. The constraint of context length in LLMs arises from the self-attention mechanism, which cannot effectively and efficiently capture the semantic relationships within infinitely long contexts via the limited pre-trained positional information and attention scope. In this work, we propose \textbf{ReAttention}, a training-free approach enabling LLM based on the self-attention mechanism to support an infinite context with a finite attention scope under sufficient memory resources. ReAttention performs the position-agnostic top-$k$ attention before the ordinary position-aware self-attention, freeing LLMs from the length extrapolation issue. We validate the performance of ReAttention on the LongBench, L-Eval, and InfiniteBench and demonstrate that it is on par with traditional methods. Furthermore, we also apply ReAttention on mainstream LLMs, including LLaMA3.1-8B and Mistral-v0.3-7B, enabling them to support context lengths of at least 1M and even expanding the context length of LLaMA3.2-3B-chat by 128$\times$ to 4M without any further training in Needle-In-A-Haystack tests. We also improve the efficiency of ReAttention with Triton and achieve an efficient extrapolation without additional overhead. The code is available at \url{https://github.com/OpenMOSS/ReAttention}. Ruixiao Li, Zhigeng Liu, Qipeng Guo, Yuerong Song, Kai Lv 0001, Hang Yan 0001, Linlin Li 0001, Qun Liu 0001, Xipeng Qiu |
ICLR | 9 |
| 2025 | ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue SynthesisabstractZezhong Wang, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong |
NAACL (Long Papers) | 8 |
| 2025 | RidgeLoRA: Matrix Ridge Enhanced Low-Rank Adaptation of Large Language ModelsabstractAs one of the state-of-the-art parameter-efficient fine-tuning~(PEFT) methods, Low-Rank Adaptation (LoRA) enables model optimization with reduced computational cost through trainable low-rank matrix. However, the low-rank nature makes it prone to produce a decrease in the representation ability, leading to suboptimal performance. In order to break this limitation, we propose RidgeLoRA, a lightweight architecture like LoRA that incorporates novel architecture and matrix ridge enhanced full-rank approximation, to match the performance of full-rank training, while eliminating the need for high memory and a large number of parameters to restore the rank of matrices. We provide a rigorous mathematical derivation to prove that RidgeLoRA has a better upper bound on the representations than vanilla LoRA. Furthermore, extensive experiments across multiple domains demonstrate that RidgeLoRA achieves better performance than other LoRA variants, and can even match or surpass full-rank training. Junda Zhu 0003, Jun Ai, Yichun Yin, Yasheng Wang, Lifeng Shang, Qun Liu 0001 |
NeurIPS | 7 |
| 2025 | End-to-End VideoQA with Frame Scoring Mechanisms and Adaptive Sampling
Jianxin Liang, Xiaojun Meng, Yueqian Wang, Chang Liu 0076, Qun Liu 0001, Dongyan Zhao 0001 |
NLPCC (2) | 5 |
| 2025 | Unsupervised Domain Adaptive Visual Question Answering in the Era of Multi-Modal Large Language ModelsabstractUnsupervised domain adaptation (UDA) for visual question answering (VQA) has attracted research interest. However, with Multi-modal Large Language Models (MLLMs) showing great performance on VQA datasets, UDA for VQA based on MLLMs remains unexplored. To fill this gap, we propose the first systematic approach to Unsupervised Domain Adaptation VQA based on MLLMs (UDAM). First, we introduce semantic context feature alignment and domain query feature alignment, which utilize a single token embedding for each modality to capture contextual domain information from unimodal inputs and conduct coarsegrained feature alignment on it, thus alleviating domain shifts in the unimodal feature space. Second, we propose the novel semantics-guided query feature alignment, which differentiates important domain-specific queries from learnable query outputs and conducts fine-grained feature alignment controlled by a semantics-guided weight map to reduce domain shifts in the cross-modal feature space. Third, we devise a pair-wise domain-aware prompt strategy, which aids UDA by prompting MLLMs to discern the commonality of tasks and the distinctiveness of domains in multi-modal inputs. Extensive experiments demonstrate UDAM's effectiveness in adapting MLLMs to unlabeled new domains. Weixi Weng, Xiaojun Meng, Jieming Zhu, Qun Liu 0001, Chun Yuan 0003 |
WACV | 5 |
| 2025 | Enhancing inter-sentence coherence of extractive summarization with multitask learning
Renlong Jie, Xiaojun Meng, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
J. Intell. Inf. Syst. | 5 |
| 2024 | Preparing Lessons for Progressive Training on Language ModelsabstractThe rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small models to improve training efficiency, but this approach may not be suitable for new model structures. On the other hand, training from scratch can be slow, and progressively stacking layers often fails to achieve significant acceleration. To address these challenges, we propose a novel method called Apollo, which prepares lessons for expanding operations by learning high-layer functionality during training of low layers. Our approach involves low-value-prioritized sampling (LVPS) to train different depths and weight sharing to facilitate efficient expansion. We also introduce an interpolation method for stable model depth extension. Experiments demonstrate that Apollo achieves state-of-the-art acceleration ratios, even rivaling methods using pretrained models, making it a universal and efficient solution for training deep models while reducing time, financial, and environmental costs. Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Jiaxin Shi, Zenglin Xu, Ming Zhang 0004, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
AAAI | 9 |
| 2024 | Unsupervised Extractive Summarization with Learnable Length Control StrategiesabstractUnsupervised extractive summarization is an important technique in information extraction and retrieval. Compared with supervised method, it does not require high-quality human-labelled summaries for training and thus can be easily applied for documents with different types, domains or languages. Most of existing unsupervised methods including TextRank and PACSUM rely on graph-based ranking on sentence centrality. However, this scorer can not be directly applied in end-to-end training, and the positional-related prior assumption is often needed for achieving good summaries. In addition, less attention is paid to length-controllable extractor, where users can decide to summarize texts under particular length constraint. This paper introduces an unsupervised extractive summarization model based on a siamese network, for which we develop a trainable bidirectional prediction objective between the selected summary and the original document. Different from the centrality-based ranking methods, our extractive scorer can be trained in an end-to-end manner, with no other requirement of positional assumption. In addition, we introduce a differentiable length control module by approximating 0-1 knapsack solver for end-to-end length-controllable extracting. Experiments show that our unsupervised method largely outperforms the centrality-based baseline using a same sentence encoder. In terms of length control ability, via our trainable knapsack module, the performance consistently outperforms the strong baseline without utilizing end-to-end training. Human evaluation further evidences that our method performs the best among baselines in terms of relevance and consistency. Renlong Jie, Xiaojun Meng, Xin Jiang 0002, Qun Liu 0001 |
AAAI | 4 |
| 2024 | EWEK-QA : Enhanced Web and Efficient Knowledge Graph Retrieval for Citation-based Question Answering SystemsabstractMohammad Dehghan, Mohammad Alomrani, Sunyam Bagga, David Alfonso-Hermelo, Khalil Bibi, Abbas Ghaddar, Yingxue Zhang, Xiaoguang Li, Jianye Hao, Qun Liu, Jimmy Lin, Boxing Chen, Prasanna Parthasarathi, Mahdi Biparva, Mehdi Rezagholizadeh. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mohammad Dehghan, Mohammad Ali Alomrani, Sunyam Bagga, David Alfonso-Hermelo, Khalil Bibi, Abbas Ghaddar, Yingxue Zhang 0001, Jianye Hao, Qun Liu 0001, Jimmy Lin, Boxing Chen, Prasanna Parthasarathi, Mahdi Biparva, Mehdi Rezagholizadeh |
ACL (1) | 10 |
| 2024 | FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language ModelsabstractYuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yufei Wang 0005, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Wei Wang 0011 |
ACL (1) | 9 |
| 2024 | Learning to Edit: Aligning LLMs with Knowledge EditingabstractYuxin Jiang, Yufei Wang, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao, Liangyou Li, Xin Jiang, Lifeng Shang, Ruiming Tang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yufei Wang 0005, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao 0002, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Qun Liu 0001, Wei Wang 0011 |
ACL (1) | 11 |
| 2024 | M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language ModelsabstractWai-Chung Kwan, Xingshan Zeng, Yufei Wang, Yusen Sun, Liangyou Li, Yuxin Jiang, Lifeng Shang, Qun Liu, Kam-Fai Wong. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wai-Chung Kwan, Xingshan Zeng, Yufei Wang 0005, Yusen Sun, Liangyou Li, Lifeng Shang, Qun Liu 0001, Kam-Fai Wong |
ACL (1) | 8 |
| 2024 | ProxyQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language ModelsabstractHaochen Tan, Zhijiang Guo, Zhan Shi, Lu Xu, Zhili Liu, Yunlong Feng, Xiaoguang Li, Yasheng Wang, Lifeng Shang, Qun Liu, Linqi Song. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Haochen Tan, Zhijiang Guo, Zhan Shi 0001, Zhili Liu, Yunlong Feng, Yasheng Wang, Lifeng Shang, Qun Liu 0001, Linqi Song |
ACL (1) | 10 |
| 2024 | LFED: A Literary Fiction Evaluation Dataset for Large Language ModelsabstractThe rapid evolution of large language models (LLMs) has ushered in the need for comprehensive assessments of their performance across various dimensions. In this paper, we propose LFED, a Literary Fiction Evaluation Dataset, which aims to evaluate the capability of LLMs on the long fiction comprehension and reasoning. We collect 95 literary fictions that are either originally written in Chinese or translated into Chinese, covering a wide range of topics across several centuries. We define a question taxonomy with 8 question categories to guide the creation of 1,304 questions. Additionally, we conduct an in-depth analysis to ascertain how specific attributes of literary fictions (e.g., novel types, character numbers, the year of publication) impact LLM performance in evaluations. Through a series of experiments involving various state-of-the-art LLMs, our findings reveal that these models face considerable challenges in effectively addressing questions related to literary fictions, with ChatGPT reaching only 57.08% under the zero-shot setting. The dataset will be publicly available at https://github.com/tjunlp-lab/LFED.git. Linhao Yu, Qun Liu 0001, Deyi Xiong |
LREC/COLING | 2 |
| 2024 | MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language ModelsabstractWai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Wai-Chung Kwan, Xingshan Zeng, Yufei Wang 0005, Liangyou Li, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong |
EMNLP | 8 |
| 2024 | CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational SearchabstractIn this paper, we study how open-source large language models (LLMs) can be effectively deployed for improving query rewriting in conversational search, especially for ambiguous queries.We introduce CHIQ, a two-step method that leverages the capabilities of LLMs to resolve ambiguities in the conversation history before query rewriting.This approach contrasts with prior studies that predominantly use closed-source LLMs to directly generate search queries from conversation history.We demonstrate on five well-established benchmarks that CHIQ leads to state-of-the-art results across most settings, showing highly competitive performances with systems leveraging closed-source LLMs.Our study provides a first step towards leveraging open-source LLMs in conversational search, as a competitive alternative to the prevailing reliance on commercial LLMs for query rewriting. Fengran Mo, Abbas Ghaddar, Kelong Mao, Mehdi Rezagholizadeh, Boxing Chen, Qun Liu 0001, Jian-Yun Nie |
EMNLP | 6 |
| 2024 | Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental ChunkabstractZhiyuan Zeng, Qipeng Guo, Xiaoran Liu, Zhangyue Yin, Wentao Shu, Mianqiu Huang, Bo Wang, Yunhua Zhou, Linlin Li, Qun Liu, Xipeng Qiu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhiyuan Zeng 0004, Qipeng Guo, Zhangyue Yin, Wentao Shu, Mianqiu Huang, Bo Wang 0084, Yunhua Zhou, Linlin Li 0001, Qun Liu 0001, Xipeng Qiu |
EMNLP | 10 |
| 2024 | Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake AnalysisabstractThe rapid development of large language models (LLMs) has not only provided numerous opportunities but also presented significant challenges. This becomes particularly evident when LLMs inadvertently generate harmful or toxic content, either unintentionally or because of intentional inducement. Existing alignment methods usually direct LLMs toward the favorable outcomes by utilizing human-annotated, flawless instruction-response pairs. Conversely, this study proposes a novel alignment technique based on mistake analysis, which deliberately exposes LLMs to erroneous content to learn the reasons for mistakes and how to avoid them. In this case, mistakes are repurposed into valuable data for alignment, effectively helping to avoid the production of erroneous responses. Without external models or human annotations, our method leverages a model's intrinsic ability to discern undesirable mistakes and improves the safety of its generated responses. Experimental results reveal that our method outperforms existing alignment approaches in enhancing model safety while maintaining the overall utility. Kai Chen 0023, Chunwei Wang, Jianhua Han, Lanqing Hong, Fei Mi, Hang Xu 0004, Zhengying Liu, Wenyong Huang, Zhenguo Li, Dit-Yan Yeung, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
ICLR | 14 |
| 2024 | Retrieval-based Disentangled Representation Learning with Natural Language SupervisionabstractDisentangled representation learning remains challenging as the underlying factors of variation in the data do not naturally exist. The inherent complexity of real-world data makes it unfeasible to exhaustively enumerate and encapsulate all its variations within a finite set of factors. However, it is worth noting that most real-world data have linguistic equivalents, typically in the form of textual descriptions. These linguistic counterparts can represent the data and effortlessly decomposed into distinct tokens. In light of this, we present Vocabulary Disentangled Retrieval (VDR), a retrieval-based framework that harnesses natural language as proxies of the underlying data variation to drive disentangled representation learning. Our approach employ a bi-encoder model to represent both data and natural language in a vocabulary space, enabling the model to distinguish dimensions that capture intrinsic characteristics within data through its natural language counterpart, thus facilitating disentanglement. We extensively assess the performance of VDR across 15 retrieval benchmark datasets, covering text-to-text and cross-modal retrieval scenarios, as well as human evaluation. Our experimental results compellingly demonstrate the superiority of VDR over previous bi-encoder retrievers with comparable model size and training costs, achieving an impressive 8.7% improvement in NDCG@10 on the BEIR benchmark, a 5.3\% increase on MS COCO, and a 6.0% increase on Flickr30k in terms of mean recall in the zero-shot setting. Moreover, The results from human evaluation indicate that interpretability of our method is on par with SOTA captioning models. Jiawei Zhou 0003, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Lei Chen 0002 |
ICLR | 5 |
| 2024 | Visually Guided Generative Text-Layout Pre-training for Document IntelligenceabstractZhiming Mao, Haoli Bai, Lu Hou, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhiming Mao, Haoli Bai, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong |
NAACL-HLT | 6 |
| 2024 | ShortcutLens: A Visual Analytics Approach for Exploring Shortcuts in Natural Language Understanding DatasetabstractBenchmark datasets play an important role in evaluating Natural Language Understanding (NLU) models. However, shortcuts-unwanted biases in the benchmark datasets-can damage the effectiveness of benchmark datasets in revealing models' real capabilities. Since shortcuts vary in coverage, productivity, and semantic meaning, it is challenging for NLU experts to systematically understand and avoid them when creating benchmark datasets. In this paper, we develop a visual analytics system, ShortcutLens, to help NLU experts explore shortcuts in NLU benchmark datasets. The system allows users to conduct multi-level exploration of shortcuts. Specifically, Statistics View helps users grasp the statistics such as coverage and productivity of shortcuts in the benchmark dataset. Template View employs hierarchical and interpretable templates to summarize different types of shortcuts. Instance View allows users to check the corresponding instances covered by the shortcuts. We conduct case studies and expert interviews to evaluate the effectiveness and usability of the system. The results demonstrate that ShortcutLens supports users in gaining a better understanding of benchmark dataset issues through shortcuts, inspiring them to create challenging and pertinent benchmark datasets. Zhihua Jin, Xingbo Wang 0001, Furui Cheng, Chunhui Sun, Qun Liu 0001, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | KPT: Keyword-Guided Pre-training for Grounded Dialog GenerationabstractIncorporating external knowledge into the response generation process is essential to building more helpful and reliable dialog agents. However, collecting knowledge-grounded conversations is often costly, calling for a better pre-trained model for grounded dialog generation that generalizes well w.r.t. different types of knowledge. In this work, we propose KPT (Keyword-guided Pre-Training), a novel self-supervised pre-training method for grounded dialog generation without relying on extra knowledge annotation. Specifically, we use a pre-trained language model to extract the most uncertain tokens in the dialog as keywords. With these keywords, we construct two kinds of knowledge and pre-train a knowledge-grounded response generation model, aiming at handling two different scenarios: (1) the knowledge should be faithfully grounded; (2) it can be selectively used. For the former, the grounding knowledge consists of keywords extracted from the response. For the latter, the grounding knowledge is additionally augmented with keywords extracted from other utterances in the same dialog. Since the knowledge is extracted from the dialog itself, KPT can be easily performed on a large volume and variety of dialogue data. We considered three data sources (open-domain, task-oriented, conversational QA) with a total of 2.5M dialogues. We conduct extensive experiments on various few-shot knowledge-grounded generation tasks, including grounding on dialog acts, knowledge graphs, persona descriptions, and Wikipedia passages. Our comprehensive experiments and analyses demonstrate that KPT consistently outperforms state-of-the-art methods on these tasks with diverse grounding knowledge. Qi Zhu 0007, Fei Mi, Zheng Zhang 0020, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Xiaoyan Zhu 0001, Minlie Huang |
AAAI | 7 |
| 2023 | MoralDial: A Framework to Train and Evaluate Moral Dialogue Systems via Moral DiscussionsabstractHao Sun, Zhexin Zhang, Fei Mi, Yasheng Wang, Wei Liu, Jianwei Cui, Bin Wang, Qun Liu, Minlie Huang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hao Sun 0012, Zhexin Zhang, Fei Mi, Yasheng Wang, Wei Liu 0005, Jianwei Cui 0002, Bin Wang 0004, Qun Liu 0001, Minlie Huang |
ACL (1) | 8 |
| 2023 | Wukong-Reader: Multi-modal Pre-training for Fine-grained Visual Document UnderstandingabstractHaoli Bai, Zhiguang Liu, Xiaojun Meng, Li Wentao, Shuang Liu, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang, Lu Hou, Jiansheng Wei, Xin Jiang, Qun Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Haoli Bai, Xiaojun Meng, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang 0004, Lu Hou 0002, Jiansheng Wei, Xin Jiang 0002, Qun Liu 0001 |
ACL (1) | 13 |
| 2023 | mCLIP: Multilingual CLIP via Cross-lingual TransferabstractGuanhua Chen, Lu Hou, Yun Chen, Wenliang Dai, Lifeng Shang, Xin Jiang, Qun Liu, Jia Pan, Wenping Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Guanhua Chen 0001, Lu Hou 0002, Yun Chen 0007, Wenliang Dai, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Jia Pan 0001, Wenping Wang 0001 |
ACL (1) | 7 |
| 2023 | DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question AnsweringabstractExisting evaluation metrics for natural language generation (NLG) tasks face the challenges on generalization ability and interpretability.Specifically, most of the wellperformed metrics are required to train on evaluation datasets of specific NLG tasks and evaluation dimensions, which may cause over-fitting to task-specific datasets.Furthermore, existing metrics only provide an evaluation score for each dimension without revealing the evidence to interpret how this score is obtained.To deal with these challenges, we propose a simple yet effective metric called DecompEval.This metric formulates NLG evaluation as an instruction-style question answering task and utilizes instruction-tuned pre-trained language models (PLMs) without training on evaluation datasets, aiming to enhance the generalization ability.To make the evaluation process more interpretable, we decompose our devised instruction-style question about the quality of generated texts into the subquestions that measure the quality of each sentence.The subquestions with their answers generated by PLMs are then recomposed as evidence to obtain the evaluation result.Experimental results show that DecompEval achieves state-of-the-art performance in untrained metrics for evaluating text summarization and dialogue generation, which also exhibits strong dimension-level / task-level generalization ability and interpretability 1 . Pei Ke, Fei Huang 0005, Fei Mi, Yasheng Wang, Qun Liu 0001, Xiaoyan Zhu 0001, Minlie Huang |
ACL (1) | 5 |
| 2023 | One Cannot Stand for Everyone! Leveraging Multiple User Simulators to train Task-oriented Dialogue SystemsabstractYajiao Liu, Xin Jiang, Yichun Yin, Yasheng Wang, Fei Mi, Qun Liu, Xiang Wan, Benyou Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yajiao Liu, Xin Jiang 0002, Yichun Yin, Yasheng Wang, Fei Mi, Qun Liu 0001, Benyou Wang |
ACL (1) | 6 |
| 2023 | TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language ModelsabstractJing Xiong, Jianhao Shen, Ye Yuan, Haiming Wang, Yichun Yin, Zhengying Liu, Lin Li, Zhijiang Guo, Qingxing Cao, Yinya Huang, Chuanyang Zheng, Xiaodan Liang, Ming Zhang, Qun Liu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jianhao Shen, Ye Yuan 0016, Yichun Yin, Zhengying Liu, Zhijiang Guo, Qingxing Cao, Yinya Huang, Chuanyang Zheng, Xiaodan Liang, Ming Zhang 0004, Qun Liu 0001 |
EMNLP | 14 |
| 2023 | Lexicon-injected Semantic Parsing for Task-Oriented DialogabstractRecently, semantic parsing using hierarchical representations for dialog systems has captured substantial attention. Task-Oriented Parse (TOP), a tree representation with intents and slots as labels of nested tree nodes, has been proposed for parsing user utterances. Previous TOP parsing methods are limited on leveraging lexicon resources, which are often used to guide the real dialog system. To mitigate this issue, we first propose a novel span-splitting representation for span-based parser that outperforms existing methods. Then we present a novel lexicon-injected semantic parser, which collects slot labels of tree representation as a lexicon, and injects lexical features to the span representation of parser. An additional slot disambiguation technique is involved to remove inappropriate span match occurrences from the lexicon. Experiments show that our best parser produces a new state-of-the-art result (87.62%) on the TOP dataset, and also confirm the effectiveness of our proposed lexicon-injected parser and slot disambiguation model. Xiaojun Meng, Wenlin Dai, Yasheng Wang, Baojun Wang, Zhiyong Wu 0003, Xin Jiang 0002, Qun Liu 0001 |
ICASSP | 7 |
| 2023 | History, Present and Future: Enhancing Dialogue Generation with Few-Shot History-Future PromptabstractDialogue history and response in open-domain dialogue are loosely coupled. Generating informative responses solely based on the original dialogue history is not easy, as dialogue history may not contain enough information or it may contain irrelevant noises. Intuitively, if a generation model can foresee possible dialogue future, or obtain real useful histories, it could generate more informative responses. In this paper, we propose a novel lightweight dialogue generation framework named few-shot history-future prompt that utilizes useful histories and simulated futures to help generate informative responses, without the need for fine-tuning or adding extra parameters. To obtain useful histories, we retrieve and combine relevant utterances from noisy multi- turn histories. Then we adopt a retrieval-generation hybrid approach to obtain diversified simulated futures. Such that our model could learn to condition on history combinations and simulated futures via few-shot learning. Experiments over publicly available datasets demonstrate that our method can help models generate better responses. Yasheng Wang, Fei Mi, Pingyi Zhou, Jin Liu 0016, Xin Jiang 0002, Qun Liu 0001 |
ICASSP | 8 |
| 2023 | Learning Summary-Worthy Visual Representation for Abstractive Summarization in VideoabstractMultimodal abstractive summarization for videos (MAS) requires generating a concise textual summary to describe the highlights of a video according to multimodal resources, in our case, the video content and its transcript. Inspired by the success of the large-scale generative pre-trained language model (GPLM) in generating high-quality textual content (e.g., summary), recent MAS methods have proposed to adapt the GPLM to this task by equipping it with the visual information, which is often obtained through a general-purpose visual feature extractor. However, the generally extracted visual features may overlook some summary-worthy visual information, which impedes model performance. In this work, we propose a novel approach to learning the summary-worthy visual representation that facilitates abstractive summarization. Our method exploits the summary-worthy information from both the cross-modal transcript data and the knowledge that distills from the pseudo summary. Extensive experiments on three public multimodal datasets show that our method outperforms all competing baselines. Furthermore, with the advantages of summary-worthy visual information, our model can have a significant improvement on small datasets or even datasets with limited training data. Zenan Xu, Xiaojun Meng, Yasheng Wang, Qinliang Su, Zexuan Qiu, Xin Jiang 0002, Qun Liu 0001 |
IJCAI | 7 |
| 2023 | Reusing Pretrained Models by Multi-linear Operators for Efficient TrainingabstractTraining large models from scratch usually costs a substantial amount of resources. Towards this problem, recent studies such as bert2BERT and LiGO have reused small pretrained models to initialize a large model (termed the ``target model''), leading to a considerable acceleration in training. Despite the successes of these previous studies, they grew pretrained models by mapping partial weights only, ignoring potential correlations across the entire model. As we show in this paper, there are inter- and intra-interactions among the weights of both the pretrained and the target models. As a result, the partial mapping may not capture the complete information and lead to inadequate growth. In this paper, we propose a method that linearly correlates each weight of the target model to all the weights of the pretrained model to further enhance acceleration ability. We utilize multi-linear operators to reduce computational and spacial complexity, enabling acceptable resource requirements. Experiments demonstrate that our method can save 76\% computational costs on DeiT-base transferred from DeiT-small, which outperforms bert2BERT by +12\% and LiGO by +21\%, respectively. Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Zenglin Xu, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
NeurIPS | 7 |
| 2023 | MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse LanguagesabstractAbstract MIRACL is a multilingual dataset for ad hoc retrieval across 18 languages that collectively encompass over three billion native speakers around the world. This resource is designed to support monolingual retrieval tasks, where the queries and the corpora are in the same language. In total, we have gathered over 726k high-quality relevance judgments for 78k queries over Wikipedia in these languages, where all annotations have been performed by native speakers hired by our team. MIRACL covers languages that are both typologically close as well as distant from 10 language families and 13 sub-families, associated with varying amounts of publicly available resources. Extensive automatic heuristic verification and manual assessments were performed during the annotation process to control data quality. In total, MIRACL represents an investment of around five person-years of human annotator effort. Our goal is to spur research on improving retrieval across a continuum of languages, thus enhancing information access capabilities for diverse populations around the world, particularly those that have traditionally been underserved. MIRACL is available at http://miracl.ai/. Xinyu Zhang 0018, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Qun Liu 0001, Mehdi Rezagholizadeh, Jimmy Lin |
Trans. Assoc. Comput. Linguistics | 7 |
| 2023 | Sub-Character Tokenization for Chinese Pretrained Language ModelsabstractAbstract Tokenization is fundamental to pretrained language models (PLMs). Existing tokenization methods for Chinese PLMs typically treat each character as an indivisible token. However, they ignore the unique feature of the Chinese writing system where additional linguistic information exists below the character level, i.e., at the sub-character level. To utilize such information, we propose sub-character (SubChar for short) tokenization. Specifically, we first encode the input text by converting each Chinese character into a short sequence based on its glyph or pronunciation, and then construct the vocabulary based on the encoded text with sub-word segmentation. Experimental results show that SubChar tokenizers have two main advantages over existing tokenizers: 1) They can tokenize inputs into much shorter sequences, thus improving the computational efficiency. 2) Pronunciation-based SubChar tokenizers can encode Chinese homophones into the same transliteration sequences and produce the same tokenization output, hence being robust to homophone typos. At the same time, models trained with SubChar tokenizers perform competitively on downstream tasks. We release our code and models at https://github.com/thunlp/SubCharTokenization to facilitate future work. Chenglei Si, Zhengyan Zhang, Yingfa Chen, Fanchao Qi, Xiaozhi Wang, Zhiyuan Liu 0001, Yasheng Wang, Qun Liu 0001, Maosong Sun 0001 |
Trans. Assoc. Comput. Linguistics | 8 |
| 2022 | From Fully Trained to Fully Random Embeddings: Improving Neural Machine Translation with Compact Word Embedding TablesabstractEmbedding matrices are key components in neural natural language processing (NLP) models that are responsible to provide numerical representations of input tokens (i.e. words or subwords). In this paper, we analyze the impact and utility of such matrices in the context of neural machine translation (NMT). We show that detracting syntactic and semantic information from word embeddings and running NMT systems with random embeddings is not as damaging as it initially sounds. We also show how incorporating only a limited amount of task-specific knowledge from fully-trained embeddings can boost the performance NMT systems. Our findings demonstrate that in exchange for negligible deterioration in performance, any NMT model can be run with partially random embeddings. Working with such structures means a minimal memory requirement as there is no longer need to store large embedding tables, which is a significant gain in industrial and on-device settings. We evaluated our embeddings in translating English into German and French and achieved a 5.3x compression rate. Despite having a considerably smaller architecture, our models in some cases are even able to outperform state-of-the-art baselines. Krtin Kumar, Peyman Passban, Mehdi Rezagholizadeh, Yiu Sing Lau, Qun Liu 0001 |
AAAI | 5 |
| 2022 | UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationabstractWith the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and image modalities to output a pictorial summary with the most relevant images. Existing methods mostly focus on either extractive or abstractive summarization and rely on the presence and quality of image captions to build image references. We are the first to propose a Unified framework for Multimodal Summarization grounding on BART, UniMS, that integrates extractive and abstractive objectives, as well as selecting the image output. Specially, we adopt knowledge distillation from a vision-language pretrained model to improve image selection, which avoids any requirement on the existence and quality of image captions. Besides, we introduce a visual guided decoder to better integrate textual and visual modalities in guiding abstractive text generation. Results show that our best model achieves a new state-of-the-art result on a large-scale benchmark dataset. The newly involved extractive objective as well as the knowledge distillation technique are proven to bring a noticeable improvement to the multimodal summarization task. Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Zhenglu Yang |
AAAI | 5 |
| 2022 | bert2BERT: Towards Reusable Pretrained Language ModelsabstractCheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang, Yujia Qin, Fengyu Wang, Zhi Wang, Xiao Chen, Zhiyuan Liu, Qun Liu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yichun Yin, Lifeng Shang, Xin Jiang 0002, Yujia Qin, Zhi Wang 0001, Xiao Chen 0012, Zhiyuan Liu 0001, Qun Liu 0001 |
ACL (1) | 10 |
| 2022 | Achieving Reliable Human Assessment of Open-Domain Dialogue SystemsabstractEvaluation of open-domain dialogue systems is highly challenging and development of better techniques is highlighted time and again as desperately needed.Despite substantial efforts to carry out reliable live evaluation of systems in recent competitions, annotations have been abandoned and reported as too unreliable to yield sensible results.This is a serious problem since automatic metrics are not known to provide a good indication of what may or may not be a high-quality conversation.Answering the distress call of competitions that have emphasized the urgent need for better evaluation techniques in dialogue, we present the successful development of human evaluation that is highly reliable while still remaining feasible and low cost.Self-replication experiments reveal almost perfectly repeatable results with a correlation of r = 0.969.Furthermore, due to the lack of appropriate methods of statistical significance testing, the likelihood of potential improvements to systems occurring due to chance is rarely taken into account in dialogue evaluation, and the evaluation we propose facilitates application of standard tests.Since we have developed a highly reliable evaluation method, new insights into system performance can be revealed.We therefore include a comparison of state-of-the-art models (i) with and without personas, to measure the contribution of personas to conversation quality, as well as (ii) prescribed versus freely chosen topics.Interestingly with respect to personas, results indicate that personas do not positively contribute to conversation quality as expected. Tianbo Ji, Yvette Graham, Gareth J. F. Jones, Chenyang Lyu, Qun Liu 0001 |
ACL (1) | 5 |
| 2022 | Universal Conditional Masked Language Pre-training for Neural Machine TranslationabstractPre-trained sequence-to-sequence models have significantly improved Neural Machine Translation (NMT).Different from prior works where pre-trained models usually adopt an unidirectional decoder, this paper demonstrates that pre-training a sequenceto-sequence model but with a bidirectional decoder can produce notable performance gains for both Autoregressive and Nonautoregressive NMT.Specifically, we propose CeMAT, a conditional masked language model pre-trained on large-scale bilingual and monolingual corpora in many languages.1 We also introduce two simple but effective methods to enhance the CeMAT, aligned code-switching & masking and dynamic dual-masking.We conduct extensive experiments and show that our CeMAT can achieve significant performance improvement for all scenarios from low-to extremely highresource languages, i.e., up to +14.4 BLEU on low-resource and +7.9 BLEU on average for Autoregressive NMT.For Non-autoregressive NMT, we demonstrate it can also produce consistent performance gains, i.e., up to +5.3 BLEU.To the best of our knowledge, this is the first work to pre-train a unified model for fine-tuning on both NMT tasks. Liangyou Li, Meng Zhang 0019, Minghao Wu, Qun Liu 0001 |
ACL (1) | 5 |
| 2022 | Compression of Generative Pre-trained Language Models via QuantizationabstractChaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang, Xin Jiang, Qun Liu, Ping Luo, Ngai Wong. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Chaofan Tao, Lu Hou 0002, Wei Zhang 0196, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Ping Luo 0002, Ngai Wong 0001 |
ACL (1) | 6 |
| 2022 | ClusterFormer: Neural Clustering Attention for Efficient and Effective TransformerabstractNingning Wang, Guobing Gan, Peng Zhang, Shuai Zhang, Junqiu Wei, Qun Liu, Xin Jiang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Guobing Gan, Peng Zhang 0002, Victor Junqiu Wei, Qun Liu 0001, Xin Jiang 0002 |
ACL (1) | 6 |
| 2022 | Hyperlink-induced Pre-training for Passage Retrieval in Open-domain Question AnsweringabstractJiawei Zhou, Xiaoguang Li, Lifeng Shang, Lan Luo, Ke Zhan, Enrui Hu, Xinyu Zhang, Hao Jiang, Zhao Cao, Fan Yu, Xin Jiang, Qun Liu, Lei Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Jiawei Zhou 0003, Lifeng Shang, Ke Zhan, Enrui Hu, Xinyu Zhang 0019, Hao Jiang 0022, Zhao Cao, Fan Yu 0004, Xin Jiang 0002, Qun Liu 0001, Lei Chen 0002 |
ACL (1) | 12 |
| 2022 | CStory: A Chinese Large-scale News Storyline DatasetabstractIn today's massive news streams, storylines can help us discover related event pairs and understand the evolution of hot events. Hence many efforts have been devoted to automatically constructing news storylines. However, the development of these methods is strongly limited by the size and quality of existing storyline datasets since news storylines are expensive to annotate as they contain a myriad of unlabeled relationships growing quadratically with the number of news events. Working around these difficulties, we propose a sophisticated pre-processing method to filter candidate news pairs by entity co-occurrence and semantic similarity. With the filter reducing annotation overhead, we construct CStory, a large-scale Chinese news storyline dataset, which contains 11,978 news articles, 112,549 manually labeled storyline relation pairs, and 49,832 evidence sentences for annotation judgment. We conduct extensive experiments on CStory using various algorithms and find that constructing news storylines is challenging even for pre-trained language models. Empirical analysis shows that the sample unbalance issue significantly influences model performance, which shall be the focus of future works. Our dataset is now publicly available at https://github.com/THU-KEG/CStory. Kaijie Shi 0001, Xiaozhi Wang, Jifan Yu, Lei Hou 0001, Juan-Zi Li, Jingtong Wu, Dingyu Yong, Jinghui Xiao, Qun Liu 0001 |
CIKM | 9 |
| 2022 | Pan More Gold from the Sand: Refining Open-domain Dialogue Training with Noisy Self-Retrieval GenerationabstractReal human conversation data are complicated, heterogeneous, and noisy, from which building open-domain dialogue systems remains a challenging task. In fact, such dialogue data still contains a wealth of information and knowledge, however, they are not fully explored. In this paper, we show existing open-domain dialogue generation methods that memorize context-response paired data with autoregressive or encode-decode language models underutilize the training data. Different from current approaches, using external knowledge, we explore a retrieval-generation training framework that can take advantage of the heterogeneous and noisy training data by considering them as “evidence”. In particular, we use BERTScore for retrieval, which gives better qualities of the evidence and generation. Experiments over publicly available datasets demonstrate that our method can help models generate better responses, even such training data are usually impressed as low-quality data. Such performance gain is comparable with those improved by enlarging the training set, even better. We also found that the model performance has a positive correlation with the relevance of the retrieved evidence. Moreover, our method performed well on zero-shot experiments, which indicates that our method can be more robust to real-world data. Yasheng Wang, Fei Mi, Pingyi Zhou, Xin Wang 0114, Jin Liu 0016, Xin Jiang 0002, Qun Liu 0001 |
COLING | 9 |
| 2022 | UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual DialogabstractVisual Dialog aims to answer multi-round, interactive questions based on the dialog history and image content. Existing methods either consider answer ranking and generating individually or only weakly capture the relation across the two tasks implicitly by two separate models. The research on a universal framework that jointly learns to rank and generate answers in a single model is seldom explored. In this paper, we propose a contrastive learning-based framework UTC to unify and facilitate both discriminative and generative tasks in visual dialog with a single model. Specifically, considering the inherent limitation of the previous learning paradigm, we devise two inter-task contrastive losses i.e., context contrastive loss and answer contrastive loss to make the discriminative and generative tasks mutually reinforce each other. These two com-plementary contrastive losses exploit dialog context and target answer as anchor points to provide representation learning signals from different perspectives. We evaluate our proposed UTC on the VisDial v1.0 dataset, where our method outperforms the state-of-the-art on both discriminative and generative tasks and surpasses previous state-of-the-art generative methods by more than 2 absolute points on Recall@1. Zhenshan Tan, Qingrong Cheng, Xin Jiang 0002, Qun Liu 0001, Yudong Zhu, Xiaodong Gu 0001 |
CVPR | 5 |
| 2022 | LiteVL: Efficient Video-Language Learning with Enhanced Spatial-Temporal ModelingabstractRecent large-scale video-language pre-trained models have shown appealing performance on various downstream tasks.However, the pretraining process is computationally expensive due to the requirement of millions of videotext pairs and the redundant data structure of each video.To mitigate these problems, we propose LiteVL, which adapts a pre-trained image-language model BLIP into a video-text model directly on downstream tasks, without heavy pre-training.To enhance the temporal modeling lacking in the image-language model, we propose to add temporal attention modules in the image encoder of BLIP with dynamic temporal scaling.Besides the model-wise adaptation, we also propose a non-parametric pooling mechanism to adaptively reweight the fine-grained video embedding conditioned on the text.Experimental results on text-video retrieval and video question answering show that the proposed LiteVL even outperforms previous video-language pre-trained models by a clear margin, though without any videolanguage pre-training. Chaofan Tao, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
EMNLP | 6 |
| 2022 | Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language ProcessingabstractAbbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Phillippe Langlais. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Yasheng Wang, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Xin Jiang 0002, Qun Liu 0001, Philippe Langlais |
EMNLP | 13 |
| 2022 | Pre-training Language Models with Deterministic Factual KnowledgeabstractPrevious works show that Pre-trained Language Models (PLMs) can capture factual knowledge.However, some analyses reveal that PLMs fail to perform it robustly, e.g., being sensitive to the changes of prompts when extracting factual knowledge.To mitigate this issue, we propose to let PLMs learn the deterministic relationship between the remaining context and the masked content.The deterministic relationship ensures that the masked factual content can be deterministically inferable based on the existing clues in the context.That would provide more stable patterns for PLMs to capture factual knowledge than randomly masking.Two pre-training tasks are further introduced to motivate PLMs to rely on the deterministic relationship when filling masks.Specifically, we use an external Knowledge Base (KB) to identify deterministic relationships and continuously pre-train PLMs with the proposed methods.The factual knowledge probing experiments indicate that the continuously pre-trained PLMs achieve better robustness in factual knowledge capturing.Further experiments on question-answering datasets show that trying to learn a deterministic relationship with the proposed methods can also help other knowledge-intensive tasks. Shaobo Li 0004, Lifeng Shang, Chengjie Sun, Bingquan Liu, Zhenzhou Ji, Xin Jiang 0002, Qun Liu 0001 |
EMNLP | 8 |
| 2022 | COPEN: Probing Conceptual Knowledge in Pre-trained Language ModelsabstractConceptual knowledge is fundamental to human cognition and knowledge bases.However, existing knowledge probing works only focus on evaluating factual knowledge of pre-trained language models (PLMs) and ignore conceptual knowledge.Since conceptual knowledge often appears as implicit commonsense behind texts, designing probes for conceptual knowledge is hard.Inspired by knowledge representation schemata, we comprehensively evaluate conceptual knowledge of PLMs by designing three tasks to probe whether PLMs organize entities by conceptual similarities, learn conceptual properties, and conceptualize entities in contexts, respectively.For the tasks, we collect and annotate 24k data instances covering 393 concepts, which is COPEN, a COnceptual knowledge Probing bENchmark.Extensive experiments on different sizes and types of PLMs show that existing PLMs systematically lack conceptual knowledge and suffer from various spurious correlations.We believe this is a critical bottleneck for realizing human-like cognition in PLMs.COPEN and our codes are publicly released at https: //github.com/THU-KEG/COPEN. Hao Peng 0015, Xiaozhi Wang, Shengding Hu, Hailong Jin, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Qun Liu 0001 |
EMNLP | 8 |
| 2022 | G-MAP: General Memory-Augmented Pre-trained Language Model for Domain TasksabstractRecently, domain-specific PLMs have been proposed to boost the task performance of specific domains (e.g., biomedical and computer science) by continuing to pre-train general PLMs with domain-specific corpora.However, this Domain-Adaptive Pre-Training (DAPT; Gururangan et al. ( 2020)) tends to forget the previous general knowledge acquired by general PLMs, which leads to a catastrophic forgetting phenomenon and sub-optimal performance.To alleviate this problem, we propose a new framework of General Memory-Augmented Pre-trained Language Model (G-MAP), which augments the domain-specific PLM by a memory representation built from the frozen general PLM without losing any general knowledge.Specifically, we propose a new memory-augmented layer, and based on it, different augmented strategies are explored to build the memory representation and then adaptively fuse it into the domain-specific PLM.We demonstrate the effectiveness of G-MAP on various domains (biomedical and computer science publications, news, and reviews) and different kinds (text classification, QA, NER) of tasks, and the extensive results show that the proposed G-MAP 1 can achieve SOTA results on all tasks. Zhongwei Wan, Yichun Yin, Wei Zhang 0196, Jiaxin Shi, Lifeng Shang, Guangyong Chen, Xin Jiang 0002, Qun Liu 0001 |
EMNLP | 8 |
| 2022 | SPIRAL: Self-supervised Perturbation-Invariant Representation Learning for Speech Pre-Training
Wenyong Huang, Zhenhe Zhang, Yu Ting Yeung, Xin Jiang 0002, Qun Liu 0001 |
ICLR | 5 |
| 2022 | Exploring extreme parameter compression for pre-trained language models
Benyou Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
ICLR | 5 |
| 2022 | CoCA-MDD: A Coupled Cross-Attention based Framework for Streaming Mispronunciation Detection and DiagnosisabstractMispronunciation detection and diagnosis (MDD) is a popular research focus in computer-aided pronunciation training (CAPT) systems.End-to-end (e2e) approaches are becoming dominant in MDD.However an e2e MDD model usually requires entire speech utterances as input context, which leads to significant time latency especially for long paragraphs.We propose a streaming e2e MDD model called CoCA-MDD.We utilize conv-transformer structure to encode input speech in a streaming manner.A coupled cross-attention (CoCA) mechanism is proposed to integrate frame-level acoustic features with encoded reference linguistic features.CoCA also enables our model to perform mispronunciation classification with whole utterances.The proposed model allows system fusion between the streaming output and mispronunciation classification output for further performance enhancement.We evaluate CoCA-MDD on publicly available corpora.CoCA-MDD achieves F1 scores of 57.03% and 60.78% for streaming and fusion modes respectively on L2-ARCTIC.For phone-level pronunciation scoring, CoCA-MDD achieves 0.58 Pearson correlation coefficient (PCC) value on SpeechOcean762. Nianzu Zheng, Liqun Deng, Wenyong Huang, Yu Ting Yeung, Baohua Xu, Yasheng Wang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001 |
INTERSPEECH | 10 |
| 2022 | TGEA 2.0: A Large-Scale Diagnostically Annotated Dataset with Benchmark Tasks for Text Generation of Pretrained Language ModelsabstractIn order to diagnostically analyze and improve the capability of pretrained language models (PLMs) in text generation, we propose TGEA 2.0, to date the largest dataset built on machine-authored texts by PLMs with fine-grained semantic annotations on a wide variety of pathological generation errors. We collect 170K nominal, phrasal and sentential prompts from 6M natural sentences in 3 domains. These prompts are fed into 4 generative PLMs with their best decoding strategy to generate paragraphs. 195,629 sentences are extracted from these generated paragraphs for manual annotation, where 36K erroneous sentences are detected, 42K erroneous spans are located and categorized into an error type defined in a two-level error taxonomy. We define a \textbf{Mi}nimal \textbf{S}et of \textbf{E}rror-related \textbf{W}ords (MiSEW) for each erroneous span, which not only provides error-associated words but also rationalizes the reasoning behind the error. Quality control with a pre-annotation and feedback loop is performed before and during the entire annotation process. With the diagnostically annotated dataset, we propose 5 diagnosis benchmark tasks (i.e., erroneous text detection, MiSEW extraction, erroneous span location and correction together with error type classification) and 2 pathology mitigation benchmark tasks (pairwise comparison and word prediction). Experiment results on these benchmark tasks demonstrate that TGEA 2.0 is a challenging dataset that could facilitate further research on automatic diagnosis and pathology mitigation over machine texts. The dataset will be publicly available at https://github.com/tjunlp-lab/TGEA/. Huibin Ge, Chuang Liu 0009, Yulong Zeng, Qun Liu 0001, Deyi Xiong |
NeurIPS | 5 |
| 2022 | Dynamic Multi-Branch Layers for On-Device Neural Machine TranslationabstractWith the rapid development of artificial intelligence (AI), there is a trend in moving AI applications, such as neural machine translation (NMT), from cloud to mobile devices. Constrained by limited hardware resources and battery, the performance of on-device NMT systems is far from satisfactory. Inspired by conditional computation, we propose to improve the performance of on-device NMT systems with dynamic multi-branch layers. Specifically, we design a layer-wise dynamic multi-branch network with only one branch activated during training and inference. As not all branches are activated during training, we propose shared-private reparameterization to ensure sufficient training for each branch. At almost the same computational cost, our method achieves improvements of up to 1.7 BLEU points on the WMT14 English-German translation task and 1.8 BLEU points on the WMT20 Chinese-English translation task over the Transformer model, respectively. Compared with a strong baseline that also uses multiple branches, the proposed method is up to 1.5 times faster with the same number of parameters. Zhixing Tan, Zeyuan Yang 0002, Meng Zhang 0019, Qun Liu 0001, Maosong Sun 0001, Yang Liu 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | HopRetriever: Retrieve Hops over Wikipedia to Answer Complex QuestionsabstractCollecting supporting evidence from large corpora of text (e.g., Wikipedia) is of great challenge for open-domain Question Answering (QA). Especially, for multi-hop open-domain QA, scattered evidence pieces are required to be gathered together to support the answer extraction. In this paper, we propose a new retrieval target, hop, to collect the hidden reasoning evidence from Wikipedia for complex question answering. Specifically, the hop in this paper is defined as the combination of a hyperlink and the corresponding outbound link document. The hyperlink is encoded as the mention embedding which models the structured knowledge of how the outbound link entity is mentioned in the textual context, and the corresponding outbound link document is encoded as the document embedding representing the unstructured knowledge within it. Accordingly, we build HopRetriever which retrieves hops over Wikipedia to answer complex questions. Experiments on the HotpotQA dataset demonstrate that HopRetriever outperforms previously published evidence retrieval methods by large margins. Moreover, our approach also yields quantifiable interpretations of the evidence collection process. Shaobo Li 0004, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Chengjie Sun, Zhenzhou Ji, Bingquan Liu |
AAAI | 5 |
| 2021 | ALP-KD: Attention-Based Layer Projection for Knowledge DistillationabstractKnowledge distillation is considered as a training and compression strategy in which two neural networks, namely a teacher and a student, are coupled together during training. The teacher network is supposed to be a trustworthy predictor and the student tries to mimic its predictions. Usually, a student with a lighter architecture is selected so we can achieve compression and yet deliver high-quality results. In such a setting, distillation only happens for final predictions whereas the student could also benefit from teacher’s supervision for internal components. Motivated by this, we studied the problem of distillation for intermediate layers. Since there might not be a one-to-one alignment between student and teacher layers, existing techniques skip some teacher layers and only distill from a subset of them. This shortcoming directly impacts quality, so we instead propose a combinatorial technique which relies on attention. Our model fuses teacher-side information and takes each layer’s significance into consideration, then it performs distillation between combined teacher layers and those of the student. Using our technique, we distilled a 12-layer BERT (Devlin et al. 2019) into 6-, 4-, and 2-layer counterparts and evaluated them on GLUE tasks (Wang et al. 2018). Experimental results show that our combinatorial approach is able to outperform other existing techniques. Peyman Passban, Yimeng Wu, Mehdi Rezagholizadeh, Qun Liu 0001 |
AAAI | 4 |
| 2021 | Towards Semantics-Enhanced Pre-Training: Can Lexicon Definitions Help Learning Sentence Meanings?abstractSelf-supervised pre-training techniques, albeit relying on large amounts of text, have enabled rapid growth in learning language representations for natural language understanding. However, as radically empirical models on sentences, they are subject to the input data distribution, inevitably incorporating data bias and reporting bias, which may lead to inaccurate understanding of sentences. To address this problem, we propose to adopt a human learner's approach: when we cannot make sense of a word in a sentence, we often consult the dictionary for specific meanings; but can the same work for empirical models? In this work, we try to inform the pre-trained masked language models of word meanings for semantics-enhanced pre-training. To achieve a contrastive and holistic view of word meanings, a definition pair of two related words is presented to the masked language model such that the model can better associate a word with its crucial semantic features. Both intrinsic and extrinsic evaluations validate the proposed approach on semantics-orientated tasks, with an almost negligible increase of training data. Xuancheng Ren, Xu Sun 0001, Houfeng Wang, Qun Liu 0001 |
AAAI | 4 |
| 2021 | BinaryBERT: Pushing the Limit of BERT QuantizationabstractHaoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jin Jin, Xin Jiang, Qun Liu, Michael Lyu, Irwin King. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Haoli Bai, Wei Zhang 0196, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Michael R. Lyu, Irwin King |
ACL/IJCNLP (1) | 7 |
| 2021 | TGEA: An Error-Annotated Dataset and Benchmark Tasks for TextGeneration from Pretrained Language ModelsabstractJie He, Bo Peng, Yi Liao, Qun Liu, Deyi Xiong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jie He 0004, Bo Peng 0041, Qun Liu 0001, Deyi Xiong |
ACL/IJCNLP (1) | 4 |
| 2021 | GhostBERT: Generate More Features with Cheap Operations for BERTabstractZhiqi Huang, Lu Hou, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhiqi Huang 0001, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
ACL/IJCNLP (1) | 6 |
| 2021 | A Mutual Information Maximization Approach for the Spurious Solution Problem in Weakly Supervised Question AnsweringabstractZhihong Shao, Lifeng Shang, Qun Liu, Minlie Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhihong Shao, Lifeng Shang, Qun Liu 0001, Minlie Huang |
ACL/IJCNLP (1) | 3 |
| 2021 | AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language ModelsabstractYichun Yin, Cheng Chen, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
ACL/IJCNLP (1) | 6 |
| 2021 | Chinese WPLC: A Chinese Dataset for Evaluating Pretrained Language Models on Word Prediction Given Long-Range ContextabstractThis paper presents a Chinese dataset for evaluating pretrained language models on Word Prediction given Long-term Context (Chinese WPLC).We propose both automatic and manual selection strategies tailored to Chinese to guarantee that target words in passages collected from over 69K novels can only be predicted with long-term context beyond the scope of sentences containing the target words.Dataset analysis reveals that the types of target words range from common nouns to Chinese 4-character idioms.We also observe that linguistic relations between target words and long-range context exhibit diversity, including lexical match, synonym, summary and reasoning.Experiment results show that the Chinese pretrained language model PanGu-α (Zeng et al., 2021) is 45 points behind human in terms of top-1 word prediction accuracy, indicating that Chinese WPLC is a challenging dataset. Huibin Ge, Deyi Xiong, Qun Liu 0001 |
EMNLP (1) | 4 |
| 2021 | Improving Unsupervised Question Answering via Summarization-Informed Question GenerationabstractQuestion Generation (QG) is the task of generating a plausible question for a given pair.Template-based QG uses linguistically-informed heuristics to transform declarative sentences into interrogatives, whereas supervised QG uses existing Question Answering (QA) datasets to train a system to generate a question given a passage and an answer.A disadvantage of the heuristic approach is that the generated questions are heavily tied to their declarative counterparts.A disadvantage of the supervised approach is that they are heavily tied to the domain/language of the QA dataset used as training data.In order to overcome these shortcomings, we propose an unsupervised QG method which uses questions generated heuristically from summaries as a source of training data for a QG system.We make use of freely available news summary data, transforming declarative summary sentences into appropriate questions using heuristics informed by dependency parsing, named entity recognition and semantic role labeling.The resulting questions are then combined with the original news articles to train an end-to-end neural QG model.We extrinsically evaluate our approach using unsupervised QA: our QG model is used to generate synthetic QA pairs for training a QA model.Experimental results show that, trained with only 20k English Wikipedia-based synthetic QA pairs, the QA model substantially outperforms previous unsupervised models on three in-domain datasets (SQuAD1.1,Natural Questions, TriviaQA) and three out-of-domain datasets (NewsQA, BioASQ, DuoRC), demonstrating the transferability of the approach. Chenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 6 |
| 2021 | Neural Machine Translation with Heterogeneous Topic Knowledge EmbeddingsabstractNeural Machine Translation (NMT) has shown a strong ability to utilize local context to disambiguate the meaning of words.However, it remains a challenge for NMT to leverage broader context information like topics.In this paper, we propose heterogeneous ways of embedding topic information at the sentence level into an NMT model to improve translation performance.Specifically, the topic information can be incorporated as pre-encoder topic embedding, post-encoder topic embedding, and decoder topic embedding to increase the likelihood of selecting target words from the same topic of the source sentence.Experimental results show that NMT models with the proposed topic knowledge embedding outperform the baselines on the English → German and English → French translation tasks. Wei Peng 0011, Meng Zhang 0019, Qun Liu 0001 |
EMNLP (1) | 4 |
| 2021 | DyLex: Incorporating Dynamic Lexicons into BERT for Sequence LabelingabstractBaojun Wang, Zhao Zhang, Kun Xu, Guang-Yuan Hao, Yuyang Zhang, Lifeng Shang, Linlin Li, Xiao Chen, Xin Jiang, Qun Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Baojun Wang, Guang-Yuan Hao, Lifeng Shang, Linlin Li 0001, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 10 |
| 2021 | Uncertainty-Aware Balancing for Multilingual and Multi-Domain Neural Machine Translation TrainingabstractLearning multilingual and multi-domain translation model is challenging as the heterogeneous and imbalanced data make the model converge inconsistently over different corpora in real world.One common practice is to adjust the share of each corpus in the training, so that the learning process is balanced and low-resource cases can benefit from the highresource ones.However, automatic balancing methods usually depend on the intra-and interdataset characteristics, which is usually agnostic or requires human priors.In this work, we propose an approach, MULTIUAT, that dynamically adjusts the training data usage based on the model's uncertainty on a small set of trusted clean data for multi-corpus machine translation.We experiment with two classes of uncertainty measures on multilingual (16 languages with 4 settings) and multi-domain settings (4 for in-domain and 2 for out-of-domain on English-German translation) and demonstrate our approach MULTIUAT substantially outperforms its baselines, including both static and dynamic strategies.We analyze the crossdomain transfer and show the deficiency of static and similarity based methods. 1 Minghao Wu, Meng Zhang 0019, Liangyou Li, Gholamreza Haffari, Qun Liu 0001 |
EMNLP (1) | 6 |
| 2021 | Document Graph for Neural Machine TranslationabstractPrevious works have shown that contextual information can improve the performance of neural machine translation (NMT).However, most existing document-level NMT methods only consider a few number of previous sentences.How to make use of the whole document as global contexts is still a challenge.To address this issue, we hypothesize that a document can be represented as a graph that connects relevant contexts regardless of their distances.We employ several types of relations, including adjacency, syntactic dependency, lexical consistency, and coreference, to construct the document graph.Then, we incorporate both source and target graphs into the conventional Transformer architecture with graph convolutional networks.Experiments on various NMT benchmarks, including IWSLT English-French, Chinese-English, WMT English-German and Opensubtitle English-Russian, demonstrate that using document graphs can significantly improve the translation quality.Extensive analysis verifies that the document graph is beneficial for capturing discourse phenomena. Mingzhou Xu, Liangyou Li, Derek F. Wong, Qun Liu 0001, Lidia S. Chao |
EMNLP (1) | 4 |
| 2021 | Self-Supervised Quality Estimation for Machine TranslationabstractYuanhang Zheng, Zhixing Tan, Meng Zhang, Mieradilijiang Maimaiti, Huanbo Luan, Maosong Sun, Qun Liu, Yang Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Yuanhang Zheng, Zhixing Tan, Meng Zhang 0019, Mieradilijiang Maimaiti, Huan-Bo Luan, Maosong Sun 0001, Qun Liu 0001, Yang Liu 0005 |
EMNLP (1) | 7 |
| 2021 | Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation
Yichun Yin, Lifeng Shang, Zhi Wang 0001, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
ICANN (3) | 7 |
| 2021 | On Position Embeddings in BERT
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang 0002, Hao Yang 0006, Qun Liu 0001, Jakob Grue Simonsen |
ICLR | 6 |
| 2021 | Reweighting Augmented Samples by Minimizing the Maximal Expected Loss
Mingyang Yi, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Zhiming Ma |
ICLR | 5 |
| 2021 | Improved OOD Generalization via Adversarial Training and PretraingabstractRecently, learning a model that generalizes well on out-of-distribution (OOD) data has attracted great attention in the machine learning community. In this paper, after defining OOD generalization by Wasserstein distance, we theoretically justify that a model robust to input perturbation also generalizes well on OOD data. Inspired by previous findings that adversarial training helps improve robustness, we show that models trained by adversarial training have converged excess risk on OOD data. Besides, in the paradigm of pre-training then fine-tuning, we theoretically justify that the input perturbation robust model in the pre-training stage provides an initialization that generalizes well on downstream OOD data. Finally, various experiments conducted on image classification and natural language understanding tasks verify our theoretical findings. Mingyang Yi, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Zhiming Ma |
ICML | 6 |
| 2021 | Improving task-agnostic BERT distillation with layer mapping search
Xiaoqi Jiao, Huating Chang, Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Linlin Li 0001, Fang Wang 0001, Qun Liu 0001 |
Neurocomputing | 9 |
| 2021 | Learning to Generate Explainable Plots for Neural Story GenerationabstractStory generation is an important natural language processing task that aims to generate coherent stories automatically. While the use of neural networks has proven effective in improving story generation, how to learn to generate an explainable high-level plot still remains a major challenge. In this article, we propose a latent variable model for neural story generation. The model treats an outline, which is a natural language sentence explainable to humans, as a latent variable to represent a high-level plot that bridges the input and output. We adopt an external summarization model to guide the latent variable model to learn how to generate outlines from training data. Experiments show that our approach achieves significant improvements over state-of-the-art methods in both automatic and human evaluations. Gang Chen 0039, Yang Liu 0005, Huan-Bo Luan, Meng Zhang 0019, Qun Liu 0001, Maosong Sun 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Dialog State Tracking with Reinforced Data AugmentationabstractNeural dialog state trackers are generally limited due to the lack of quantity and diversity of annotated training data. In this paper, we address this difficulty by proposing a reinforcement learning (RL) based framework for data augmentation that can generate high-quality data to improve the neural state tracker. Specifically, we introduce a novel contextual bandit generator to learn fine-grained augmentation policies that can generate new effective instances by choosing suitable replacements for specific context. Moreover, by alternately learning between the generator and the state tracker, we can keep refining the generative policies to generate more high-quality training data for neural state tracker. Experimental results on the WoZ and MultiWoZ (restaurant) datasets demonstrate that the proposed framework significantly improves the performance over the state-of-the-art models, especially with limited training data. Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
AAAI | 5 |
| 2020 | Multi-Channel Reverse Dictionary ModelabstractA reverse dictionary takes the description of a target word as input and outputs the target word together with other words that match the description. Existing reverse dictionary methods cannot deal with highly variable input queries and low-frequency target words successfully. Inspired by the description-to-word inference process of humans, we propose the multi-channel reverse dictionary model, which can mitigate the two problems simultaneously. Our model comprises a sentence encoder and multiple predictors. The predictors are expected to identify different characteristics of the target word from the input query. We evaluate our model on English and Chinese datasets including both dictionary definitions and human-written descriptions. Experimental results show that our model achieves the state-of-the-art performance, and even outperforms the most popular commercial reverse dictionary system on the human-written description dataset. We also conduct quantitative analyses and a case study to demonstrate the effectiveness and robustness of our model. All the code and data of this work can be obtained on https://github.com/thunlp/MultiRD. Fanchao Qi, Zhiyuan Liu 0001, Yasheng Wang, Qun Liu 0001, Maosong Sun 0001 |
AAAI | 5 |
| 2020 | Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word OrderabstractMasked language model and autoregressive language model are two types of language models.While pretrained masked language models such as BERT (Devlin et al., 2019) overwhelm the line of natural language understanding (NLU) tasks, autoregressive language models such as GPT (Radford et al., 2018) are especially capable in natural language generation (NLG).In this paper, we propose a probabilistic masking scheme for the masked language model, which we call probabilistically masked language model (PMLM).We implement a specific PMLM with a uniform prior distribution on the masking ratio named u-PMLM.We prove that u-PMLM is equivalent to an autoregressive permutated language model.One main advantage of the model is that it supports text generation in arbitrary order with surprisingly good quality, which could potentially enable new applications over traditional unidirectional generation.Besides, the pretrained u-PMLM also outperforms BERT on a set of downstream NLU tasks. Xin Jiang 0002, Qun Liu 0001 |
ACL | 3 |
| 2020 | Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTabstractBy introducing a small set of additional parameters, a probe learns to solve specific linguistic tasks (e.g., dependency parsing) in a supervised manner using feature representations (e.g., contextualized embeddings).The effectiveness of such probing tasks is taken as evidence that the pre-trained model encodes linguistic knowledge.However, this approach of evaluating a language model is undermined by the uncertainty of the amount of knowledge that is learned by the probe itself.Complementary to those works, we propose a parameter-free probing technique for analyzing pre-trained language models (e.g., BERT).Our method does not require direct supervision from the probing tasks, nor do we introduce additional parameters to the probing process.Our experiments on BERT show that syntactic trees recovered from BERT using our method are significantly better than linguistically-uninformed baselines.We further feed the empirically induced dependency structures into a downstream sentiment classification task and find its improvement compatible with or even superior to a human-designed dependency schema. Zhiyong Wu 0003, Yun Chen 0007, Ben Kao, Qun Liu 0001 |
ACL | 4 |
| 2020 | Word-level Textual Adversarial Attacking as Combinatorial OptimizationabstractAdversarial attacks are carried out to reveal the vulnerability of deep neural networks.Textual adversarial attacking is challenging because text is discrete and a small perturbation can bring significant change to the original input.Word-level attacking, which can be regarded as a combinatorial optimization problem, is a well-studied class of textual attack methods.However, existing word-level attack models are far from perfect, largely because unsuitable search space reduction methods and inefficient optimization algorithms are employed.In this paper, we propose a novel attack model, which incorporates the sememebased word substitution method and particle swarm optimization-based search algorithm to solve the two problems separately.We conduct exhaustive experiments to evaluate our attack model by attacking BiLSTM and BERT on three benchmark datasets.Experimental results demonstrate that our model consistently achieves much higher attack success rates and crafts more high-quality adversarial examples as compared to baseline methods.Also, further experiments show our model has higher transferability and can bring more robustness enhancement to victim models by adversarial training. Yuan Zang, Fanchao Qi, Chenghao Yang 0001, Zhiyuan Liu 0001, Meng Zhang 0019, Qun Liu 0001, Maosong Sun 0001 |
ACL | 6 |
| 2020 | Accurate Word Alignment Induction from Neural Machine TranslationabstractDespite its original goal to jointly learn to align and translate, prior researches suggest that Transformer captures poor word alignments through its attention mechanism.In this paper, we show that attention weights DO capture accurate word alignments and propose two novel word alignment induction methods SHIFT-ATT and SHIFT-AET.The main idea is to induce alignments at the step when the to-be-aligned target token is the decoder input rather than the decoder output as in previous work.SHIFT-ATT is an interpretation method that induces alignments from the attention weights of Transformer and does not require parameter update or architecture change.SHIFT-AET extracts alignments from an additional alignment module which is tightly integrated into Transformer and trained in isolation with supervision from symmetrized SHIFT-ATT alignments.Experiments on three publicly available datasets demonstrate that both methods perform better than their corresponding neural baselines and SHIFT-AET significantly outperforms GIZA++ by 1.4-4.8AER points. 1 Yun Chen 0007, Yang Liu 0005, Guanhua Chen 0001, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 5 |
| 2020 | Why Skip If You Can Combine: A Simple Knowledge Distillation Technique for Intermediate LayersabstractWith the growth of computing power neural machine translation (NMT) models also grow accordingly and become better.However, they also become harder to deploy on edge devices due to memory constraints.To cope with this problem, a common practice is to distill knowledge from a large and accurately-trained teacher network (T ) into a compact student network (S).Although knowledge distillation (KD) is useful in most cases, our study shows that existing KD techniques might not be suitable enough for deep NMT engines, so we propose a novel alternative.In our model, besides matching T and S predictions we have a combinatorial mechanism to inject layer-level supervision from T to S. In this paper, we target low-resource settings and evaluate our translation engines for Portuguese→English, Turkish→English, and English→German directions.Students trained using our technique have 50% fewer parameters and can still deliver comparable results to those of 12-layer teachers. Yimeng Wu, Peyman Passban, Mehdi Rezagholizadeh, Qun Liu 0001 |
EMNLP (1) | 4 |
| 2020 | TernaryBERT: Distillation-aware Ultra-low Bit BERTabstractTransformer-based pre-training models like BERT have achieved remarkable performance in many natural language processing tasks.However, these models are both computation and memory expensive, hindering their deployment to resource-constrained devices.In this work, we propose TernaryBERT, which ternarizes the weights in a fine-tuned BERT model.Specifically, we use both approximation-based and loss-aware ternarization methods and empirically investigate the ternarization granularity of different parts of BERT.Moreover, to reduce the accuracy degradation caused by the lower capacity of low bits, we leverage the knowledge distillation technique (Jiao et al., 2019) in the training process.Experiments on the GLUE benchmark and SQuAD show that our proposed TernaryBERT outperforms the other BERT quantization methods, and even achieves comparable performance as the fullprecision model while being 14.9x smaller. Wei Zhang 0196, Lu Hou 0002, Yichun Yin, Lifeng Shang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 7 |
| 2020 | From Unsupervised Machine Translation to Adversarial Text GenerationabstractWe present a self-attention based bilingual adversarial text generator (B-GAN) which can learn to generate text from the encoder representation of an unsupervised neural machine translation system. B-GAN is able to generate a distributed latent space representation which can be paired with an attention based decoder to generate fluent sentences. When trained on an encoder shared between two languages and paired with the appropriate decoder, it can generate sentences in either language. B-GAN is trained using a combination of reconstruction loss for auto-encoder, a cross domain loss for translation and a GAN based adversarial loss for text generation. We demonstrate that B-GAN, trained on monolingual corpora only using multiple losses, generates more fluent sentences compared to monolingual baselines while effectively using half the number of parameters. Ahmad Rashid, Alan Do-Omri, Md. Akmal Haidar, Qun Liu 0001, Mehdi Rezagholizadeh |
ICASSP | 4 |
| 2020 | Bridging the Gap between Training and Inference for Neural Machine Translation (Extended Abstract)abstractNeural Machine Translation (NMT) generates target words sequentially in the way of predicting the next word conditioned on the context words. At training time, it predicts with the ground truth words as context while at inference it has to generate the entire sequence from scratch. This discrepancy of the fed context leads to error accumulation among the translation. Furthermore, word-level training requires strict matching between the generated sequence and the ground truth sequence which leads to overcorrection over different but reasonable translations. In this paper, we address these issues by sampling context words not only from the ground truth sequence but also from the predicted sequence during training. Experimental results on NIST Chinese->English and WMT2014 English->German translation tasks demonstrate that our method can achieve significant improvements on multiple data sets compared to strong baselines. Wen Zhang 0009, Yang Feng 0004, Qun Liu 0001 |
IJCAI | 3 |
| 2020 | DynaBERT: Dynamic BERT with Adaptive Width and DepthabstractThe pre-trained language models like BERT, though powerful in many natural language processing tasks, are both computation and memory expensive. To alleviate this problem, one approach is to compress them for specific tasks before deployment. However, recent works on BERT compression usually compress the large BERT model to a fixed smaller size, and can not fully satisfy the requirements of different edge devices with various hardware performances. In this paper, we propose a novel dynamic BERT model (abbreviated as DynaBERT), which can flexibly adjust the size and latency by selecting adaptive width and depth. The training process of DynaBERT includes first training a width-adaptive BERT and then allowing both adaptive width and depth, by distilling knowledge from the full-sized model to small sub-networks. Network rewiring is also used to keep the more important attention heads and neurons shared by more sub-networks. Comprehensive experiments under various efficiency constraints demonstrate that our proposed dynamic BERT (or RoBERTa) at its largest size has comparable performance as BERT-base (or RoBERTa-base), while at smaller widths and depths consistently outperforms existing BERT compression methods. Code is available at https://github.com/huawei-noah/Pretrained-Language-Model/tree/master/DynaBERT. Lu Hou 0002, Zhiqi Huang 0001, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
NeurIPS | 6 |
| 2020 | The Solution of Huawei Cloud & Noah's Ark Lab to the NLPCC-2020 Challenge: Light Pre-Training Chinese Language Model for NLP Task
Yichun Yin, Qun Liu 0001 |
NLPCC (2) | 6 |
| 2020 | PERQ: Predicting, Explaining, and Rectifying Failed Questions in KB-QA SystemsabstractA knowledge-based question-answering (KB-QA) system is one that answers natural-language questions by accessing information stored in a knowledge base (KB). Existing KB-QA systems generally register an accuracy of 70-80% for simple questions and less for more complex ones. We observe that certain questions are intrinsically difficult to answer correctly with existing systems. We propose the PERQ framework to address this issue. Given a question q, we perform three steps to boost answer accuracy: (1) (Prediction) We predict if q can be answered correctly by a KB-QA system S. (2) (Explanation) If S is predicted to fail q, we analyze them to determine the most likely reasons of the failure. (3) (Rectification) We use the prediction and explanation results to rectify the answer. We put forward tools to achieve the three steps and analyze their effectiveness. Our experiments show that the PERQ framework can significantly improve KB-QA systems' accuracies over simple questions. Zhiyong Wu 0003, Ben Kao, Tien-Hsuan Wu, Qun Liu 0001 |
WSDM | 5 |
| 2020 | Improving Sequence Modeling Ability of Recurrent Neural Networks via SememesabstractSememes, the minimum semantic units of human languages, have been successfully utilized in various natural language processing applications. However, most existing studies exploit sememes in specific tasks and few efforts are made to utilize sememes more fundamentally. In this paper, we propose to incorporate sememes into recurrent neural networks (RNNs) to improve their sequence modeling ability, which is beneficial to all kinds of downstream tasks. We design three different sememe incorporation methods and employ them in typical RNNs including LSTM, GRU and their bidirectional variants. In evaluation, we use several benchmark datasets involving PTB and WikiText-2 for language modeling, SNLI for natural language inference and another two datasets for sentiment analysis and paraphrase detection. Experimental results show evident and consistent improvement of our sememe-incorporated models compared with vanilla RNNs, which proves the effectiveness of our sememe incorporation methods. Moreover, we find the sememe-incorporated models have higher robustness and outperform adversarial training in defending adversarial attack. All the code and data of this work can be obtained at https://github.com/thunlp/SememeRNN. Yujia Qin, Fanchao Qi, Sicong Ouyang, Zhiyuan Liu 0001, Cheng Yang 0002, Yasheng Wang, Qun Liu 0001, Maosong Sun 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2019 | Decomposable Neural Paraphrase GenerationabstractParaphrasing exists at different granularity levels, such as lexical level, phrasal level and sentential level.This paper presents Decomposable Neural Paraphrase Generator (DNPG), a Transformer-based model that can learn and generate paraphrases of a sentence at different levels of granularity in a disentangled way.Specifically, the model is composed of multiple encoders and decoders with different structures, each of which corresponds to a specific granularity.The empirical study shows that the decomposition mechanism of DNPG makes paraphrase generation more interpretable and controllable.Based on DNPG, we further develop an unsupervised domain adaptation method for paraphrase generation.Experimental results show that the proposed model achieves competitive in-domain performance compared to the state-of-the-art neural models, and significantly better performance when adapting to a new domain.What is the population of New York?How many people is there in NYC?Who wrote the Winnie the Pooh books?Who is the author of winnie the pooh?What is the best phone to buy below 15k?Which are best mobile phones to buy under 15000?How can I be a good geologist?What should I do to be a great geologist?How do I reword a sentence to avoid plagiarism?How can I paraphrase my essay and avoid plagiarism? Zichao Li 0001, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001 |
ACL (1) | 4 |
| 2019 | Modeling Semantic Compositionality with Sememe KnowledgeabstractSemantic compositionality (SC) refers to the phenomenon that the meaning of a complex linguistic unit can be composed of the meanings of its constituents.Most related works focus on using complicated compositionality functions to model SC while few works consider external knowledge in models.In this paper, we verify the effectiveness of sememes, the minimum semantic units of human languages, in modeling SC by a confirmatory experiment.Furthermore, we make the first attempt to incorporate sememe knowledge into SC models, and employ the sememeincorporated models in learning representations of multiword expressions, a typical task of SC.In experiments, we implement our models by incorporating knowledge from a famous sememe knowledge base HowNet and perform both intrinsic and extrinsic evaluations.Experimental results show that our models achieve significant performance boost as compared to the baseline methods without considering sememe knowledge.We further conduct quantitative analysis and case studies to demonstrate the effectiveness of applying sememe knowledge in modeling SC.All the code and data of this paper can be obtained on https: //github.com/thunlp/Sememe-SC. Fanchao Qi, Junjie Huang 0003, Chenghao Yang 0001, Zhiyuan Liu 0001, Xiao Chen 0012, Qun Liu 0001, Maosong Sun 0001 |
ACL (1) | 6 |
| 2019 | Bridging the Gap between Training and Inference for Neural Machine TranslationabstractNeural Machine Translation (NMT) generates target words sequentially in the way of predicting the next word conditioned on the context words.At training time, it predicts with the ground truth words as context while at inference it has to generate the entire sequence from scratch.This discrepancy of the fed context leads to error accumulation among the way.Furthermore, word-level training requires strict matching between the generated sequence and the ground truth sequence which leads to overcorrection over different but reasonable translations.In this paper, we address these issues by sampling context words not only from the ground truth sequence but also from the predicted sequence by the model during training, where the predicted sequence is selected with a sentence-level optimum.Experiment results on Chinese→English and WMT'14 English→German translation tasks demonstrate that our approach can achieve significant improvements on multiple datasets. Wen Zhang 0009, Yang Feng 0004, Fandong Meng, Di You, Qun Liu 0001 |
ACL (1) | 5 |
| 2019 | ERNIE: Enhanced Language Representation with Informative EntitiesabstractNeural language representation models such as BERT pre-trained on large-scale corpora can well capture rich semantic patterns from plain text, and be fine-tuned to consistently improve the performance of various NLP tasks.However, the existing pre-trained language models rarely consider incorporating knowledge graphs (KGs), which can provide rich structured knowledge facts for better language understanding.We argue that informative entities in KGs can enhance language representation with external knowledge.In this paper, we utilize both large-scale textual corpora and KGs to train an enhanced language representation model (ERNIE), which can take full advantage of lexical, syntactic, and knowledge information simultaneously.The experimental results have demonstrated that ERNIE achieves significant improvements on various knowledge-driven tasks, and meanwhile is comparable with the state-of-the-art model BERT on other common NLP tasks.The source code and experiment details of this paper can be obtained from https:// github.com/thunlp/ERNIE. Zhengyan Zhang, Xu Han 0007, Zhiyuan Liu 0001, Xin Jiang 0002, Maosong Sun 0001, Qun Liu 0001 |
ACL (1) | 6 |
| 2019 | An error analysis for image-based multi-modal neural machine translationabstractIn this article, we conduct an extensive quantitative error analysis of different multi-modal neural machine translation (MNMT) models which integrate visual features into different parts of both the encoder and the decoder. We investigate the scenario where models are trained on an in-domain training data set of parallel sentence pairs with images. We analyse two different types of MNMT models, that use global and local image features: the latter encode an image globally, i.e. there is one feature vector representing an entire image, whereas the former encode spatial information, i.e. there are multiple feature vectors, each encoding different portions of the image. We conduct an error analysis of translations generated by different MNMT models as well as text-only baselines, where we study how multi-modal models compare when translating both visual and non-visual terms. In general, we find that the additional multi-modal signals consistently improve translations, even more so when using simpler MNMT models that use global visual features. We also find that not only translations of terms with a strong visual connotation are improved, but almost all kinds of errors decreased when using multi-modal models. Iacer Calixto, Qun Liu 0001 |
Mach. Transl. | 2 |
| 2019 | Machine Translation Evaluation Metric Based on Dependency Parsing ModelabstractMost of the syntax-based metrics obtain the similarity by comparing the sub-structures extracted from the trees of hypothesis and reference. These sub-structures cannot represent all the information in the trees because their lengths are limited. To sufficiently use the reference syntax information, a new automatic evaluation metric is proposed based on the dependency parsing model. First, a dependency parsing model is trained using the reference dependency tree for each sentence. Then, the hypothesis is parsed by this dependency parsing model and the corresponding hypothesis dependency tree is generated. The quality of hypothesis can be judged by the quality of the hypothesis dependency tree. Unigram F-score is included in the new metric so that lexicon similarity is obtained. According to experimental results, the proposed metric can perform better than METEOR and BLEU on system level and get comparable results with METEOR on sentence level. To further improve the performance, we also propose a combined metric which gets the best performance on the sentence level and on the system level. Hui Yu 0010, Weizhi Xu 0001, Shouxun Lin, Qun Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2018 | Translating Pro-Drop Languages With Reconstruction ModelsabstractPronouns are frequently omitted in pro-drop languages, such as Chinese, generally leading to significant challenges with respect to the production of complete translations. To date, very little attention has been paid to the dropped pronoun (DP) problem within neural machine translation (NMT). In this work, we propose a novel reconstruction-based approach to alleviating DP translation problems for NMT models. Firstly, DPs within all source sentences are automatically annotated with parallel information extracted from the bilingual training corpus. Next, the annotated source sentence is reconstructed from hidden representations in the NMT model. With auxiliary training objectives, in the terms of reconstruction scores, the parameters associated with the NMT model are guided to produce enhanced hidden representations that are encouraged as much as possible to embed annotated DP information. Experimental results on both Chinese-English and Japanese-English dialogue translation tasks show that the proposed approach significantly and consistently improves translation performance over a strong NMT baseline, which is directly built on the training data annotated with DPs. Longyue Wang, Zhaopeng Tu, Shuming Shi 0001, Tong Zhang 0001, Yvette Graham, Qun Liu 0001 |
AAAI | 6 |
| 2018 | Knowledge Diffusion for Neural Dialogue GenerationabstractEnd-to-end neural dialogue generation has shown promising results recently, but it does not employ knowledge to guide the generation and hence tends to generate short, general, and meaningless responses.In this paper, we propose a neural knowledge diffusion (NKD) model to introduce knowledge into dialogue generation.This method can not only match the relevant facts for the input utterance but diffuse them to similar entities.With the help of facts matching and entity diffusion, the neural dialogue generation is augmented with the ability of convergent and divergent thinking over the knowledge base.Our empirical study on a real-world dataset proves that our model is capable of generating meaningful, diverse and natural responses for both factoid-questions and knowledge grounded chi-chats.The experiment results also show that our model outperforms competitive baseline models significantly. Shuman Liu, Hongshen Chen, Zhaochun Ren, Yang Feng 0004, Qun Liu 0001, Dawei Yin 0001 |
ACL (1) | 5 |
| 2018 | Tailoring Neural Architectures for Translating from Morphologically Rich LanguagesabstractA morphologically complex word (MCW) is a hierarchical constituent with meaning-preserving subunits, so word-based models which rely on surface forms might not be powerful enough to translate such structures. When translating from morphologically rich languages (MRLs), a source word could be mapped to several words or even a full sentence on the target side, which means an MCW should not be treated as an atomic unit. In order to provide better translations for MRLs, we boost the existing neural machine translation (NMT) architecture with a double- channel encoder and a double-attentive decoder. The main goal targeted in this research is to provide richer information on the encoder side and redesign the decoder accordingly to benefit from such information. Our experimental results demonstrate that we could achieve our goal as the proposed model outperforms existing subword- and character-based architectures and showed significant improvements on translating from German, Russian, and Turkish into English. Peyman Passban, Andy Way, Qun Liu 0001 |
COLING | 3 |
| 2018 | Refining Source Representations with Relation Networks for Neural Machine TranslationabstractAlthough neural machine translation with the encoder-decoder framework has achieved great success recently, it still suffers drawbacks of forgetting distant information, which is an inherent disadvantage of recurrent neural network structure, and disregarding relationship between source words during encoding step. Whereas in practice, the former information and relationship are often useful in current step. We target on solving these problems and thus introduce relation networks to learn better representations of the source. The relation networks are able to facilitate memorization capability of recurrent neural network via associating source words with each other, this would also help retain their relationships. Then the source representations and all the relations are fed into the attention component together while decoding, with the main encoder-decoder framework unchanged. Experiments on several datasets show that our method can improve the translation performance significantly over the conventional encoder-decoder model and even outperform the approach involving supervised syntactic knowledge. Wen Zhang 0009, Yang Feng 0004, Qun Liu 0001 |
COLING | 4 |
| 2018 | Learning to Jointly Translate and Predict Dropped Pronouns with a Shared Reconstruction MechanismabstractPronouns are frequently omitted in pro-drop languages, such as Chinese, generally leading to significant challenges with respect to the production of complete translations.Recently, Wang et al. (2018) proposed a novel reconstruction-based approach to alleviating dropped pronoun (DP) translation problems for neural machine translation models.In this work, we improve the original model from two perspectives.First, we employ a shared reconstructor to better exploit encoder and decoder representations.Second, we jointly learn to translate and predict DPs in an end-to-end manner, to avoid the errors propagated from an external DP prediction model.Experimental results show that our approach significantly improves both translation performance and DP prediction accuracy. Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001 |
EMNLP | 4 |
| 2018 | Speeding Up Neural Machine Translation Decoding by Cube PruningabstractAlthough neural machine translation has achieved promising results, it suffers from slow translation speed.The direct consequence is that a trade-off has to be made between translation quality and speed, thus its performance can not come into full play.We apply cube pruning, a popular technique to speed up dynamic programming, into neural machine translation to speed up the translation.To construct the equivalence class, similar target hidden states are combined, leading to less RNN expansion operations on the target side and less softmax operations over the large target vocabulary.The experiments show that, at the same or even better translation quality, our method can translate faster compared with naive beam search by 3.3× on GPUs and 3.5× on CPUs. Wen Zhang 0009, Liang Huang 0001, Yang Feng 0004, Lei Shen 0001, Qun Liu 0001 |
EMNLP | 5 |
| 2018 | Learning Tag Dependencies for Sequence TaggingabstractSequence tagging is the basis for multiple applications in natural language processing. Despite successes in learning long term token sequence dependencies with neural network, tag dependencies are rarely considered previously. Sequence tagging actually possesses complex dependencies and interactions among the input tokens and the output tags. We propose a novel multi-channel model, which handles different ranges of token-tag dependencies and their interactions simultaneously. A tag LSTM is augmented to manage the output tag dependencies and word-tag interactions, while three mechanisms are presented to efficiently incorporate token context representation and tag dependency. Extensive experiments on part-of-speech tagging and named entity recognition tasks show that the proposed model outperforms the BiLSTM-CRF baseline by effectively incorporating the tag dependency feature. Yuan Zhang 0015, Hongshen Chen, Yihong Eric Zhao, Qun Liu 0001, Dawei Yin 0001 |
IJCAI | 4 |
| 2018 | E2E NLG Challenge Submission: Towards Controllable Generation of Diverse Natural LanguageabstractIn natural language generation (NLG), the task is to generate utterances from a more abstract input, such as structured data.An added challenge is to generate utterances that contain an accurate representation of the input, while reflecting the fluency and variety of human-generated text.In this paper, we report experiments with NLG models that can be used in task oriented dialogue systems.We explore the use of additional input to the model to encourage diversity and control of outputs.While our submission does not rank highly using automated metrics, qualitative investigation of generated utterances suggests the use of additional information in neural network NLG systems to be a promising research direction. Henry Elder, Sebastian Gehrmann, Alexander O'Connor, Qun Liu 0001 |
INLG | 4 |
| 2018 | Improving Character-Based Decoding Using Target-Side Morphological Information for Neural Machine TranslationabstractPeyman Passban, Qun Liu, Andy Way. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Peyman Passban, Qun Liu 0001, Andy Way |
NAACL-HLT | 2 |
| 2017 | Doubly-Attentive Decoder for Multi-modal Neural Machine TranslationabstractWe introduce a Multi-modal Neural Machine Translation model in which a doubly-attentive decoder naturally incorporates spatial visual features obtained using pre-trained convolutional neural networks, bridging the gap between image description and translation.Our decoder learns to attend to source-language words and parts of an image independently by means of two separate attention mechanisms as it generates words in the target language.We find that our model can efficiently exploit not just back-translated in-domain multi-modal data but also large general-domain text-only MT corpora.We also report state-of-the-art results on the Multi30k data set. Iacer Calixto, Qun Liu 0001, Nick Campbell 0001 |
ACL (1) | 2 |
| 2017 | Lexically Constrained Decoding for Sequence Generation Using Grid Beam SearchabstractWe present Grid Beam Search (GBS), an algorithm which extends beam search to allow the inclusion of pre-specified lexical constraints.The algorithm can be used with any model that generates a sequence ŷ = {y 0 . . .y T }, by maximizing p(y|x) = t p(y t |x; {y 0 . . .y t-1 }).Lexical constraints take the form of phrases or words that must be present in the output sequence.This is a very general way to incorporate additional knowledge into a model's output without requiring any modification of the model parameters or training data.We demonstrate the feasibility and flexibility of Lexically Constrained Decoding by conducting experiments on Neural Interactive-Predictive Translation, as well as Domain Adaptation for Neural Machine Translation.Experiments show that GBS can provide large improvements in translation quality in interactive scenarios, and that, even without any user input, GBS can be used to achieve significant gains in performance in domain adaptation scenarios. Chris Hokamp, Qun Liu 0001 |
ACL (1) | 2 |
| 2017 | Deep Neural Machine Translation with Linear Associative UnitabstractDeep Neural Networks (DNNs) have provably enhanced the state-of-the-art Neural Machine Translation (NMT) with their capability in modeling complex functions and capturing complex linguistic structures.However NMT systems with deep architecture in their encoder or decoder RNNs often suffer from severe gradient diffusion due to the non-linear recurrent activations, which often make the optimization much more difficult.To address this problem we propose novel linear associative units (LAU) to reduce the gradient propagation length inside the recurrent unit.Different from conventional approaches (LSTM unit and GRU), LAUs utilizes linear associative connections between input and output of the recurrent unit, which allows unimpeded information flow through both space and time direction.The model is quite simple, but it is surprisingly effective.Our empirical study on Chinese-English translation shows that our model with proper configuration can improve by 11.7 BLEU upon Groundhog and the best reported results in the same setting.On WMT14 English-German task and a larger WMT14 English-French task, our model achieves comparable results with the state-of-the-art. Mingxuan Wang, Zhengdong Lu, Jie Zhou 0016, Qun Liu 0001 |
ACL (1) | 4 |
| 2017 | Incorporating Word Reordering Knowledge into Attention-based Neural Machine TranslationabstractThis paper proposes three distortion models to explicitly incorporate the word reordering knowledge into attention-based Neural Machine Translation (NMT) for further improving translation performance.Our proposed models enable attention mechanism to attend to source words regarding both the semantic requirement and the word reordering penalty.Experiments on Chinese-English translation show that the approaches can improve word alignment quality and achieve significant translation improvements over a basic attention-based N-MT by large margins.Compared with previous works on identical corpora, our system achieves the state-of-the-art performance on translation quality. Jinchao Zhang 0001, Mingxuan Wang, Qun Liu 0001, Jie Zhou 0016 |
ACL (1) | 3 |
| 2017 | If You Can't Beat Them Join Them: Handcrafted Features Complement Neural Nets for Non-Factoid Answer RerankingabstractWe show that a neural approach to the task of non-factoid answer reranking can benefit from the inclusion of tried-and-tested handcrafted features.We present a novel neural network architecture based on a combination of recurrent neural networks that are used to encode questions and answers, and a multilayer perceptron.We show how this approach can be combined with additional features, in particular, the discourse features presented by Jansen et al. (2014).Our neural approach achieves state-of-the-art performance on a public dataset from Yahoo! Answers and its performance is further improved by incorporating the discourse features.Additionally, we present a new dataset of Ask Ubuntu questions where the hybrid approach also achieves good results. Dasha Bogdanova, Jennifer Foster, Daria Dzendzik, Qun Liu 0001 |
EACL (1) | 4 |
| 2017 | Incorporating Global Visual Features into Attention-based Neural Machine TranslationabstractWe introduce multi-modal, attentionbased Neural Machine Translation (NMT) models which incorporate visual features into different parts of both the encoder and the decoder.Global image features are extracted using a pre-trained convolutional neural network and are incorporated (i) as words in the source sentence, (ii) to initialise the encoder hidden state, and (iii) as additional data to initialise the decoder hidden state.In our experiments, we evaluate translations into English and German, how different strategies to incorporate global image features compare and which ones perform best.We also study the impact that adding synthetic multi-modal, multilingual data brings and find that the additional data have a positive impact on multi-modal models.We report new state-of-the-art results and our best models also significantly improve on a comparable Phrase-Based Statistical MT (PBSMT) model trained on the Multi30k data set according to all metrics evaluated.To the best of our knowledge, it is the first time a purely neural model significantly improves over a PBSMT model on all metrics evaluated on this data set. Iacer Calixto, Qun Liu 0001 |
EMNLP | 2 |
| 2017 | Further Investigation into Reference Bias in Monolingual Evaluation of Machine TranslationabstractMonolingual evaluation of Machine Translation (MT) aims to simplify human assessment by requiring assessors to compare the meaning of the MT output with a reference translation, opening up the task to a much larger pool of genuinely qualified evaluators.Monolingual evaluation runs the risk, however, of bias in favour of MT systems that happen to produce translations superficially similar to the reference and, consistent with this intuition, previous investigations have concluded monolingual assessment to be strongly biased in this respect.On re-examination of past analyses, we identify a series of potential analytical errors that force some important questions to be raised about the reliability of past conclusions, however.We subsequently carry out further investigation into reference bias via direct human assessment of MT adequacy via quality controlled crowd-sourcing.Contrary to both intuition and past conclusions, results show no significant evidence of reference bias in monolingual evaluation of MT. Qingsong Ma, Yvette Graham, Timothy Baldwin, Qun Liu 0001 |
EMNLP | 4 |
| 2017 | Exploiting Cross-Sentence Context for Neural Machine TranslationabstractIn translation, considering the document as a whole can help to resolve ambiguities and inconsistencies.In this paper, we propose a cross-sentence context-aware approach and investigate the influence of historical contextual information on the performance of neural machine translation (NMT).First, this history is summarized in a hierarchical way.We then integrate the historical representation into NMT in two strategies: 1) a warm-start of encoder and decoder states, and 2) an auxiliary context source for updating decoder states.Experimental results on a large Chinese-English translation task show that our approach significantly improves upon a strong attention-based NMT system by up to +2.1 BLEU points. Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001 |
EMNLP | 4 |
| 2017 | ME-MD: An Effective Framework for Neural Machine Translation with Multiple Encoders and DecodersabstractThe encoder-decoder neural framework is widely employed for Neural Machine Translation (NMT) with a single encoder to represent the source sentence and a single decoder to generate target words. The translation performance heavily relies on the representation ability of the encoder and the generation ability of the decoder. To further enhance NMT, we propose to extend the original encoder-decoder framework to a novel one, which has multiple encoders and decoders (ME-MD). Through this way, multiple encoders extract more diverse features to represent the source sequence and multiple decoders capture more complicated translation knowledge. Our proposed ME-MD framework is convenient to integrate heterogeneous encoders and decoders with multiple depths and multiple types. Experiment on Chinese-English translation task shows that our ME-MD system surpasses the state-of-the-art NMT system by 2.1 BLEU points and surpasses the phrase-based Moses by 7.38 BLEU points. Our framework is general and can be applied to other sequence to sequence tasks. Jinchao Zhang 0001, Qun Liu 0001, Jie Zhou 0016 |
IJCAI | 2 |
| 2017 | Editorial
Qun Liu 0001, Xiaodong He 0001, Hermann Ney |
Mach. Transl. | 1 |
| 2017 | A novel and robust approach for pro-drop language translationabstractA significant challenge for machine translation (MT) is the phenomena of dropped pronouns (DPs), where certain classes of pronouns are frequently dropped in the source language but should be retained in the target language. In response to this common problem, we propose a semi-supervised approach with a universal framework to recall missing pronouns in translation. Firstly, we build training data for DP generation in which the DPs are automatically labelled according to the alignment information from a parallel corpus. Secondly, we build a deep learning-based DP generator for input sentences in decoding when no corresponding references exist. More specifically, the generation has two phases: (1) DP position detection, which is modeled as a sequential labelling task with recurrent neural networks; and (2) DP prediction, which employs a multilayer perceptron with rich features. Finally, we integrate the above outputs into our statistical MT (SMT) system to recall missing pronouns by both extracting rules from the DP-labelled training data and translating the DP-generated input sentences. To validate the robustness of our approach, we investigate our approach on both Chinese–English and Japanese–English corpora extracted from movie subtitles. Compared with an SMT baseline system, experimental results show that our approach achieves a significant improvement of $$+$$ 1.58 BLEU points in translation performance with 66% F-score for DP generation accuracy for Chinese–English, and nearly $$+$$ 1 BLEU point with 58% F-score for Japanese–English. We believe that this work could help both MT researchers and industries to boost the performance of MT systems between pro-drop and non-pro-drop languages. Longyue Wang, Zhaopeng Tu, Siyou Liu, Hang Li 0001, Andy Way, Qun Liu 0001 |
Mach. Transl. | 7 |
| 2017 | Translating Low-Resource Languages by Vocabulary Adaptation from Close CounterpartsabstractSome natural languages belong to the same family or share similar syntactic and/or semantic regularities. This property persuades researchers to share computational models across languages and benefit from high-quality models to boost existing low-performance counterparts. In this article, we follow a similar idea, whereby we develop statistical and neural machine translation (MT) engines that are trained on one language pair but are used to translate another language. First we train a reliable model for a high-resource language, and then we exploit cross-lingual similarities and adapt the model to work for a close language with almost zero resources. We chose Turkish (Tr) and Azeri or Azerbaijani (Az) as the proposed pair in our experiments. Azeri suffers from lack of resources as there is almost no bilingual corpus for this language. Via our techniques, we are able to train an engine for the Az → English (En) direction, which is able to outperform all other existing models. Peyman Passban, Qun Liu 0001, Andy Way |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2016 | Graph-Based Translation Via Graph SegmentationabstractOne major drawback of phrase-based translation is that it segments an input sentence into continuous phrases.To support linguistically informed source discontinuity, in this paper we construct graphs which combine bigram and dependency relations and propose a graph-based translation model.The model segments an input graph into connected subgraphs, each of which may cover a discontinuous phrase.We use beam search to combine translations of each subgraph left-to-right to produce a complete translation.Experiments on Chinese-English and German-English tasks show that our system is significantly better than the phrase-based model by up to +1.5/+0.5 BLEU scores.By explicitly modeling the graph segmentation, our system obtains further improvement, especially on German-English. Liangyou Li, Andy Way, Qun Liu 0001 |
ACL (1) | 3 |
| 2016 | Interactive Attention for Neural Machine TranslationabstractConventional attention-based Neural Machine Translation (NMT) conducts dynamic alignment in generating the target sentence. By repeatedly reading the representation of source sentence, which keeps fixed after generated by the encoder (Bahdanau et al., 2015), the attention mechanism has greatly enhanced state-of-the-art NMT. In this paper, we propose a new attention mechanism, called INTERACTIVE ATTENTION, which models the interaction between the decoder and the representation of source sentence during translation by both reading and writing operations. INTERACTIVE ATTENTION can keep track of the interaction history and therefore improve the translation performance. Experiments on NIST Chinese-English translation task show that INTERACTIVE ATTENTION can achieve significant improvements over both the previous attention-based NMT baseline and some state-of-the-art variants of attention-based NMT (i.e., coverage models (Tu et al., 2016)). And neural machine translator with our INTERACTIVE ATTENTION can outperform the open source attention-based NMT system Groundhog by 4.22 BLEU points and the open source phrase-based system Moses by 3.94 BLEU points averagely on multiple test sets. Fandong Meng, Zhengdong Lu, Hang Li 0001, Qun Liu 0001 |
COLING | 4 |
| 2016 | Enriching Phrase Tables for Statistical Machine Translation Using Mixed EmbeddingsabstractThe phrase table is considered to be the main bilingual resource for the phrase-based statistical machine translation (PBSMT) model. During translation, a source sentence is decomposed into several phrases. The best match of each source phrase is selected among several target-side counterparts within the phrase table, and processed by the decoder to generate a sentence-level translation. The best match is chosen according to several factors, including a set of bilingual features. PBSMT engines by default provide four probability scores in phrase tables which are considered as the main set of bilingual features. Our goal is to enrich that set of features, as a better feature set should yield better translations. We propose new scores generated by a Convolutional Neural Network (CNN) which indicate the semantic relatedness of phrase pairs. We evaluate our model in different experimental settings with different language pairs. We observe significant improvements when the proposed features are incorporated into the PBSMT pipeline. Peyman Passban, Qun Liu 0001, Andy Way |
COLING | 2 |
| 2016 | Topic-Informed Neural Machine TranslationabstractIn recent years, neural machine translation (NMT) has demonstrated state-of-the-art machine translation (MT) performance. It is a new approach to MT, which tries to learn a set of parameters to maximize the conditional probability of target sentences given source sentences. In this paper, we present a novel approach to improve the translation performance in NMT by conveying topic knowledge during translation. The proposed topic-informed NMT can increase the likelihood of selecting words from the same topic and domain for translation. Experimentally, we demonstrate that topic-informed NMT can achieve a 1.15 (3.3% relative) and 1.67 (5.4% relative) absolute improvement in BLEU score on the Chinese-to-English language pair using NIST 2004 and 2005 test sets, respectively, compared to NMT without topic information. Jian Zhang 0003, Liangyou Li, Andy Way, Qun Liu 0001 |
COLING | 4 |
| 2016 | Fast Gated Neural Domain Adaptation: Language Model as a Case StudyabstractNeural network training has been shown to be advantageous in many natural language processing applications, such as language modelling or machine translation. In this paper, we describe in detail a novel domain adaptation mechanism in neural network training. Instead of learning and adapting the neural network on millions of training sentences – which can be very time-consuming or even infeasible in some cases – we design a domain adaptation gating mechanism which can be used in recurrent neural networks and quickly learn the out-of-domain knowledge directly from the word vector representations with little speed overhead. In our experiments, we use the recurrent neural network language model (LM) as a case study. We show that the neural LM perplexity can be reduced by 7.395 and 12.011 using the proposed domain adaptation mechanism on the Penn Treebank and News data, respectively. Furthermore, we show that using the domain-adapted neural LM to re-rank the statistical machine translation n-best list on the French-to-English language pair can significantly improve translation quality. Jian Zhang 0003, Andy Way, Qun Liu 0001 |
COLING | 4 |
| 2016 | A subtree-based factorization of dependency parsingabstractWe propose a dependency parsing pipeline, in which the parsing of long-distance projections and localized dependencies are explicitly decomposed at the input level. A chosen baseline dependency parsing model performs only on ‘carved’ sequences at the second stage, which are transformed from coarse constituent parsing outputs at the first stage. When k-best constituent parsing outputs are kept, a third-stage is required to search for an optimal combination of the overlapped dependency subtrees. In this sense, our dependency model is subtree-factored. We explore alternative approaches for scoring subtrees, including feature-based models as well as continuous representations. The search for optimal subset to combine is formulated as an ILP problem. This framework especially benefits the models poor on long sentences, generally improving baselines by 0.75-1.28 (UAS) on English, achieving comparable performance with high-order models but faster. For Chinese, the most notable increase is as high as 3.63 (UAS) when the proposed framework is applied to first-order parsing models. Qiuye Zhao, Qun Liu 0001 |
COLING | 2 |
| 2016 | Combining Translation Memories and Syntax-Based SMT: Experiments with Real Industrial Data
Liangyou Li, Carla Parra Escartín, Qun Liu 0001 |
EAMT | 3 |
| 2016 | Improving Phrase-Based SMT Using Cross-Granularity Embedding Similarity
Peyman Passban, Chris Hokamp, Andy Way, Qun Liu 0001 |
EAMT | 4 |
| 2016 | Neural Network for Heterogeneous AnnotationsabstractMultiple treebanks annotated under heterogeneous standards give rise to the research question of best utilizing multiple resources for improving statistical models.Prior research has focused on discrete models, leveraging stacking and multi-view learning to address the problem.In this paper, we empirically investigate heterogeneous annotations using neural network models, building a neural network counterpart to discrete stacking and multiview learning, respectively, finding that neural models have their unique advantages thanks to the freedom from manual feature engineering.Neural model achieves not only better accuracy improvements, but also an order of magnitude faster speed compared to its discrete baseline, adding little time cost compared to a neural model trained on a single treebank. Hongshen Chen, Yue Zhang 0004, Qun Liu 0001 |
EMNLP | 3 |
| 2016 | Memory-enhanced Decoder for Neural Machine TranslationabstractWe propose to enhance the RNN decoder in a neural machine translator (NMT) with external memory, as a natural but powerful extension to the state in the decoding RNN.This memory-enhanced RNN decoder is called MEMDEC.At each time during decoding, MEMDEC will read from this memory and write to this memory once, both with content-based addressing.Unlike the unbounded memory in previous work (Bahdanau et al., 2014) to store the representation of source sentence, the memory in MEMDEC is a matrix with predetermined size designed to better capture the information important for the decoding process at each time step.Our empirical study on Chinese-English translation shows that it can improve by 4.8 BLEU upon Groundhog and 5.3 BLEU upon on Moses, yielding the best performance achieved with the same training set. Mingxuan Wang, Zhengdong Lu, Hang Li 0001, Qun Liu 0001 |
EMNLP | 4 |
| 2016 | Variational Neural Discourse Relation RecognizerabstractImplicit discourse relation recognition is a crucial component for automatic discourselevel analysis and nature language understanding.Previous studies exploit discriminative models that are built on either powerful manual features or deep discourse representations.In this paper, instead, we explore generative models and propose a variational neural discourse relation recognizer.We refer to this model as VarNDRR.VarNDRR establishes a directed probabilistic model with a latent continuous variable that generates both a discourse and the relation between the two arguments of the discourse.In order to perform efficient inference and learning, we introduce neural discourse relation models to approximate the prior and posterior distributions of the latent variable, and employ these approximated distributions to optimize a reparameterized variational lower bound.This allows VarNDRR to be trained with standard stochastic gradient methods.Experiments on the benchmark data set show that VarNDRR can achieve comparable results against stateof-the-art baselines without using any manual features. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Qun Liu 0001, Rongrong Ji, Hong Duan, Min Zhang 0005 |
EMNLP | 4 |
| 2016 | Dropped pronoun generation for dialogue machine translationabstractDropped pronoun (DP) is a common problem in dialogue machine translation, in which pronouns are frequently dropped in the source sentence and thus are missing in its translation. In response to this problem, we propose a novel approach to improve the translation of DPs for dialogue machine translation. Firstly, we build a training data for DP generation, in which the DPs are automatically added according to the alignment information from a parallel corpus. Then we model the DP generation problem as a sequence labelling task, and develop a generation model based on recurrent neural networks and language models. Finally, we apply the DP generator to machine translation task by completing the source sentences with the missing pronouns. Experimental results show that our approach achieves a significant improvement of 1.7 BLEU points by recalling possible DPs in the source sentences. Longyue Wang, Zhaopeng Tu, Hang Li 0001, Qun Liu 0001 |
ICASSP | 5 |
| 2016 | Automatic Construction of Discourse Corpora for Dialogue Translation
Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001 |
LREC | 5 |
| 2016 | ProphetMT: A Tree-based SMT-driven Controlled Language Authoring/Post-Editing Tool
Jinhua Du, Qun Liu 0001, Andy Way |
LREC | 3 |
| 2016 | Achieving Accurate Conclusions in Evaluation of Automatic Machine Translation MetricsabstractAutomatic Machine Translation metrics, such as BLEU, are widely used in empirical evaluation as a substitute for human assessment.Subsequently, the performance of a given metric is measured by its strength of correlation with human judgment.When a newly proposed metric achieves a stronger correlation over that of a baseline, it is important to take into account the uncertainty inherent in correlation point estimates prior to concluding improvements in metric performance.Confidence intervals for correlations with human judgment are rarely reported in metric evaluations, however, and when they have been reported, the most suitable methods have unfortunately not been applied.For example, incorrect assumptions about correlation sampling distributions made in past evaluations risk over-estimation of significant differences in metric performance.In this paper, we provide analysis of each of the issues that may lead to inaccuracies before providing detail of a method that overcomes previous challenges.Additionally, we propose a new method of translation sampling that in contrast achieves genuine high conclusivity in evaluation of the relative performance of metrics. Yvette Graham, Qun Liu 0001 |
HLT-NAACL | 2 |
| 2016 | A Novel Approach to Dropped Pronoun TranslationabstractLongyue Wang, Zhaopeng Tu, Xiaojun Zhang, Hang Li, Andy Way, Qun Liu. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Longyue Wang, Zhaopeng Tu, Hang Li 0001, Andy Way, Qun Liu 0001 |
HLT-NAACL | 6 |
| 2016 | Topic-based term translation models for statistical machine translation
Deyi Xiong, Fandong Meng, Qun Liu 0001 |
Artif. Intell. | 3 |
| 2016 | Combining translation memories and statistical machine translation using sparse features
Liangyou Li, Carla Parra Escartín, Andy Way, Qun Liu 0001 |
Mach. Transl. | 4 |
| 2016 | Boosting Neural POS Tagger for Farsi Using Morphological InformationabstractFarsi (Persian) is a low-resource language that suffers from the data sparsity problem and a lack of efficient processing tools. Due to their broad application in natural language processing tasks, part-of-speech (POS) taggers are one of those important tools that should be considered in this respect. Despite recent work on Farsi tagging, there is still room for improvement. The best reported accuracy so far is 96%, which in special cases can rise to 96.9%. The main problem with existing taggers is their inefficiency in coping with out-of-vocabulary (OOV) words. Addressing both problems of accuracy and OOV words, we developed a neural network-based POS tagger (NPT) that performs efficiently on Farsi. Despite using less data, NPT provides better results in comparison to state-of-the-art systems. Our proposed tagger performs with an accuracy of 97.4%, with performance highly influenced by morphological features. We carry out a shallow morphological analysis and show considerable improvement over the baseline configuration. Peyman Passban, Qun Liu 0001, Andy Way |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2016 | A Semisupervised Tag-Transition-Based Markovian Model for Uyghur Morphology AnalysisabstractMorphological analysis, which includes analysis of part-of-speech (POS) tagging, stemming, and morpheme segmentation, is one of the key components in natural language processing (NLP), particularly for agglutinative languages. In this article, we investigate the morphological analysis of the Uyghur language, which is the native language of the people in the Xinjiang Uyghur autonomous region of western China. Morphological analysis of Uyghur is challenging primarily because of factors such as (1) ambiguities arising due to the likelihood of association of a multiple number of POS tags with a word stem or a multiple number of functional tags with a word suffix, (2) ambiguous morpheme boundaries, and (3) complex morphopholonogy of the language. Further, the unavailability of a manually annotated training set in the Uyghur language for the purpose of word segmentation makes Uyghur morphological analysis more difficult. In our proposed work, we address these challenges by undertaking a semisupervised approach of learning a Markov model with the help of a manually constructed dictionary of “suffix to tag” mappings in order to predict the most likely tag transitions in the Uyghur morpheme sequence. Due to the linguistic characteristics of Uyghur, we incorporate a prior belief in our model for favoring word segmentations with a lower number of morpheme units. Empirical evaluation of our proposed model shows an accuracy of about 82%. We further improve the effectiveness of the tag transition model with an active learning paradigm. In particular, we manually investigated a subset of words for which the model prediction ambiguity was within the top 20%. Manually incorporating rules to handle these erroneous cases resulted in an overall accuracy of 93.81%. Eziz Tursun, Debasis Ganguly, Osman Turghun, Yating Yang, Ghalip Abdukerim, Junlin Zhou, Qun Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 7 |
| 2015 | Encoding Source Language with Convolutional Neural Network for Machine TranslationabstractFandong Meng, Zhengdong Lu, Mingxuan Wang, Hang Li, Wenbin Jiang, Qun Liu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Fandong Meng, Zhengdong Lu, Mingxuan Wang, Hang Li 0001, Wenbin Jiang 0002, Qun Liu 0001 |
ACL (1) | 6 |
| 2015 | genCNN: A Convolutional Architecture for Word Sequence PredictionabstractMingxuan Wang, Zhengdong Lu, Hang Li, Wenbin Jiang, Qun Liu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Mingxuan Wang, Zhengdong Lu, Hang Li 0001, Wenbin Jiang 0002, Qun Liu 0001 |
ACL (1) | 5 |
| 2015 | HandyCAT - An Open-Source Platform for CAT Tool Research
Chris Hokamp, Qun Liu 0001 |
EAMT | 2 |
| 2015 | Benchmarking SMT Performance for Farsi Using the TEP++ Corpus
Peyman Passban, Andy Way, Qun Liu 0001 |
EAMT | 3 |
| 2015 | Dependency Graph-to-String TranslationabstractCompared to tree grammars, graph grammars have stronger generative capacity over structures.Based on an edge replacement grammar, in this paper we propose to use a synchronous graph-to-string grammar for statistical machine translation.The graph we use is directly converted from a dependency tree by labelling edges.We build our translation model in the log-linear framework with standard features.Large-scale experiments on Chinese-English and German-English tasks show that our model is significantly better than the state-of-the-art hierarchical phrase-based (HPB) model and a recently improved dependency tree-to-string model on BLEU, METEOR and TER scores.Experiments also suggest that our model has better capability to perform long-distance reordering and is more suitable for translating long sentences. Liangyou Li, Andy Way, Qun Liu 0001 |
EMNLP | 3 |
| 2015 | Joint Learning of Constituency and Dependency Grammars by Decomposed Cross-Lingual Induction
Wenbin Jiang 0002, Qun Liu 0001, Thepchai Supnithi |
IJCAI | 2 |
| 2015 | Syntax-Based Deep Matching of Short Texts
Mingxuan Wang, Zhengdong Lu, Hang Li 0001, Qun Liu 0001 |
IJCAI | 4 |
| 2015 | Bilingual distributed phrase representations for statistical machin translation
Peyman Passban, Chris Hokamp, Qun Liu 0001 |
MTSummit | 3 |
| 2015 | Automatic Adaptation of AnnotationsabstractManually annotated corpora are indispensable resources, yet for many annotation tasks, such as the creation of treebanks, there exist multiple corpora with different and incompatible annotation guidelines. This leads to an inefficient use of human expertise, but it could be remedied by integrating knowledge across corpora with different annotation guidelines. In this article we describe the problem of annotation adaptation and the intrinsic principles of the solutions, and present a series of successively enhanced models that can automatically adapt the divergence between different annotation formats. We evaluate our algorithms on the tasks of Chinese word segmentation and dependency parsing. For word segmentation, where there are no universal segmentation guidelines because of the lack of morphology in Chinese, we perform annotation adaptation from the much larger People's Daily corpus to the smaller but more popular Penn Chinese Treebank. For dependency parsing, we perform annotation adaptation from the Penn Chinese Treebank to a semantics-oriented Dependency Treebank, which is annotated using significantly different annotation guidelines. In both experiments, automatic annotation adaptation brings significant improvement, achieving state-of-the-art performance despite the use of purely local features in training. Wenbin Jiang 0002, Yajuan Lü, Liang Huang 0001, Qun Liu 0001 |
Comput. Linguistics | 4 |
| 2014 | Joint Morphological Generation and Syntactic LinearizationabstractThere has been growing interest in stochastic methods to natural language generation (NLG). While most NLG pipelines separate morphological generation and syntactic linearization, the two tasks are closely related. In this paper, we study joint morphological generation and linearization, making use of word order and inflections information for both tasks and reducing error propagation. Experiments show that the joint method significantly outperforms a strong pipelined baseline (by 1.1 BLEU points). It also achieves the best reported result on the Generation Challenge 2011 shared task. Linfeng Song, Yue Zhang 0004, Qun Liu 0001 |
AAAI | 4 |
| 2014 | A Dependency Edge-based Transfer Model for Statistical Machine Translation
Hongshen Chen, Fandong Meng, Wenbin Jiang 0002, Qun Liu 0001 |
COLING | 5 |
| 2014 | Annotation Adaptation and Language Adaptation in NLP
Qun Liu 0001 |
COLING | 1 |
| 2014 | Augment Dependency-to-String Translation with Fixed and Floating Structures
Jin An Xu, Qun Liu 0001 |
COLING | 3 |
| 2014 | A Structured Language Model for Incremental Tree-to-String Translation
Heng Yu 0006, Haitao Mi, Liang Huang 0001, Qun Liu 0001 |
COLING | 4 |
| 2014 | RED: A Reference Dependency Based MT Evaluation Metric
Hui Yu 0010, Wenbin Jiang 0002, Qun Liu 0001, Shouxun Lin |
COLING | 5 |
| 2014 | Active Learning for Post-Editing Based Incrementally Retrained MTabstractAswarth Abhilash Dara, Josef van Genabith, Qun Liu, John Judge, Antonio Toral. Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers. 2014. Aswarth Abhilash Dara, Josef van Genabith, Qun Liu 0001, John Judge, Antonio Toral |
EACL | 3 |
| 2014 | Modeling Term Translation for Document-informed Machine TranslationabstractTerm translation is of great importance for statistical machine translation (SMT), especially document-informed SMT.In this paper, we investigate three issues of term translation in the context of documentinformed SMT and propose three corresponding models: (a) a term translation disambiguation model which selects desirable translations for terms in the source language with domain information, (b) a term translation consistency model that encourages consistent translations for terms with a high strength of translation consistency throughout a document, and (c) a term bracketing model that rewards translation hypotheses where bracketable source terms are translated as a whole unit.We integrate the three models into hierarchical phrase-based SMT and evaluate their effectiveness on NIST Chinese-English translation tasks with large-scale training data.Experiment results show that all three models can achieve significant improvements over the baseline.Additionally, we can obtain a further improvement when combining the three models. Fandong Meng, Deyi Xiong, Wenbin Jiang 0002, Qun Liu 0001 |
EMNLP | 4 |
| 2014 | Syntactic SMT Using a Discriminative Text Generation ModelabstractWe study a novel architecture for syntactic SMT. In contrast to the dominant approach in the literature, the system does not rely on translation rules, but treat translation as an unconstrained target sentence gen-eration task, using soft features to cap-ture lexical and syntactic correspondences between the source and target languages. Target syntax features and bilingual trans-lation features are trained consistently in a discriminative model. Experiments us-ing the IWSLT 2010 dataset show that the system achieves BLEU comparable to the state-of-the-art syntactic SMT systems. 1 Yue Zhang 0004, Linfeng Song, Qun Liu 0001 |
EMNLP | 5 |
| 2014 | A Novel Rule Refinement Method for SMT through Simulated Post-Editing
Sitong Yang, Heng Yu 0006, Qun Liu 0001 |
NLPCC | 3 |
| 2014 | Topic-Based Dissimilarity and Sensitivity Models for Translation Rule SelectionabstractTranslation rule selection is a task of selecting appropriate translation rules for an ambiguous source-language segment. As translation ambiguities are pervasive in statistical machine translation, we introduce two topic-based models for translation rule selection which incorporates global topic information into translation disambiguation. We associate each synchronous translation rule with source- and target-side topic distributions.With these topic distributions, we propose a topic dissimilarity model to select desirable (less dissimilar) rules by imposing penalties for rules with a large value of dissimilarity of their topic distributions to those of given documents. In order to encourage the use of non-topic specific translation rules, we also present a topic sensitivity model to balance translation rule selection between generic rules and topic-specific rules. Furthermore, we project target-side topic distributions onto the source-side topic model space so that we can benefit from topic information of both the source and target language. We integrate the proposed topic dissimilarity and sensitivity model into hierarchical phrase-based machine translation for synchronous translation rule selection. Experiments show that our topic-based translation rule selection model can substantially improve translation quality. Min Zhang 0005, Xinyan Xiao, Deyi Xiong, Qun Liu 0001 |
J. Artif. Intell. Res. | 4 |
| 2013 | Discriminative Learning with Natural Annotations: Word Segmentation as a Case Study
Wenbin Jiang 0002, Yajuan Lü, Yating Yang, Qun Liu 0001 |
ACL (1) | 5 |
| 2013 | Bilingually-Guided Monolingual Dependency Grammar Induction
Yajuan Lü, Wenbin Jiang 0002, Qun Liu 0001 |
ACL (1) | 4 |
| 2013 | Translation with Source Constituency and Dependency TreesabstractWe present a novel translation model, which simultaneously exploits the constituency and dependency trees on the source side, to combine the advantages of two types of trees.We take head-dependents relations of dependency trees as backbone and incorporate phrasal nodes of constituency trees as the source side of our translation rules, and the target side as strings.Our rules hold the property of long distance reorderings and the compatibility with phrases.Large-scale experimental results show that our model achieves significantly improvements over the constituency-to-string (+2.45 BLEU on average) and dependencyto-string (+0.91 BLEU on average) models, which only employ single type of trees, and significantly outperforms the state-of-theart hierarchical phrase-based model (+1.12BLEU on average), on three Chinese-English NIST test sets. Fandong Meng, Linfeng Song, Yajuan Lü, Qun Liu 0001 |
EMNLP | 5 |
| 2013 | Improving Alignment of System Combination by Using Multi-objective OptimizationabstractThis paper proposes a multi-objective optimization framework which supports heterogeneous information sources to improve alignment in machine translation system combination techniques.In this area, most of techniques usually utilize confusion networks (CN) as their central data structure to compact an exponential number of an potential hypotheses, and because better hypothesis alignment may benefit constructing better quality confusion networks, it is natural to add more useful information to improve alignment results.However, these information may be heterogeneous, so the widely-used Viterbi algorithm for searching the best alignment may not apply here.In the multi-objective optimization framework, each information source is viewed as an independent objective, and a new goal of improving all objectives can be searched by mature algorithms.The solutions from this framework, termed Pareto optimal solutions, are then combined to construct confusion networks.Experiments on two Chinese-to-English translation datasets show significant improvements, 0.97 and 1.06 BLEU points over a strong Indirected Hidden Markov Model-based (IHMM) system, and 4.75 and 3.53 points over the best single machine translation systems. Tian Xia 0004, Zongcheng Ji, Shaodan Zhai, Yidong Chen 0001, Qun Liu 0001 |
EMNLP | 5 |
| 2013 | Modeling Lexical Cohesion for Document-Level Machine Translation
Deyi Xiong, Guosheng Ben, Min Zhang 0005, Yajuan Lü, Qun Liu 0001 |
IJCAI | 5 |
| 2013 | A Topic-Triggered Language Model for Statistical Machine Translation
Heng Yu 0006, Jinsong Su, Yajuan Lü, Qun Liu 0001 |
IJCNLP | 4 |
| 2013 | A Simple, Fast Strategy for Weighted Alignment Hypergraph
Zhaopeng Tu, Yajuan Lü, Qun Liu 0001 |
NLPCC | 4 |
| 2012 | Hierarchical Chunk-to-String Translation
Yang Feng 0004, Dongdong Zhang 0001, Mu Li 0001, Qun Liu 0001 |
ACL (1) | 4 |
| 2012 | Translation Model Adaptation for Statistical Machine Translation with Monolingual Topic Information
Jinsong Su, Hua Wu 0003, Haifeng Wang 0001, Yidong Chen 0001, Xiaodong Shi, Huailin Dong, Qun Liu 0001 |
ACL (1) | 7 |
| 2012 | A Topic Similarity Model for Hierarchical Phrase-based Translation
Xinyan Xiao, Deyi Xiong, Min Zhang 0005, Qun Liu 0001, Shouxun Lin |
ACL (1) | 4 |
| 2012 | Unsupervised Discriminative Induction of Synchronous Grammar for Machine Translation
Xinyan Xiao, Deyi Xiong, Yang Liu 0005, Qun Liu 0001, Shouxun Lin |
COLING | 4 |
| 2012 | Left-to-Right Tree-to-String Decoding with Prediction
Yang Feng 0004, Yang Liu 0005, Qun Liu 0001, Trevor Cohn |
EMNLP-CoNLL | 3 |
| 2012 | Iterative Annotation Transformation with Predict-Self Reestimation for Chinese Word Segmentation
Wenbin Jiang 0002, Fandong Meng, Qun Liu 0001, Yajuan Lü |
EMNLP-CoNLL | 3 |
| 2011 | Adjoining Tree-to-String Translation
Yang Liu 0005, Qun Liu 0001, Yajuan Lü |
ACL | 2 |
| 2011 | Relaxed Cross-lingual Projection of Constituent Syntax
Wenbin Jiang 0002, Qun Liu 0001, Yajuan Lü |
EMNLP | 2 |
| 2011 | Fast Generation of Translation Forest for Large-Scale SMT Discriminative Training
Xinyan Xiao, Yang Liu 0005, Qun Liu 0001, Shouxun Lin |
EMNLP | 3 |
| 2011 | A novel dependency-to-string model for statistical machine translation
Haitao Mi, Qun Liu 0001 |
EMNLP | 3 |
| 2011 | Extracting Hierarchical Rules from a Weighted Alignment Matrix
Zhaopeng Tu, Yang Liu 0005, Qun Liu 0001, Shouxun Lin |
IJCNLP | 3 |
| 2011 | Bagging-based System Combination for Domain Adaption
Linfeng Song, Haitao Mi, Yajuan Lü, Qun Liu 0001 |
MTSummit | 4 |
| 2011 | Multi-granularity Word Alignment and Decoding for Agglutinative Language Translation
Yajuan Lü, Qun Liu 0001 |
MTSummit | 3 |
| 2011 | Maximum Rank Correlation Training for Statistical Machine Translation
Daqi Zheng, Yifan He 0007, Yang Liu 0005, Qun Liu 0001 |
MTSummit | 4 |
| 2011 | Introduction to the Special Issue on Chinese Language Processingabstractintroduction Share on Introduction to the Special Issue on Chinese Language Processing Authors: Keh-Jiann Chen Institute of Information Science, Academia Sinica Institute of Information Science, Academia SinicaView Profile , Qun Liu Institute of Computing Technology, Chinese Academy of Sciences Institute of Computing Technology, Chinese Academy of SciencesView Profile , Nianwen Xue Brandeis University Brandeis UniversityView Profile , Le Sun Institute of Software, Chinese Academy of Sciences Institute of Software, Chinese Academy of SciencesView Profile Authors Info & Claims ACM Transactions on Asian Language Information ProcessingVolume 10Issue 3September 2011 Article No.: 11pp 1–3https://doi.org/10.1145/2002980.2002981Published:01 September 2011Publication History 0citation284DownloadsMetricsTotal Citations0Total Downloads284Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Keh-Jiann Chen, Qun Liu 0001, Nianwen Xue, Le Sun 0001 |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2010 | Forest-Based Semantic Role LabelingabstractParsing plays an important role in semantic role labeling (SRL) because most SRL systems infer semantic relations from 1-best parses. Therefore, parsing errors inevitably lead to labeling mistakes. To alleviate this problem, we propose to use packed forest, which compactly encodes all parses for a sentence. We design an algorithm to exploit exponentially many parses to learn semantic relations efciently. Experimental results on the CoNLL-2005 shared task show that using forests achieves an absolute improvement of 1.2% in terms of F1 score over using 1-best parses and 0.6% over using 50-best parses. Haitao Mi, Yang Liu 0005, Qun Liu 0001 |
AAAI | 4 |
| 2010 | Dependency Parsing and Projection Based on Word-Pair Classification
Wenbin Jiang 0002, Qun Liu 0001 |
ACL | 2 |
| 2010 | Constituency to Dependency Translation with Forests
Haitao Mi, Qun Liu 0001 |
ACL | 2 |
| 2010 | Joint Parsing and Translation
Yang Liu 0005, Qun Liu 0001 |
COLING | 2 |
| 2010 | Dependency Forest for Statistical Machine Translation
Zhaopeng Tu, Yang Liu 0005, Young-Sook Hwang, Qun Liu 0001, Shouxun Lin |
COLING | 4 |
| 2010 | Joint Tokenization and Translation
Xinyan Xiao, Yang Liu 0005, Young-Sook Hwang, Qun Liu 0001, Shouxun Lin |
COLING | 4 |
| 2010 | Statistical Translation Model Based On Source Syntax Structure
Qun Liu 0001, Yang Liu 0005, Haitao Mi |
PACLIC | 1 |
| 2010 | Discriminative Word Alignment by Linear ModelingabstractWord alignment plays an important role in many NLP tasks as it indicates the correspondence between words in a parallel text. Although widely used to align large bilingual corpora, generative models are hard to extend to incorporate arbitrary useful linguistic information. This article presents a discriminative framework for word alignment based on a linear model. Within this framework, all knowledge sources are treated as feature functions, which depend on a source language sentence, a target language sentence, and the alignment between them. We describe a number of features that could produce symmetric alignments. Our model is easy to extend and can be optimized with respect to evaluation metrics directly. The model achieves state-of-the-art alignment quality on three word alignment shared tasks for five language pairs with varying divergence and richness of resources. We further show that our approach improves translation performance for various statistical machine translation systems. Yang Liu 0005, Qun Liu 0001, Shouxun Lin |
Comput. Linguistics | 2 |
| 2009 | Automatic Adaptation of Annotation Standards: Chinese Word Segmentation and POS Tagging - A Case Study
Wenbin Jiang 0002, Liang Huang 0001, Qun Liu 0001 |
ACL/IJCNLP | 3 |
| 2009 | Improving Tree-to-Tree Translation with Packed Forests
Yang Liu 0005, Yajuan Lü, Qun Liu 0001 |
ACL/IJCNLP | 3 |
| 2009 | Joint Decoding with Multiple Translation Models
Yang Liu 0005, Haitao Mi, Yang Feng 0004, Qun Liu 0001 |
ACL/IJCNLP | 4 |
| 2009 | Lattice-based System Combination for Statistical Machine Translation
Yang Feng 0004, Yang Liu 0005, Haitao Mi, Qun Liu 0001, Yajuan Lü |
EMNLP | 4 |
| 2009 | Bilingually-Constrained (Monolingual) Shift-Reduce Parsing
Liang Huang 0001, Wenbin Jiang 0002, Qun Liu 0001 |
EMNLP | 3 |
| 2009 | Weighted Alignment Matrices for Statistical Machine Translation
Yang Liu 0005, Tian Xia 0004, Xinyan Xiao, Qun Liu 0001 |
EMNLP | 4 |
| 2009 | Introduction to China's CWMT2008 Machine Translation Evaluation
Hongmei Zhao, Qun Liu 0001, Yajuan Lü, Dongdong Zhang 0001, Mu Li 0001 |
MTSummit | 3 |
| 2008 | A Cascaded Linear Model for Joint Chinese Word Segmentation and Part-of-Speech Tagging
Wenbin Jiang 0002, Liang Huang 0001, Qun Liu 0001, Yajuan Lü |
ACL | 3 |
| 2008 | Forest-Based Translation
Haitao Mi, Liang Huang 0001, Qun Liu 0001 |
ACL | 3 |
| 2008 | Improving Statistical Machine Translation using Lexicalized Rule Selection
Zhongjun He, Qun Liu 0001, Shouxun Lin |
COLING | 2 |
| 2008 | Word Lattice Reranking for Chinese Word Segmentation and Part-of-Speech Tagging
Wenbin Jiang 0002, Haitao Mi, Qun Liu 0001 |
COLING | 3 |
| 2008 | Maximum Entropy based Rule Selection Model for Syntax-based Statistical Machine Translation
Qun Liu 0001, Zhongjun He, Yang Liu 0005, Shouxun Lin |
EMNLP | 1 |
| 2008 | Fast commercial detection based on audio retrievalabstractAutomatic detection of commercials in digital multimedia material is a challenging task with many applications. This paper presents a novel approach to fast commercial detection based on audio retrieval. It is based on the idea of segmenting energy envelope of audio into units, using only audio signal for matching on a commercial database. Fast searching and matching can be performed with high accuracy, by searching and by novel similarity function based on units. Experimental results show that 96.8% recall rate and 98.7% precision rate can be achieved under 0.125 real-time. Yueliang Qian, Qun Liu 0001, Shouxun Lin |
ICME | 4 |
| 2008 | Refinements in BTG-based Statistical Machine Translation
Deyi Xiong, Min Zhang 0005, AiTi Aw, Haitao Mi, Qun Liu 0001, Shouxun Lin |
IJCNLP | 5 |
| 2007 | Forest-to-String Statistical Translation Rules
Yang Liu 0005, Qun Liu 0001, Shouxun Lin |
ACL | 3 |
| 2007 | Improving Statistical Machine Translation Performance by Training Data Selection and Optimization
Yajuan Lü, Qun Liu 0001 |
EMNLP-CoNLL | 3 |
| 2007 | The PICA Framework for Performance Analysis of Pattern Recognition Systems and Its Application in Broadcast News Segmentation
Meiyin Li, Shouxun Lin, Yueliang Qian, Qun Liu 0001 |
IEA/AIE | 5 |
| 2007 | HTRDP evaluations on Chinese information processing and intelligent human-machine interface
Qun Liu 0001, Hong Liu 0007, Le Sun 0001, Sheng Tang, Deyi Xiong, Hongxu Hou, Yuanhua Lv, Shouxun Lin, Yueliang Qian |
Frontiers Comput. Sci. China | 1 |
| 2006 | Tree-to-String Alignment Template for Statistical Machine TranslationabstractWe present a novel translation model based on tree-to-string alignment template (TAT) which describes the alignment between a source parse tree and a target string. A TAT is capable of generating both terminals and non-terminals and performing reordering at both low and high levels. The model is linguistically syntax-based because TATs are extracted automatically from word-aligned, source side parsed parallel texts. To translate a source sentence, we first employ a parser to produce a source parse tree and then apply TATs to transform the tree into a target string. Our experiments show that the TAT-based model significantly outperforms Pharaoh, a state-of-the-art decoder for phrase-based models. Yang Liu 0005, Qun Liu 0001, Shouxun Lin |
ACL | 2 |
| 2006 | Maximum Entropy Based Phrase Reordering Model for Statistical Machine TranslationabstractWe propose a novel reordering model for phrase-based statistical machine translation (SMT) that uses a maximum entropy (MaxEnt) model to predicate reorderings of neighbor blocks (phrase pairs).The model provides content-dependent, hierarchical phrasal reordering with generalization based on features automatically learned from a real-world bitext.We present an algorithm to extract all reordering events of neighbor blocks from bilingual data.In our experiments on Chineseto-English translation, this MaxEnt-based reordering model obtains significant improvements in BLEU score on the NIST MT-05 and IWSLT-04 tasks. Deyi Xiong, Qun Liu 0001, Shouxun Lin |
ACL | 2 |
| 2006 | An Approximate Analytic Performance Model of Object-Based Storage
Qun Liu 0001, Dan Feng 0001 |
ICCSA (1) | 1 |
| 2006 | Storage challenge - HUSt: a heterogeneous unified storage system for GIS gridabstractGeographic Information System Grid integrates geographic information systems and Grid technology for data gathering, accessing, transmitting and service, in different I/O patterns, built upon massive storage systems. Existing non-standardized multi-source and multi-scale data lack spatial information sharing either internally or externally between organizations or departments, especially in national or global applications. HUSt is a massive storage system that was built at Wuhan National Laboratory for Optoelectronics, in China. There are heterogeneous storage areas in the system, including Object-based Storage System for the main data storing especially for the data searched frequently, Virtual Interface based Storage System for the data required at high transfer speed, and InfiniBand based SAN for high performance. HUSt is primarily meant for research on the organization and key technologies of storage systems for the next generation Internet. The goal is to unify network storage and construct a peta-byte storage system, which supports GIS Grid and applications. Lingfang Zeng, Ke Zhou 0001, Zhan Shi 0001, Dan Feng 0001, Fang Wang 0001, Changsheng Xie 0001, Zhitang Li, Zhanwu Yu, Jianya Gong, Qiang Cao 0001, Zhongying Niu, Lingjun Qin, Qun Liu 0001, Yao Li 0002 |
SC | 13 |
| 2005 | Log-Linear Models for Word AlignmentabstractWe present a framework for word alignment based on log-linear models.All knowledge sources are treated as feature functions, which depend on the source langauge sentence, the target language sentence and possible additional variables.Log-linear models allow statistical alignment models to be easily extended by incorporating syntactic information.In this paper, we use IBM Model 3 alignment probabilities, POS correspondence, and bilingual dictionary coverage as features.Our experiments show that log-linear models significantly outperform IBM translation models. Yang Liu 0005, Qun Liu 0001, Shouxun Lin |
ACL | 2 |
| 2005 | Adaptive Policy Trigger Mechanism for OBSSabstractTraditional storage systems, such as NAS, SAN, are largely unaware of the users and applications actually using the storage, because block-based storage devices manage opaque data blocks. But, with OBSS (object-based storage system), the attributes and methods among the storage devices can be adopted in the storage system, the data can be distributed on some of the storage devices and organized better to anticipate users demand. In this paper, the scalability of object storage (including object attributes, object methods and OBSS) is studied. And a self-managing approach, denoted adaptive policy trigger mechanism (APTM), is presented. APTM borrows proven machine learning techniques and takes the perspective scalable object storage. The implementation reveals that APTM is the embodiment of the idea about smart storage device and facilitates to self-manage mass storage system. Dan Feng 0001, Lingfang Zeng, Fang Wang 0001, Lingjun Qin, Qun Liu 0001 |
AINA | 5 |
| 2005 | Lexicalized Beam Thresholding Parsing with Prior and Boundary Estimates
Deyi Xiong, Qun Liu 0001, Shouxun Lin |
CICLing | 2 |
| 2005 | Parsing the Penn Chinese Treebank with Semantic Knowledge
Deyi Xiong, Shuanglong Li, Qun Liu 0001, Shouxun Lin, Yueliang Qian |
IJCNLP | 3 |
| 2005 | A Multi-aligner for Japanese-Chinese Parallel CorporaabstractAutomatic word alignment is an important technology for extracting translation knowledge from parallel corpora. However, automatic techniques cannot resolve this problem completely because of variances in translations. We therefore need to investigate the performance potential of automatic word alignment and then decide how to suitably apply it. In this paper we first propose a lexical knowledge-based approach to word alignment on a Japanese-Chinese corpus. Then we evaluate the performance of the proposed approach on the corpus. At the same time we also apply a statistics-based approach, the well-known toolkit GIZA++, to the same test data. Through comparison of the performances of the two approaches, we propose a multi-aligner, exploiting the lexical knowledge-based aligner and the statistics-based aligner at the same time. Quantitative results confirmed the effectiveness of the multi-aligner. Qun Liu 0001, Hitoshi Isahara |
MTSummit | 2 |
| 2004 | Tagging Complex NEs with MaxEnt Models: Layered Structures Versus Extended Tagset
Deyi Xiong, Hongkui Yu, Qun Liu 0001 |
IJCNLP | 3 |
| 2002 | Semantic Computation in a Chinese Question-Answering System
Sujian Li, Jian Zhang 0003, Huang Xiong, Shuo Bai, Qun Liu 0001 |
J. Comput. Sci. Technol. | 5 |
| 2001 | Automatic extraction of lexical relations from Chinese machine readable dictionaryabstractLexical relations are very important for NLP. Most previous work to get them is done by hand. In this paper, we describe an automated strategy which exploits a machine readable dictionary (MRD) to construct a richly-structured network of lexical relations. In our system lexical relations include five basic semantic relations, two phonetic relations and one orthographic relation. These relations constitute the basic framework of our lexical network. Then we present an approach to use heuristic functions to extract semantic relations while we conduct syntactic parsing. Experimental results demonstrate that our method is effective. Sujian Li, Qun Liu 0001, Shuo Bai, Xueqi Cheng 0001 |
SMC | 2 |
| 1995 | Efficient realization of frequently used bijections on cube-connected cycles
Qun Liu 0001 |
J. Comput. Sci. Technol. | 2 |