Shuming Shi 0001

dblp:s/ShumingShi-1 · DBLP profile ↗
← Back
119ranked-venue papers
6as first author
56since 2021 · last 2025
0009-0003-1712-5619ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 105 · 5 first-author · 51 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 since 2021Databases, data management, data science and information retrieval · 17 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models
abstract
Abstract The evolution of Neural Machine Translation (NMT) has been significantly influenced by six core challenges (Koehn and Knowles, 2017) that have acted as benchmarks for progress in this field. This study revisits these challenges, offering insights into their ongoing relevance in the context of advanced Large Language Models (LLMs): domain mismatch, amount of parallel data, rare word prediction, translation of long sentences, attention model as word alignment, and sub-optimal beam search. Our empirical findings show that LLMs effectively reduce reliance on parallel data for major languages during pretraining and significantly improve translation of long sentences containing approximately 80 words, even translating documents up to 512 words. Despite these improvements, challenges in domain mismatch and rare word prediction persist. While NMT-specific challenges like word alignment and beam search may not apply to LLMs, we identify three new challenges in LLM-based translation: inference efficiency, translation of low-resource languages during pretraining, and human-aligned evaluation.
Jianhui Pang, Fanghua Ye 0001, Derek F. Wong, Dian Yu 0001, Shuming Shi 0001, Zhaopeng Tu, Longyue Wang
Trans. Assoc. Comput. Linguistics5
2024 Advancement in Graph Understanding: A Multimodal Benchmark and Fine-Tuning of Vision-Language Models
abstract
Qihang Ai, Jiafan Li, Jincheng Dai, Jianwu Zhou, Lemao Liu, Haiyun Jiang, Shuming Shi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Qihang Ai, Jiafan Li, Jincheng Dai, Jianwu Zhou, Lemao Liu, Haiyun Jiang, Shuming Shi 0001
ACL (1)7
2024 MAGE: Machine-generated Text Detection in the Wild
abstract
Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, Yue Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi 0001, Yue Zhang 0004
ACL (1)8
2024 A Frustratingly Simple Decoding Method for Neural Text Generation
abstract
We introduce a frustratingly simple, highly efficient, and surprisingly effective decoding method, termed Frustratingly Simple Decoding (FSD), for neural text generation. The idea behind FSD is straightforward: We construct an anti-language model (anti-LM) based on previously generated text, which is employed to penalize the future generation of repetitive content. The anti-LM can be implemented as simple as an n-gram language model or a vectorized variant. In this way, FSD incurs no additional model parameters and negligible computational overhead (FSD can be as fast as greedy search). Despite its simplicity, FSD is surprisingly effective and generalizes across different datasets, models, and languages. Extensive experiments show that FSD outperforms established strong baselines in terms of generation quality, decoding speed, and universality.
Deng Cai 0002, Wei Bi, Wai Lam, Shuming Shi 0001
LREC/COLING6
2024 On the Cultural Gap in Text-to-Image Generation
abstract
One challenge in text-to-image (T2I) generation is the inadvertent reflection of culture gaps present in the training data, which signifies the disparity in generated image quality when the cultural elements of the input text are rarely collected in the training set. Although various T2I models have shown impressive but arbitrary examples, there is no benchmark to systematically evaluate a T2I model’s ability to generate cross-cultural images. To bridge the gap, we propose a Challenging Cross-Cultural (C3) benchmark with comprehensive evaluation criteria, which can assess how well-suited a model is to a target culture. By analyzing the flawed images generated by the Stable Diffusion model on the C3 benchmark, we find that the model often fails to generate certain cultural objects. Accordingly, we propose a novel multi-modal metric that considers object-text alignment to filter the fine-tuning data in the target culture, which is used to fine-tune a T2I model to improve cross-cultural generation. Experimental results show that our multi-modal metric provides stronger data selection performance on the C3 benchmark than existing metrics, in which the object-text alignment is crucial. We release the benchmark, data, code, and generated images to facilitate future research on culturally diverse T2I generation.
Bingshuai Liu, Longyue Wang, Chenyang Lyu, Yong Zhang 0034, Jinsong Su, Shuming Shi 0001, Zhaopeng Tu
ECAI6
2024 Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
abstract
Modern large language models (LLMs) like ChatGPT have shown remarkable performance on general language tasks but still struggle on complex reasoning tasks, which drives the research on cognitive behaviors of LLMs to explore human-like problem-solving strategies.Along this direction, one representative strategy is self-reflection, which asks an LLM to refine the solution with the feedback generated by itself iteratively.However, our study shows that such reflection-style methods suffer from the Degeneration-of-Thought (DoT) problem: once the LLM has established confidence in its solutions, it is unable to generate novel thoughts later through reflection even if its initial stance is incorrect.To address the DoT problem, we propose a Multi-Agent Debate (MAD) framework, in which multiple agents express their arguments in the state of "tit for tat" and a judge manages the debate process to obtain a final solution.Clearly, our MAD framework encourages divergent thinking in LLMs which would be helpful for tasks that require deep levels of contemplation.Experiment results on two challenging datasets, commonsense machine translation and counterintuitive arithmetic reasoning, demonstrate the effectiveness of our MAD framework.Extensive analyses suggest that the adaptive break of debate and the modest level of "tit for tat" state are required for MAD to obtain good performance.Moreover, we find that LLMs might not be a fair judge if different LLMs are used for agents.Code is available at https://github. com/Skytliang/Multi-Agents-Debate.
Zhiwei He 0002, Wenxiang Jiao, Xing Wang 0007, Yan Wang 0060, Rui Wang 0015, Yujiu Yang 0001, Shuming Shi 0001, Zhaopeng Tu
EMNLP8
2024 Knowledge Verification to Nip Hallucination in the Bud
abstract
While large language models (LLMs) have demonstrated exceptional performance across various tasks following human alignment, they may still generate responses that sound plausible but contradict factual knowledge, a phenomenon known as hallucination.In this paper, we demonstrate the feasibility of mitigating hallucinations by verifying and minimizing the inconsistency between external knowledge present in the alignment data and the intrinsic knowledge embedded within foundation LLMs.Specifically, we propose a novel approach called Knowledge Consistent Alignment (KCA), which employs a well-aligned LLM to automatically formulate assessments based on external knowledge to evaluate the knowledge boundaries of foundation LLMs.To address knowledge inconsistencies in the alignment data, KCA implements several specific strategies to deal with these data instances.We demonstrate the superior efficacy of KCA in reducing hallucinations across six benchmarks, utilizing foundation LLMs of varying backbones and scales.This confirms the effectiveness of mitigating hallucinations by reducing knowledge inconsistency.Our code, model weights, and data are openly accessible at https://github.com/fanqiwan/KCA.* Part of the work was done during his internship at Tencent AI Lab.
Fanqi Wan, Xinting Huang, Leyang Cui, Xiaojun Quan, Wei Bi, Shuming Shi 0001
EMNLP6
2024 SkillNet-X: A Multilingual Multitask Model with Sparsely Activated Skills
abstract
Traditional multitask learning methods typically can only leverage shared knowledge within specific tasks or languages, resulting in a loss of either cross-language or cross-task knowledge. This paper proposes a general multilingual multitask model, named SkillNet-X, which enables a single model to tackle many different tasks from different languages. To this end, we define several language-specific skills and task-specific skills, each of which corresponds to a skill module. SkillNet-X sparsely activates parts of the skill modules which are relevant to eitherthe target task or the target language. Acting as knowledge transit hubs, skill modules are capable of absorbing task-related knowledge and language-related knowledge consecutively. We evaluate SkillNet-X on eleven natural language understanding datasets in four languages. Results show that SkillNet-X performs better than task-specific and two multitask learning baselines.To investigate the generalization of our model, we conduct experiments on two new tasks and find that SkillNet-X significantly outperforms baselines.
Zhangyin Feng, Yong Dai 0001, Fan Zhang 0092, Duyu Tang, Shuangzhi Wu, Bing Qin 0001, Yunbo Cao, Shuming Shi 0001
ICASSP9
2024 Retrieval is Accurate Generation
abstract
Standard language models generate text by selecting tokens from a fixed, finite, and standalone vocabulary. We introduce a novel method that selects context-aware phrases from a collection of supporting documents. One of the most significant challenges for this paradigm shift is determining the training oracles, because a string of text can be segmented in various ways and each segment can be retrieved from numerous possible documents. To address this, we propose to initialize the training oracles using linguistic heuristics and, more importantly, bootstrap the oracles through iterative self-reinforcement. Extensive experiments show that our model not only outperforms standard language models on a variety of knowledge-intensive tasks but also demonstrates improved generation quality in open-ended text generation. For instance, compared to the standard language model counterpart, our model raises the accuracy from 23.47% to 36.27% on OpenbookQA, and improves the MAUVE score from 42.61% to 81.58% in open-ended text generation. Remarkably, our model also achieves the best performance and the lowest latency among several retrieval-augmented baselines. In conclusion, we assert that retrieval is more accurate generation and hope that our work will encourage further research on this new paradigm shift.
Bowen Cao, Deng Cai 0002, Leyang Cui, Xuxin Cheng, Wei Bi, Yuexian Zou, Shuming Shi 0001
ICLR7
2024 The Reasonableness Behind Unreasonable Translation Capability of Large Language Model
abstract
Multilingual large language models trained on non-parallel data yield impressive translation capabilities. Existing studies demonstrate that incidental sentence-level bilingualism within pre-training data contributes to the LLM's translation abilities. However, it has also been observed that LLM's translation capabilities persist even when incidental sentence-level bilingualism are excluded from the training corpus. In this study, we comprehensively investigate the unreasonable effectiveness and the underlying mechanism for LLM's translation abilities, specifically addressing the question why large language models learn to translate without parallel data, using the BLOOM model series as a representative example. Through extensive experiments, our findings suggest the existence of unintentional bilingualism in the pre-training corpus, especially word alignment data significantly contributes to the large language model's acquisition of translation ability. Moreover, the translation signal derived from word alignment data is comparable to that from sentence-level bilingualism. Additionally, we study the effects of monolingual data and parameter-sharing in assisting large language model to learn to translate. Together, these findings present another piece of the broader puzzle of trying to understand how large language models acquire translation capability.
Tingchen Fu, Lemao Liu, Deng Cai 0002, Guoping Huang, Shuming Shi 0001, Rui Yan 0001
ICLR5
2024 Knowledge Fusion of Large Language Models
abstract
While training large language models (LLMs) from scratch can generate models with distinct functionalities and strengths, it comes at significant costs and may result in redundant capabilities. Alternatively, a cost-effective and compelling approach is to merge existing pre-trained LLMs into a more potent model. However, due to the varying architectures of these LLMs, directly blending their weights is impractical. In this paper, we introduce the notion of knowledge fusion for LLMs, aimed at combining the capabilities of existing LLMs and transferring them into a single LLM. By leveraging the generative distributions of source LLMs, we externalize their collective knowledge and unique strengths, thereby potentially elevating the capabilities of the target model beyond those of any individual source LLM. We validate our approach using three popular LLMs with different architectures—Llama-2, MPT, and OpenLLaMA—across various benchmarks and tasks. Our findings confirm that the fusion of LLMs can improve the performance of the target model across a range of capabilities such as reasoning, commonsense, and code generation. Our code, model weights, and data are public at \url{https://github.com/fanqiwan/FuseLLM}.
Fanqi Wan, Xinting Huang, Deng Cai 0002, Xiaojun Quan, Wei Bi, Shuming Shi 0001
ICLR6
2024 GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
abstract
Safety lies at the core of the development of Large Language Models (LLMs). There is ample work on aligning LLMs with human ethics and preferences, including data filtering in pretraining, supervised fine-tuning, reinforcement learning from human feedback, red teaming, etc. In this study, we discover that chat in cipher can bypass the safety alignment techniques of LLMs, which are mainly conducted in natural languages. We propose a novel framework CipherChat to systematically examine the generalizability of safety alignment to non-natural languages -- ciphers. CipherChat enables humans to chat with LLMs through cipher prompts topped with system role descriptions and few-shot enciphered demonstrations. We use CipherChat to assess state-of-the-art LLMs, including ChatGPT and GPT-4 for different representative human ciphers across 11 safety domains in both English and Chinese. Experimental results show that certain ciphers succeed almost 100% of the time in bypassing the safety alignment of GPT-4 in several safety domains, demonstrating the necessity of developing safety alignment for non-natural languages. Notably, we identify that LLMs seem to have a ''secret cipher'', and propose a novel SelfCipher that uses only role play and several unsafe demonstrations in natural language to evoke this capability. SelfCipher surprisingly outperforms existing human ciphers in almost all cases.
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang 0001, Jen-tse Huang 0001, Pinjia He, Shuming Shi 0001, Zhaopeng Tu
ICLR6
2024 GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation
Zhanyu Wang, Longyue Wang, Zhen Zhao 0001, Minghao Wu, Chenyang Lyu, Deng Cai 0002, Luping Zhou, Shuming Shi 0001, Zhaopeng Tu
ACM Multimedia9
2024 Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward Model
abstract
Zhiwei He, Xing Wang, Wenxiang Jiao, Zhuosheng Zhang, Rui Wang, Shuming Shi, Zhaopeng Tu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhiwei He 0002, Xing Wang 0007, Wenxiang Jiao, Zhuosheng Zhang 0001, Rui Wang 0015, Shuming Shi 0001, Zhaopeng Tu
NAACL-HLT6
2024 DoG-Instruct: Towards Premium Instruction-Tuning Data via Text-Grounded Instruction Wrapping
abstract
Yongrui Chen, Haiyun Jiang, Xinting Huang, Shuming Shi, Guilin Qi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yongrui Chen 0002, Haiyun Jiang, Xinting Huang, Shuming Shi 0001, Guilin Qi
NAACL-HLT4
2024 Benchmarking LLMs via Uncertainty Quantification
abstract
The proliferation of open-source Large Language Models (LLMs) from various institutions has highlighted the urgent need for comprehensive evaluation methods. However, current evaluation platforms, such as the widely recognized HuggingFace open LLM leaderboard, neglect a crucial aspect -- uncertainty, which is vital for thoroughly assessing LLMs. To bridge this gap, we introduce a new benchmarking approach for LLMs that integrates uncertainty quantification. Our examination involves nine LLMs (LLM series) spanning five representative natural language processing tasks. Our findings reveal that: I) LLMs with higher accuracy may exhibit lower certainty; II) Larger-scale LLMs may display greater uncertainty compared to their smaller counterparts; and III) Instruction-finetuning tends to increase the uncertainty of LLMs. These results underscore the significance of incorporating uncertainty in the evaluation of LLMs. Our implementation is available at https://github.com/smartyfh/LLM-Uncertainty-Bench.
Fanghua Ye 0001, Jianhui Pang, Longyue Wang, Derek F. Wong, Emine Yilmaz, Shuming Shi 0001, Zhaopeng Tu
NeurIPS7
2024 StrategyLLM: Large Language Models as Strategy Generators, Executors, Optimizers, and Evaluators for Problem Solving
abstract
Most existing prompting methods suffer from the issues of generalizability and consistency, as they often rely on instance-specific solutions that may not be applicable to other instances and lack task-level consistency across the selected few-shot examples. To address these limitations, we propose a comprehensive framework, StrategyLLM, allowing LLMs to perform inductive reasoning, deriving general strategies from specific task instances, and deductive reasoning, applying these general strategies to particular task examples, for constructing generalizable and consistent few-shot prompts. It employs four LLM-based agents: strategy generator, executor, optimizer, and evaluator, working together to generate, evaluate, and select promising strategies for a given task. Experimental results demonstrate that StrategyLLM outperforms the competitive baseline CoT-SC that requires human-annotated solutions on 13 datasets across 4 challenging tasks without human involvement, including math reasoning (34.2\% $\rightarrow$ 38.8\%), commonsense reasoning (70.3\% $\rightarrow$ 72.5\%), algorithmic reasoning (73.7\% $\rightarrow$ 85.0\%), and symbolic reasoning (30.0\% $\rightarrow$ 79.2\%). Further analysis reveals that StrategyLLM is applicable to various LLMs and demonstrates advantages across numerous scenarios.
Haiyun Jiang, Deng Cai 0002, Shuming Shi 0001, Wai Lam
NeurIPS4
2024 Exploring Human-Like Translation Strategy with Large Language Models
abstract
Abstract Large language models (LLMs) have demonstrated impressive capabilities in general scenarios, exhibiting a level of aptitude that approaches, in some aspects even surpasses, human-level intelligence. Among their numerous skills, the translation abilities of LLMs have received considerable attention. Compared to typical machine translation that focuses solely on source-to-target mapping, LLM-based translation can potentially mimic the human translation process, which might take preparatory steps to ensure high-quality translation. This work explores this possibility by proposing the MAPS framework, which stands for Multi-Aspect Prompting and Selection. Specifically, we enable LLMs first to analyze the given source sentence and induce three aspects of translation-related knowledge (keywords, topics, and relevant demonstrations) to guide the final translation process. Moreover, we employ a selection mechanism based on quality estimation to filter out noisy and unhelpful knowledge. Both automatic (3 LLMs × 11 directions × 2 automatic metrics) and human evaluation (preference study and MQM) demonstrate the effectiveness of MAPS. Further analysis shows that by mimicking the human translation process, MAPS reduces various translation errors such as hallucination, ambiguity, mistranslation, awkward style, untranslated text, and omission. Source code is available at https://github.com/zwhe99/MAPS-mt.
Zhiwei He 0002, Wenxiang Jiao, Zhuosheng Zhang 0001, Yujiu Yang 0001, Rui Wang 0015, Zhaopeng Tu, Shuming Shi 0001, Xing Wang 0007
Trans. Assoc. Comput. Linguistics8
2024 An Energy-based Model for Word-level AutoCompletion in Computer-aided Translation
abstract
Abstract Word-level AutoCompletion (WLAC) is a rewarding yet challenging task in Computer-aided Translation. Existing work addresses this task through a classification model based on a neural network that maps the hidden vector of the input context into its corresponding label (i.e., the candidate target word is treated as a label). Since the context hidden vector itself does not take the label into account and it is projected to the label through a linear classifier, the model cannot sufficiently leverage valuable information from the source sentence as verified in our experiments, which eventually hinders its overall performance. To alleviate this issue, this work proposes an energy-based model for WLAC, which enables the context hidden vector to capture crucial information from the source sentence. Unfortunately, training and inference suffer from efficiency and effectiveness challenges, therefore we employ three simple yet effective strategies to put our model into practice. Experiments on four standard benchmarks demonstrate that our reranking-based approach achieves substantial improvements (about 6.07%) over the previous state-of-the-art model. Further analyses show that each strategy of our approach contributes to the final performance.1
Cheng Yang 0007, Guoping Huang, Mo Yu, Zhirui Zhang, Siheng Li, Shuming Shi 0001, Yujiu Yang 0001, Lemao Liu
Trans. Assoc. Comput. Linguistics7
2023 Zero-Shot Rumor Detection with Propagation Structure via Prompt Learning
abstract
The spread of rumors along with breaking events seriously hinders the truth in the era of social media. Previous studies reveal that due to the lack of annotated resources, rumors presented in minority languages are hard to be detected. Furthermore, the unforeseen breaking events not involved in yesterday's news exacerbate the scarcity of data resources. In this work, we propose a novel zero-shot framework based on prompt learning to detect rumors falling in different domains or presented in different languages. More specifically, we firstly represent rumor circulated on social media as diverse propagation threads, then design a hierarchical prompt encoding mechanism to learn language-agnostic contextual representations for both prompts and rumor data. To further enhance domain adaptation, we model the domain-invariant structural features from the propagation threads, to incorporate structural position representations of influential community response. In addition, a new virtual response augmentation method is used to improve model training. Extensive experiments conducted on three real-world datasets demonstrate that our proposed model achieves much better performance than state-of-the-art methods and exhibits a superior capacity for detecting rumors at early stages.
Hongzhan Lin 0001, Pengyao Yi, Jing Ma 0004, Haiyun Jiang, Shuming Shi 0001, Ruifang Liu
AAAI6
2023 Enhancing Grammatical Error Correction Systems with Explanations
abstract
Grammatical error correction systems improve written communication by detecting and correcting language mistakes.To help language learners better understand why the GEC system makes a certain correction, the causes of errors (evidence words) and the corresponding error types are two key factors.To enhance GEC systems with explanations, we introduce EXPECT, a large dataset annotated with evidence words and grammatical error types.We propose several baselines and analysis to understand this task.Furthermore, human evaluation verifies our explainable GEC system's explanations can assist second-language learners in determining whether to accept a correction suggestion and in understanding the associated grammar rule.
Yuejiao Fei, Leyang Cui, Sen Yang 0005, Wai Lam, Zhen-Zhong Lan, Shuming Shi 0001
ACL (1)6
2023 Explicit Syntactic Guidance for Neural Text Generation
abstract
Most existing text generation models follow the sequence-to-sequence paradigm.Generative Grammar suggests that humans generate natural language texts by learning language grammar.We propose a syntax-guided generation schema, which generates the sequence guided by a constituency parse tree in a topdown direction.The decoding process can be decomposed into two parts: (1) predicting the infilling texts for each constituent in the lexicalized syntax context given the source sentence;(2) mapping and expanding each constituent to construct the next-level syntax context.Accordingly, we propose a structural beam search method to find possible syntax structures hierarchically.Experiments on paraphrase generation and machine translation show that the proposed method outperforms autoregressive baselines, while also demonstrating effectiveness in terms of interpretability, controllability, and diversity.
Yafu Li, Leyang Cui, Jianhao Yan, Yongjing Yin, Wei Bi, Shuming Shi 0001, Yue Zhang 0004
ACL (1)6
2023 A Survey on Zero Pronoun Translation
abstract
Zero pronouns (ZPs) are frequently omitted in pro-drop languages (e.g.Chinese, Hungarian, and Hindi), but should be recalled in nonpro-drop languages (e.g.English).This phenomenon has been studied extensively in machine translation (MT), as it poses a significant challenge for MT systems due to the difficulty in determining the correct antecedent for the pronoun.This survey paper highlights the major works that have been undertaken in zero pronoun translation (ZPT) after the neural revolution so that researchers can recognize the current state and future directions of this field.We provide an organization of the literature based on evolution, dataset, method, and evaluation.In addition, we compare and analyze competing models and evaluation metrics on different benchmarks.We uncover a number of insightful findings such as: 1) ZPT is in line with the development trend of large language model; 2) data limitation causes learning bias in languages and domains; 3) performance improvements are often reported on single benchmarks, but advanced methods are still far from realworld use; 4) general-purpose metrics are not reliable on nuances and complexities of ZPT, emphasizing the necessity of targeted metrics; 5) apart from commonly-cited errors, ZPs will cause risks of gender bias.
Longyue Wang, Siyou Liu, Mingzhou Xu, Linfeng Song, Shuming Shi 0001, Zhaopeng Tu
ACL (1)5
2023 RobustGEC: Robust Grammatical Error Correction Against Subtle Context Perturbation
abstract
Grammatical Error Correction (GEC) systems play a vital role in assisting people with their daily writing tasks.However, users may sometimes come across a GEC system that initially performs well but fails to correct errors when the inputs are slightly modified.To ensure an ideal user experience, a reliable GEC system should have the ability to provide consistent and accurate suggestions when encountering irrelevant context perturbations, which we refer to as context robustness.In this paper, we introduce RobustGEC, a benchmark designed to evaluate the context robustness of GEC systems.RobustGEC comprises 5,000 GEC cases, each with one original error-correct sentence pair and five variants carefully devised by human annotators.Utilizing RobustGEC, we reveal that state-of-the-art GEC systems still lack sufficient robustness against context perturbations.In addition, we propose a simple yet effective method for remitting this issue.
Yue Zhang 0004, Leyang Cui, Enbo Zhao, Wei Bi, Shuming Shi 0001
EMNLP5
2023 Rethinking Word-Level Auto-Completion in Computer-Aided Translation
abstract
Word-Level Auto-Completion (WLAC) plays a crucial role in Computer-Assisted Translation.It aims at providing word-level autocompletion suggestions for human translators.While previous studies have primarily focused on designing complex model architectures, this paper takes a different perspective by rethinking the fundamental question: what kind of words are good auto-completions?We introduce a measurable criterion to answer this question and discover that existing WLAC models often fail to meet this criterion.Building upon this observation, we propose an effective approach to enhance WLAC performance by promoting adherence to the criterion.Notably, the proposed approach is general and can be applied to various encoder-based architectures.Through extensive experiments, we demonstrate that our approach outperforms the top-performing system submitted to the WLAC shared tasks in WMT2022, while utilizing significantly smaller model sizes ¶ .
Lemao Liu, Guoping Huang, Zhirui Zhang, Shuming Shi 0001, Rui Wang 0015
EMNLP6
2023 IMTLab: An Open-Source Platform for Building, Evaluating, and Diagnosing Interactive Machine Translation Systems
abstract
Xu Huang, Zhirui Zhang, Ruize Gao, Yichao Du, Lemao Liu, Guoping Huang, Shuming Shi, Jiajun Chen, Shujian Huang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Zhirui Zhang, Yichao Du, Lemao Liu, Guoping Huang, Shuming Shi 0001, Jiajun Chen 0001, Shujian Huang
EMNLP7
2023 Explore-Instruct: Enhancing Domain-Specific Instruction Coverage through Active Exploration
abstract
Instruction-tuning can be substantially optimized through enhanced diversity, resulting in models capable of handling a broader spectrum of tasks.However, existing data employed for such tuning often exhibit an inadequate coverage of individual domains, limiting the scope for nuanced comprehension and interactions within these areas.To address this deficiency, we propose EXPLORE-INSTRUCT, a novel approach to enhance the data coverage to be used in domain-specific instruction-tuning through active exploration via Large Language Models (LLMs).Built upon representative domain use cases, EXPLORE-INSTRUCT explores a multitude of variations or possibilities by implementing a search algorithm to obtain diversified and domain-focused instruction-tuning data.Our data-centric analysis validates the effectiveness of this proposed approach in improving domain-specific instruction coverage.Moreover, our model's performance demonstrates considerable advancements over multiple baselines, including those utilizing domainspecific data enhancement.Our findings offer a promising opportunity to improve instruction coverage, especially in domain-specific contexts, thereby advancing the development of adaptable language models.Our code, model weights, and data are public at https:// github.com/fanqiwan/Explore-Instruct.
Fanqi Wan, Xinting Huang, Tao Yang 0033, Xiaojun Quan, Wei Bi, Shuming Shi 0001
EMNLP6
2023 Document-Level Machine Translation with Large Language Models
abstract
Large language models (LLMs) such as Chat-GPT can produce coherent, cohesive, relevant, and fluent answers for various natural language processing (NLP) tasks.Taking documentlevel machine translation (MT) as a testbed, this paper provides an in-depth evaluation of LLMs' ability on discourse modeling.The study focuses on three aspects: 1) Effects of Context-Aware Prompts, where we investigate the impact of different prompts on document-level translation quality and discourse phenomena; 2) Comparison of Translation Models, where we compare the translation performance of Chat-GPT with commercial MT systems and advanced document-level MT methods; 3) Analysis of Discourse Modelling Abilities, where we further probe discourse knowledge encoded in LLMs and shed light on impacts of training techniques on discourse modeling.By evaluating on a number of benchmarks, we surprisingly find that LLMs have demonstrated superior performance and show potential to become a new paradigm for document-level translation: 1) leveraging their powerful long-text modeling capabilities, GPT-3.5 and GPT-4 outperform commercial MT systems in terms of human evaluation; 1 2) GPT-4 demonstrates a stronger ability for probing linguistic knowledge than GPT-3.5.This work highlights the challenges and opportunities of LLMs for MT, which we hope can inspire the future design and evaluation of LLMs. 2 * Equal contribution.
Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu 0001, Shuming Shi 0001, Zhaopeng Tu
EMNLP6
2023 Skillnet-NLG: General-Purpose Natural Language Generation with a Sparsely Activated Approach
abstract
We present SkillNet-NLG, a sparsely activated approach that handles many natural language generation tasks with one model. Different from traditional dense models that always activate all the parameters, SkillNet-NLG selectively activates relevant parts of the parameters to accomplish a task, where the relevance is controlled by a set of predefined skills. The strength of such model design is that it provides an opportunity to precisely adapt relevant skills to learn new tasks effectively. We evaluate on Chinese natural language generation tasks. Results show that, with only one model file, SkillNet-NLG outperforms previous best performance methods on four of five tasks. SkillNet-NLG performs better than two multitask learning baselines (a dense model and a Mixture-of-Expert model) and achieves comparable performance to task-specific models. Lastly, SkillNet-NLG surpasses baseline systems when adapted to new tasks.
Junwei Liao, Duyu Tang, Fan Zhang 0092, Shuming Shi 0001
ICASSP4
2023 A Simple Yet Effective Approach to Structured Knowledge Distillation
abstract
Structured prediction models aim at solving tasks where the output is a complex structure, rather than a single variable. Performing knowledge distillation for such problems is non- trivial due to their exponentially large output space. Previous works address this problem by developing particular distillation strategies (e.g., dynamic programming) that are both complicated and of low run-time efficiency. In this work, we propose an approach that is much simpler in its formulation, far more efficient for training than existing methods, and even performs better than our baselines. Specifically, we transfer the knowledge from a teacher model to its student by locally matching their computations on all internal structures rather than the final outputs. In this manner, we avoid time-consuming techniques like Monte Carlo Sampling for decoding output structures, permitting parallel computation and efficient training. Besides, we show that it encourages the student model to better mimic the internal behavior of the teacher model. Experiments on two structured prediction tasks demonstrate that our approach not only halves the time cost, but also outperforms previous methods on two widely adopted benchmark datasets.1 2
Wenye Lin, Yangming Li, Lemao Liu, Shuming Shi 0001, Hai-Tao Zheng 0002
ICASSP4
2023 MarkBERT: Marking Word Boundaries Improves Chinese BERT
Linyang Li, Yong Dai 0001, Duyu Tang, Xipeng Qiu, Shuming Shi 0001
NLPCC (1)6
2023 Predicting Events in MOBA Games: Prediction, Attribution, and Evaluation
abstract
The multiplayer online battle arena (MOBA) games have become increasingly popular in recent years. Consequently, many efforts have been devoted to providing pregame or in-game predictions for them. These predictions can be used in many MOBA esports-related applications, such as artificial intelligence commentator systems, in-game data analysis, and game-assistant bots. However, these works are limited in the following two aspects: the lack of sufficient in-game features and the absence of interpretability in the prediction results. These two limitations greatly restrict the practical performance and industrial application of the current works. In this work, we collect a large-scale dataset containing rich in-game features for the popular MOBA gameHonor of Kings. We then propose to predict four types of prediction tasks in an interpretable way by attributing the predictions to the input features using two gradient-based attribution methods:Integrated GradientsandSmoothGrad. To evaluate the explanatory power of different models and attribution methods, a fidelity-based evaluation metric is further proposed. Finally, we evaluate the accuracy and fidelity of several competitive methods to assess how well machines predict events in MOBA games.
Zelong Yang 0002, Yan Wang 0060, Piji Li, Shaobin Lin, Shuming Shi 0001, Shao-Lun Huang, Wei Bi
IEEE Trans. Games5
2022 Redistributing Low-Frequency Words: Making the Most of Monolingual Data in Non-Autoregressive Translation
abstract
Knowledge distillation (KD) is the preliminary step for training non-autoregressive translation (NAT) models, which eases the training of NAT models at the cost of losing important information for translating low-frequency words.In this work, we provide an appealing alternative for NAT -monolingual KD, which trains NAT student on external monolingual data with AT teacher trained on the original bilingual data.Monolingual KD is able to transfer both the knowledge of the original bilingual data (implicitly encoded in the trained AT teacher model) and that of the new monolingual data to the NAT student model.Extensive experiments on eight WMT benchmarks over two advanced NAT models show that monolingual KD consistently outperforms the standard KD by improving lowfrequency word translation, without introducing any computational cost.Monolingual KD enjoys desirable expandability, which can be further enhanced (when given more computational budget) by combining with the standard KD, a reverse monolingual KD, or enlarging the scale of monolingual data.Extensive analyses demonstrate that these techniques can be used together profitably to further recall the useful information lost in the standard KD.Encouragingly, combining with standard KD, our approach achieves 30.4 and 34.1 BLEU points on the WMT14 English-German and German-English datasets, respectively.Our code and trained models are freely available at https://github.com/ alphadl/RLFW-NAT.mono.
Liang Ding 0006, Longyue Wang, Shuming Shi 0001, Dacheng Tao, Zhaopeng Tu
ACL (1)3
2022 Learning from Sibling Mentions with Scalable Graph Inference in Fine-Grained Entity Typing
abstract
Yi Chen, Jiayang Cheng, Haiyun Jiang, Lemao Liu, Haisong Zhang, Shuming Shi, Ruifeng Xu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yi Chen 0019, Cheng Jiayang, Haiyun Jiang, Lemao Liu, Haisong Zhang, Shuming Shi 0001, Ruifeng Xu 0001
ACL (1)6
2022 A Model-agnostic Data Manipulation Method for Persona-based Dialogue Generation
abstract
Towards building intelligent dialogue agents, there has been a growing interest in introducing explicit personas in generation models.However, with limited persona-based dialogue data at hand, it may be difficult to train a dialogue generation model well.We point out that the data challenges of this generation task lie in two aspects: first, it is expensive to scale up current persona-based dialogue datasets; second, each data sample in this task is more complex to learn with than conventional dialogue data.To alleviate the above data issues, we propose a data manipulation method, which is model-agnostic to be packed with any personabased dialogue generation model to improve its performance.The original training samples will first be distilled and thus expected to be fitted more easily.Next, we show various effective ways that can diversify such easier distilled data.A given base model will then be trained via the constructed data curricula, i.e. first on augmented distilled samples and then on original ones.Experiments illustrate the superiority of our method with two strong base dialogue models (Transformer encoderdecoder and GPT2).
Yu Cao 0014, Wei Bi, Shuming Shi 0001, Dacheng Tao
ACL (1)4
2022 Bridging the Data Gap between Training and Inference for Unsupervised Neural Machine Translation
abstract
Back-translation is a critical component of Unsupervised Neural Machine Translation (UNMT), which generates pseudo parallel data from target monolingual data.A UNMT model is trained on the pseudo parallel data with translated source, and translates natural source sentences in inference.The source discrepancy between training and inference hinders the translation performance of UNMT models.By carefully designing experiments, we identify two representative characteristics of the data gap in source: (1) style gap (i.e., translated vs. natural text style) that leads to poor generalization capability; (2) content gap that induces the model to produce hallucination content biased towards the target language.To narrow the data gap, we propose an online self-training approach, which simultaneously uses the pseudo parallel data {natural source, translated target} to mimic the inference scenario.Experimental results on several widelyused language pairs show that our approach outperforms two strong baselines (XLM and MASS) by remedying the style and content gaps. 1 Model En-Fr En-De En-Ro Avg.⇒ ⇐ ⇒ ⇐ ⇒ ⇐ Full Test Set SNMT 38.4 33.6 29.5 33.9 33.7 32.5 33.6 XLM 37.4 34.5 27.2 34.3 34.6 32.7 33.5 MASS 37.8 34.9 27.1 35.2 35.1 33.4 33.9 Model En-Fr En-De En-Ro Avg.⇒ ⇐ ⇒ ⇐ ⇒ ⇐ Full Test Set SNMT 37.3 33.4 29.7 33.8 33.8 32.4 33.4 XLM 36.3 34.3 27.4 34.1 34.8 32.4 33.2 MASS 36.6 34.7 27.3 35.1 35.2 33.0 33.7
Zhiwei He 0002, Xing Wang 0007, Rui Wang 0015, Shuming Shi 0001, Zhaopeng Tu
ACL (1)4
2022 Rethinking Negative Sampling for Handling Missing Entity Annotations
abstract
Negative sampling is highly effective in handling missing annotations for named entity recognition (NER).One of our contributions is an analysis on how it makes sense through introducing two insightful concepts: missampling and uncertainty.Empirical studies show low missampling rate and high uncertainty are both essential for achieving promising performances with negative sampling.Based on the sparsity of named entities, we also theoretically derive a lower bound for the probability of zero missampling rate, which is only relevant to sentence length.The other contribution is an adaptive and weighted sampling distribution that further improves negative sampling via our former analysis.Experiments on synthetic datasets and well-annotated datasets (e.g., CoNLL-2003) show that our proposed approach benefits negative sampling in terms of F1 score and loss convergence.Besides, models with improved negative sampling have achieved new state-of-the-art results on realworld datasets (e.g., EC).
Yangming Li, Lemao Liu, Shuming Shi 0001
ACL (1)3
2022 Exploring and Adapting Chinese GPT to Pinyin Input Method
abstract
Minghuan Tan, Yong Dai, Duyu Tang, Zhangyin Feng, Guoping Huang, Jing Jiang, Jiwei Li, Shuming Shi. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Minghuan Tan, Yong Dai 0001, Duyu Tang, Zhangyin Feng, Guoping Huang, Jing Jiang 0001, Shuming Shi 0001
ACL (1)8
2022 Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine Translation
abstract
Wenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang, Shuming Shi, Zhaopeng Tu, Michael Lyu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Wenxuan Wang 0001, Wenxiang Jiao, Yongchang Hao, Xing Wang 0007, Shuming Shi 0001, Zhaopeng Tu, Michael R. Lyu
ACL (1)5
2022 BiTIIMT: A Bilingual Text-infilling Method for Interactive Machine Translation
abstract
Yanling Xiao, Lemao Liu, Guoping Huang, Qu Cui, Shujian Huang, Shuming Shi, Jiajun Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yanling Xiao, Lemao Liu, Guoping Huang, Qu Cui, Shujian Huang, Shuming Shi 0001, Jiajun Chen 0001
ACL (1)6
2022 On the Evaluation Metrics for Paraphrase Generation
abstract
In this paper we revisit automatic metrics for paraphrase evaluation and obtain two findings that disobey conventional wisdom:(1) Reference-free metrics achieve better performance than their reference-based counterparts.(2) Most commonly used metrics do not align well with human annotation.Underlying reasons behind the above findings are explored through additional experiments and in-depth analyses.Based on the experiments and analyses, we propose ParaScore, a new evaluation metric for paraphrase generation.It possesses the merits of referencebased and reference-free metrics and explicitly models lexical divergence.Based on our analysis and improvements, our proposed reference-based outperforms than referencefree metrics.Experimental results demonstrate that ParaScore significantly outperforms existing metrics.Our codes and toolkit are released in https://github.com/ shadowkiller33/ParaScore.
Lingfeng Shen, Lemao Liu, Haiyun Jiang, Shuming Shi 0001
EMNLP4
2022 GuoFeng: A Benchmark for Zero Pronoun Recovery and Translation
abstract
Mingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu, Linfeng Song, Lidia S. Chao, Shuming Shi, Zhaopeng Tu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Mingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu, Linfeng Song, Lidia S. Chao, Shuming Shi 0001, Zhaopeng Tu
EMNLP7
2022 Towards Efficient Dialogue Pre-training with Transferable and Interpretable Latent Structure
abstract
With the availability of massive generaldomain dialogue data, pre-trained dialogue generation appears to be super appealing to transfer knowledge from the general domain to downstream applications.In most existing work, such transferable ability is mainly obtained by fitting a large model with hundreds of millions of parameters on massive data in an exhaustive way, leading to inefficient running and poor interpretability.This paper proposes a novel dialogue generation model with a latent structure that is easily transferable from the general domain to downstream tasks in a lightweight and transparent way.Experiments on two benchmarks validate the effectiveness of the proposed model.Thanks to the transferable latent structure, our model is able to yield better dialogue responses than four strong baselines in terms of both automatic and human evaluations, and our model with about 22% parameters particularly delivers a 5x speedup in running time compared with the strongest baseline.Moreover, the proposed model is explainable by interpreting the discrete latent variables.
Xueliang Zhao, Lemao Liu, Tingchen Fu, Shuming Shi 0001, Dongyan Zhao 0001, Rui Yan 0001
EMNLP4
2022 On Synthetic Data for Back Translation
abstract
Jiahao Xu, Yubin Ruan, Wei Bi, Guoping Huang, Shuming Shi, Lihui Chen, Lemao Liu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jiahao Xu 0001, Yubin Ruan, Wei Bi, Guoping Huang, Shuming Shi 0001, Lihui Chen 0001, Lemao Liu
NAACL-HLT5
2022 Recent Advances in Retrieval-Augmented Text Generation
abstract
Recently retrieval-augmented text generation has achieved state-of-the-art performance in many NLP tasks and has attracted increasing attention of the NLP and IR community, this tutorial thereby aims to present recent advances in retrieval-augmented text generation comprehensively and comparatively. It firstly highlights the generic paradigm of retrieval-augmented text generation, then reviews notable works for different text generation tasks including dialogue generation, machine translation, and other generation tasks, and finally points out some limitations and shortcomings to facilitate future research.
Deng Cai 0002, Yan Wang 0060, Lemao Liu, Shuming Shi 0001
SIGIR4
2022 Interpretable Real-Time Win Prediction for Honor of Kings - A Popular Mobile MOBA Esport
abstract
With the rapid prevalence and explosive development of Multiplayer Online Battle Arena electronic sports (MOBA esports), much research effort has been devoted to automatically predicting game results (win predictions). While this task has great potential in various applications, such as esports live streaming and game commentator artificial intelligence systems, previous studies fail to investigate the methods tointerpretthese win predictions. To mitigate this issue, we collected a large-scale dataset that contains real-time game records with rich input features of the popular MOBA gameHonor of Kings. For interpretable predictions, we proposed a two-stage spatial–temporal network (TSSTN) that can not only provide accurate real-time win predictions but also attribute the ultimate prediction results to the contributions of different features for interpretability. Experiment results and applications in real-world live streaming scenarios showed that the proposed TSSTN model is effective in both prediction accuracy and interpretability.
Zelong Yang 0002, Zhufeng Pan, Yan Wang 0060, Deng Cai 0002, Shuming Shi 0001, Shao-Lun Huang, Wei Bi, Xiaojiang Liu
IEEE Trans. Games5
2021 Self-Training Sampling with Monolingual Data Uncertainty for Neural Machine Translation
abstract
Wenxiang Jiao, Xing Wang, Zhaopeng Tu, Shuming Shi, Michael Lyu, Irwin King. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wenxiang Jiao, Xing Wang 0007, Zhaopeng Tu, Shuming Shi 0001, Michael R. Lyu, Irwin King
ACL/IJCNLP (1)4
2021 Tail-to-Tail Non-Autoregressive Sequence Prediction for Chinese Grammatical Error Correction
abstract
Piji Li, Shuming Shi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Piji Li, Shuming Shi 0001
ACL/IJCNLP (1)2
2021 GWLAN: General Word-Level AutocompletioN for Computer-Aided Translation
abstract
Huayang Li, Lemao Liu, Guoping Huang, Shuming Shi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Lemao Liu, Guoping Huang, Shuming Shi 0001
ACL/IJCNLP (1)4
2021 Dialogue Response Selection with Hierarchical Curriculum Learning
abstract
Yixuan Su, Deng Cai, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi, Nigel Collier, Yan Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yixuan Su, Deng Cai 0002, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi 0001, Nigel Collier, Yan Wang 0060
ACL/IJCNLP (1)7
2021 An Empirical Study on Multiple Information Sources for Zero-Shot Fine-Grained Entity Typing
abstract
Auxiliary information from multiple sources has been demonstrated to be effective in zeroshot fine-grained entity typing (ZFET).However, there lacks a comprehensive understanding about how to make better use of the existing information sources and how they affect the performance of ZFET.In this paper, we empirically study three kinds of auxiliary information: context consistency, type hierarchy and background knowledge (e.g., prototypes and descriptions) of types, and propose a multi-source fusion model (MSF) targeting these sources.The performance obtains up to 11.42% and 22.84% absolute gains over stateof-the-art baselines on BBN and Wiki respectively with regard to macro F1 scores.More importantly, we further discuss the characteristics, merits and demerits of each information source and provide an intuitive understanding of the complementarity among them.
Yi Chen 0019, Haiyun Jiang, Lemao Liu, Shuming Shi 0001, Chuang Fan, Min Yang 0007, Ruifeng Xu 0001
EMNLP (1)4
2021 Fine-grained Entity Typing without Knowledge Base
abstract
Existing work on Fine-grained Entity Typing (FET) typically trains automatic models on the datasets obtained by using Knowledge Bases (KB) as distant supervision.However, the reliance on KB means this training setting can be hampered by the lack of or the incompleteness of the KB.To alleviate this limitation, we propose a novel setting for training FET models: FET without accessing any knowledge base.Under this setting, we propose a two-step framework to train FET models.In the first step, we automatically create pseudo data with fine-grained labels from a large unlabeled dataset.Then a neural network model is trained based on the pseudo data, either in an unsupervised way or using self-training under the weak guidance from a coarse-grained Named Entity Recognition (NER) model.Experimental results show that our method achieves competitive performance with respect to the models trained on the original KB-supervised datasets.* The first two authors (Jing and Yibin) contributed equally to this work during the internships at Tencent AI Lab.
Lemao Liu, Yangming Li, Haiyun Jiang, Haisong Zhang, Shuming Shi 0001
EMNLP (1)7
2021 Empirical Analysis of Unlabeled Entity Problem in Named Entity Recognition
Yangming Li, Lemao Liu, Shuming Shi 0001
ICLR3
2021 Context-aware Self-Attention Networks for Natural Language Processing
Baosong Yang, Longyue Wang, Derek F. Wong, Shuming Shi 0001, Zhaopeng Tu
Neurocomputing4
2021 Attending From Foresight: A Novel Attention Mechanism for Neural Machine Translation
abstract
Machines translation (MT) is an essential task in natural language processing or even in artificial intelligence. Statistical machine translation has been the dominant approach to MT for decades, but recently neural machine translation achieves increasing interest because of its appealing model architecture and impressive translation performance. In neural machine translation, an attention model is used to identify the aligned source words for the next target word, i.e., target foresight word, to select translation context. However, it does not make use of any information about this target foresight word at all. Previous work proposed an approach to improve the attention model by explicitly accessing this target foresight word and demonstrating substantial alignment tasks. However, this approach cannot be applied in machine translation tasks where the target foresight word is unavailable. This paper proposes several novel enhanced attention models by introducing hidden information (such as part-of-speech) of the target foresight word for the translation task. We incorporate the novel enhanced attention employing hidden information about the target foresight word into both recurrent and self-attention-based neural translation models and theoretically justify that such hidden information can make translation prediction easier. Empirical experiments on four datasets further verify that the proposed attention models deliver significant improvements in translation quality.
Lemao Liu, Zhaopeng Tu, Shuming Shi 0001, Max Q.-H. Meng
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Detecting Source Contextual Barriers for Understanding Neural Machine Translation
abstract
In machine translation evaluation, the traditional wisdom measures model's generalization ability in an average sense, for example by using corpus BLEU. However, the statistics of corpus BLEU cannot provide comprehensive understanding and fine-grained analysis on model's generalization ability. As a remedy, this paper attempts to understand NMT at fine-grained level, by detecting contextual barriers within an unseen input sentence that \textit{cause} the degradation in model's translation quality. It proposes a principled definition of source contextual barriers as well as its modified version which is tractable in computation and operates at word-level. Based on the modified one, three simple methods are proposed for barrier detection by search-aware risk estimation through counterfactual generation. Extensive analyses are conducted on those detected contextual barrier words on both Zh$\Leftrightarrow$En NIST benchmarks. Potential usages motivated from barrier words are also discussed.
Lemao Liu, Conghui Zhu, Rui Wang 0015, Tiejun Zhao, Shuming Shi 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2020 CASE: Context-Aware Semantic Expansion
abstract
In this paper, we define and study a new task called Context-Aware Semantic Expansion (CASE). Given a seed term in a sentential context, we aim to suggest other terms that well fit the context as the seed. CASE has many interesting applications such as query suggestion, computer-assisted writing, and word sense disambiguation, to name a few. Previous explorations, if any, only involve some similar tasks, and all require human annotations for evaluation. In this study, we demonstrate that annotations for this task can be harvested at scale from existing corpora, in a fully automatic manner. On a dataset of 1.8 million sentences thus derived, we propose a network architecture that encodes the context and seed term separately before suggesting alternative terms. The context encoder in this architecture can be easily extended by incorporating seed-aware attention. Our experiments demonstrate that competitive results are achieved with appropriate choices of context encoder and attention scoring function.
Jialong Han, Aixin Sun, Haisong Zhang, Chenliang Li 0005, Shuming Shi 0001
AAAI5
2020 Neuron Interaction Based Representation Composition for Neural Machine Translation
abstract
Recent NLP studies reveal that substantial linguistic information can be attributed to single neurons, i.e., individual dimensions of the representation vectors. We hypothesize that modeling strong interactions among neurons helps to better capture complex information by composing the linguistic properties embedded in individual neurons. Starting from this intuition, we propose a novel approach to compose representations learned by different components in neural machine translation (e.g., multi-layer networks or multi-head attention), based on modeling strong interactions among neurons in the representation vectors. Specifically, we leverage bilinear pooling to model pairwise multiplicative interactions among individual neurons, and a low-rank approximation to make the model computationally feasible. We further propose extended bilinear pooling to incorporate first-order representations. Experiments on WMT14 English⇒German and English⇒French translation tasks show that our model consistently improves performances over the SOTA Transformer baseline. Further analyses demonstrate that our approach indeed captures more syntactic and semantic information as expected.
Jian Li 0054, Xing Wang 0007, Baosong Yang, Shuming Shi 0001, Michael R. Lyu, Zhaopeng Tu
AAAI4
2020 Go From the General to the Particular: Multi-Domain Translation with Domain Transformation Networks
abstract
The key challenge of multi-domain translation lies in simultaneously encoding both the general knowledge shared across domains and the particular knowledge distinctive to each domain in a unified model. Previous work shows that the standard neural machine translation (NMT) model, trained on mixed-domain data, generally captures the general knowledge, but misses the domain-specific knowledge. In response to this problem, we augment NMT model with additional domain transformation networks to transform the general representations to domain-specific representations, which are subsequently fed to the NMT decoder. To guarantee the knowledge transformation, we also propose two complementary supervision signals by leveraging the power of knowledge distillation and adversarial learning. Experimental results on several language pairs, covering both balanced and unbalanced multi-domain translation, demonstrate the effectiveness and universality of the proposed approach. Encouragingly, the proposed unified model achieves comparable results with the fine-tuning approach that requires multiple models to preserve the particular knowledge. Further analyses reveal that the domain transformation networks successfully capture the domain-specific knowledge as expected.1
Yong Wang 0032, Longyue Wang, Shuming Shi 0001, Victor O. K. Li, Zhaopeng Tu
AAAI3
2020 Balancing Quality and Human Involvement: An Effective Approach to Interactive Neural Machine Translation
abstract
Conventional interactive machine translation typically requires a human translator to validate every generated target word, even though most of them are correct in the advanced neural machine translation (NMT) scenario. Previous studies have exploited confidence approaches to address the intensive human involvement issue, which request human guidance only for a few number of words with low confidences. However, such approaches do not take the history of human involvement into account, and optimize the models only for the translation quality while ignoring the cost of human involvement. In response to these pitfalls, we propose a novel interactive NMT model, which explicitly accounts the history of human involvements and particularly is optimized towards two objectives corresponding to the translation quality and the cost of human involvement, respectively. Specifically, the model jointly predicts a target word and a decision on whether to request human guidance, which is based on both the partial translation and the history of human involvements. Since there is no explicit signals on the decisions of requesting human guidance in the bilingual corpus, we optimize the model with the reinforcement learning technique which enables our model to accurately predict when to request human guidance. Simulated and real experiments show that the proposed model can achieve higher translation quality with similar or less human involvement over the confidence-based baseline.
Tianxiang Zhao 0001, Lemao Liu, Guoping Huang, Yingling Liu, Guiquan Liu, Shuming Shi 0001
AAAI7
2020 Evaluating Explanation Methods for Neural Machine Translation
abstract
Recently many efforts have been devoted to interpreting the black-box NMT models, but little progress has been made on metrics to evaluate explanation methods.Word Alignment Error Rate can be used as such a metric that matches human understanding, however, it can not measure explanation methods on those target words that are not aligned to any source word.This paper thereby makes an initial attempt to evaluate explanation methods from an alternative viewpoint.To this end, it proposes a principled metric based on fidelity in regard to the predictive behavior of the NMT model.As the exact computation for this metric is intractable, we employ an efficient approach as its approximation.On six standard translation tasks, we quantitatively evaluate several explanation methods in terms of the proposed metric and we reveal some valuable findings for these explanation methods in our experiments.
Jierui Li, Lemao Liu, Guoping Huang, Shuming Shi 0001
ACL6
2020 Rigid Formats Controlled Text Generation
abstract
Neural text generation has made tremendous progress in various tasks.One common characteristic of most of the tasks is that the texts are not restricted to some rigid formats when generating.However, we may confront some special text paradigms such as Lyrics (assume the music score is given), Sonnet, SongCi (classical Chinese poetry of the Song dynasty), etc.The typical characteristics of these texts are in three folds: (1) They must comply fully with the rigid predefined formats.(2) They must obey some rhyming schemes.(3) Although they are restricted to some formats, the sentence integrity must be guaranteed.To the best of our knowledge, text generation based on the predefined rigid formats has not been well investigated.Therefore, we propose a simple and elegant framework named SongNet to tackle this problem.The backbone of the framework is a Transformer-based auto-regressive language model.Sets of symbols are tailor-designed to improve the modeling performance especially on format, rhyme, and sentence integrity.We improve the attention mechanism to impel the model to capture some future information on the format.A pre-training and fine-tuning framework is designed to further improve the generation quality.Extensive experiments conducted on two collected corpora demonstrate that our proposed framework generates significantly better results in terms of both automatic metrics and the human evaluation. 1
Piji Li, Haisong Zhang, Xiaojiang Liu, Shuming Shi 0001
ACL4
2020 On the Inference Calibration of Neural Machine Translation
abstract
Confidence calibration, which aims to make model predictions equal to the true correctness measures, is important for neural machine translation (NMT) because it is able to offer useful indicators of translation errors in the generated output.While prior studies have shown that NMT models trained with label smoothing are well-calibrated on the groundtruth training data, we find that miscalibration still remains a severe challenge for NMT during inference due to the discrepancy between training and inference.By carefully designing experiments on three language pairs, our work provides in-depth analyses of the correlation between calibration and translation performance as well as linguistic properties of miscalibration and reports a number of interesting findings that might help humans better analyze, understand and improve NMT models.Based on these observations, we further propose a new graduated label smoothing method that can improve both inference calibration and translation performance.1
Shuo Wang 0013, Zhaopeng Tu, Shuming Shi 0001, Yang Liu 0005
ACL3
2020 The World is Not Binary: Learning to Rank with Grayscale Data for Dialogue Response Selection
abstract
Response selection plays a vital role in building retrieval-based conversation systems.Despite that response selection is naturally a learning-to-rank problem, most prior works take a point-wise view and train binary classifiers for this task: each response candidate is labeled either relevant (one) or irrelevant (zero).On the one hand, this formalization can be sub-optimal due to its ignorance of the diversity of response quality.On the other hand, annotating grayscale data for learning-to-rank can be prohibitively expensive and challenging.In this work, we show that grayscale data can be automatically constructed without human effort.Our method employs off-the-shelf response retrieval models and response generation models as automatic grayscale data generators.With the constructed grayscale data, we propose multi-level ranking objectives for training, which can (1) teach a matching model to capture more fine-grained context-response relevance difference and (2) reduce the traintest discrepancy in terms of distractor strength.Our method is simple, effective, and universal.Experiments on three benchmark datasets and four state-of-the-art matching models show that the proposed approach brings significant and consistent performance improvements.
Zibo Lin, Deng Cai 0002, Yan Wang 0060, Xiaojiang Liu, Hai-Tao Zheng 0002, Shuming Shi 0001
EMNLP (1)6
2020 When Hearst Is not Enough: Improving Hypernymy Detection from Corpus with Distributional Models
abstract
We address hypernymy detection, i.e., whether an is-a relationship exists between words (x, y), with the help of large textual corpora.Most conventional approaches to this task have been categorized to be either pattern-based or distributional.Recent studies suggest that pattern-based ones are superior, if large-scale Hearst pairs are extracted and fed, with the sparsity of unseen (x, y) pairs relieved.However, they become invalid in some specific sparsity cases, where x or y is not involved in any pattern.For the first time, this paper quantifies the non-negligible existence of those specific cases.We also demonstrate that distributional methods are ideal to make up for patternbased ones in such cases.We devise a complementary framework, under which a patternbased and a distributional model collaborate seamlessly in cases which they each prefer.On several benchmark datasets, our framework achieves competitive improvements and the case study shows its better interpretability.
Changlong Yu, Jialong Han, Peifeng Wang, Yangqiu Song, Hongming Zhang 0009, Wilfred Ng, Shuming Shi 0001
EMNLP (1)7
2020 Exploiting deep representations for natural language processing
Zi-Yi Dou, Xing Wang 0007, Shuming Shi 0001, Zhaopeng Tu
Neurocomputing3
2019 Dynamic Layer Aggregation for Neural Machine Translation with Routing-by-Agreement
abstract
With the promising progress of deep neural networks, layer aggregation has been used to fuse information across layers in various fields, such as computer vision and machine translation. However, most of the previous methods combine layers in a static fashion in that their aggregation strategy is independent of specific hidden states. Inspired by recent progress on capsule networks, in this paper we propose to use routing-by-agreement strategies to aggregate layers dynamically. Specifically, the algorithm learns the probability of a part (individual layer representations) assigned to a whole (aggregated representations) in an iterative way and combines parts accordingly. We implement our algorithm on top of the state-of-the-art neural machine translation model TRANSFORMER and conduct experiments on the widely-used WMT14 sh⇒German and WMT17 Chinese⇒English translation datasets. Experimental results across language pairs show that the proposed approach consistently outperforms the strong baseline model and a representative static aggregation model.
Zi-Yi Dou, Zhaopeng Tu, Xing Wang 0007, Longyue Wang, Shuming Shi 0001, Tong Zhang 0001
AAAI5
2019 Generating Multiple Diverse Responses for Short-Text Conversation
abstract
Neural generative models have become popular and achieved promising performance on short-text conversation tasks. They are generally trained to build a 1-to-1 mapping from the input post to its output response. However, a given post is often associated with multiple replies simultaneously in real applications. Previous research on this task mainly focuses on improving the relevance and informativeness of the top one generated response for each post. Very few works study generating multiple accurate and diverse responses for the same post. In this paper, we propose a novel response generation model, which considers a set of responses jointly and generates multiple diverse responses simultaneously. A reinforcement learning algorithm is designed to solve our model. Experiments on two short-text conversation tasks validate that the multiple responses generated by our model obtain higher quality and larger diversity compared with various state-ofthe-art generative models.
Wei Bi, Xiaojiang Liu, Junhui Li 0001, Shuming Shi 0001
AAAI5
2019 Neural Machine Translation with Adequacy-Oriented Learning
abstract
Although Neural Machine Translation (NMT) models have advanced state-of-the-art performance in machine translation, they face problems like the inadequate translation. We attribute this to that the standard Maximum Likelihood Estimation (MLE) cannot judge the real translation quality due to its several limitations. In this work, we propose an adequacyoriented learning mechanism for NMT by casting translation as a stochastic policy in Reinforcement Learning (RL), where the reward is estimated by explicitly measuring translation adequacy. Benefiting from the sequence-level training of RL strategy and a more accurate reward designed specifically for translation, our model outperforms multiple strong baselines, including (1) standard and coverage-augmented attention models with MLE-based training, and (2) advanced reinforcement and adversarial training strategies with rewards based on both word-level BLEU and character-level CHRF3. Quantitative and qualitative analyses on different language pairs and NMT architectures demonstrate the effectiveness and universality of the proposed approach.
Xiang Kong, Zhaopeng Tu, Shuming Shi 0001, Eduard H. Hovy, Tong Zhang 0001
AAAI3
2019 Graph Based Translation Memory for Neural Machine Translation
abstract
A translation memory (TM) is proved to be helpful to improve neural machine translation (NMT). Existing approaches either pursue the decoding efficiency by merely accessing local information in a TM or encode the global information in a TM yet sacrificing efficiency due to redundancy. We propose an efficient approach to making use of the global information in a TM. The key idea is to pack a redundant TM into a compact graph and perform additional attention mechanisms over the packed graph for integrating the TM representation into the decoding network. We implement the model by extending the state-of-the-art NMT, Transformer. Extensive experiments on three language pairs show that the proposed approach is efficient in terms of running time and space occupation, and particularly it outperforms multiple strong baselines in terms of BLEU scores.
Mengzhou Xia, Guoping Huang, Lemao Liu, Shuming Shi 0001
AAAI4
2019 Fine-Grained Sentence Functions for Short-Text Conversation
abstract
Sentence function is an important linguistic feature referring to a user's purpose in uttering a specific sentence.The use of sentence function has shown promising results to improve the performance of conversation models.However, there is no large conversation dataset annotated with sentence functions.In this work, we collect a new Short-Text Conversation dataset with manually annotated SEntence FUNctions (STC-Sefun).Classification models are trained on this dataset to (i) recognize the sentence function of new data in a large corpus of short-text conversations; (ii) estimate a proper sentence function of the response given a test query.We later train conversation models conditioned on the sentence functions, including information retrieval-based and neural generative models.Experimental results demonstrate that the use of sentence functions can help improve the quality of the returned responses.
Wei Bi, Xiaojiang Liu, Shuming Shi 0001
ACL (1)4
2019 On the Word Alignment from Neural Machine Translation
abstract
Prior researches suggest that neural machine translation (NMT) captures word alignment through its attention mechanism, however, this paper finds attention may almost fail to capture word alignment for some NMT models.This paper thereby proposes two methods to induce word alignment which are general and agnostic to specific NMT models.Experiments show that both methods induce much better word alignment than attention.This paper further visualizes the translation through the word alignment induced by NMT.In particular, it analyzes the effect of alignment errors on translation errors at word level and its quantitative analysis over many testing examples consistently demonstrate that alignment errors are likely to lead to translation errors measured by different metrics.
Lemao Liu, Max Meng, Shuming Shi 0001
ACL (1)5
2019 Topic-Aware Neural Keyphrase Generation for Social Media Language
abstract
A huge volume of user-generated content is daily produced on social media.To facilitate automatic language understanding, we study keyphrase prediction, distilling salient information from massive posts.While most existing methods extract words from source posts to form keyphrases, we propose a sequence-to-sequence (seq2seq) based neural keyphrase generation framework, enabling absent keyphrases to be created.Moreover, our model, being topic-aware, allows joint modeling of corpus-level latent topic representations, which helps alleviate the data sparsity that widely exhibited in social media language.Experiments on three datasets collected from English and Chinese social media platforms show that our model significantly outperforms both extraction and generation models that do not exploit latent topics. 1 Further discussions show that our model learns meaningful topics, which interprets its superiority in social media keyphrase generation.
Yue Wang 0034, Jing Li 0049, Hou Pong Chan, Irwin King, Michael R. Lyu, Shuming Shi 0001
ACL (1)6
2019 Exploiting Sentential Context for Neural Machine Translation
abstract
In this work, we present novel approaches to exploit sentential context for neural machine translation (NMT).Specifically, we first show that a shallow sentential context extracted from the top encoder layer only, can improve translation performance via contextualizing the encoding representations of individual words.Next, we introduce a deep sentential context, which aggregates the sentential context representations from all the internal layers of the encoder to form a more comprehensive context representation.Experimental results on the WMT14 English⇒German and English⇒French benchmarks show that our model consistently improves performance over the strong TRANSFORMER model (Vaswani et al., 2017), demonstrating the necessity and effectiveness of exploiting sentential context for NMT.
Xing Wang 0007, Zhaopeng Tu, Longyue Wang, Shuming Shi 0001
ACL (1)4
2019 Retrieval-guided Dialogue Response Generation via a Matching-to-Generation Framework
abstract
Deng Cai, Yan Wang, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Deng Cai 0002, Yan Wang 0060, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Shuming Shi 0001
EMNLP/IJCNLP (1)6
2019 A Discrete CVAE for Response Generation on Short-Text Conversation
abstract
Jun Gao, Wei Bi, Xiaojiang Liu, Junhui Li, Guodong Zhou, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Wei Bi, Xiaojiang Liu, Junhui Li 0001, Guodong Zhou 0001, Shuming Shi 0001
EMNLP/IJCNLP (1)6
2019 Multi-Granularity Self-Attention for Neural Machine Translation
abstract
Jie Hao, Xing Wang, Shuming Shi, Jinfeng Zhang, Zhaopeng Tu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xing Wang 0007, Shuming Shi 0001, Zhaopeng Tu
EMNLP/IJCNLP (1)3
2019 Towards Better Modeling Hierarchical Structure for Self-Attention with Ordered Neurons
abstract
Jie Hao, Xing Wang, Shuming Shi, Jinfeng Zhang, Zhaopeng Tu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xing Wang 0007, Shuming Shi 0001, Zhaopeng Tu
EMNLP/IJCNLP (1)3
2019 Towards Understanding Neural Machine Translation with Word Importance
abstract
Shilin He, Zhaopeng Tu, Xing Wang, Longyue Wang, Michael Lyu, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shilin He, Zhaopeng Tu, Xing Wang 0007, Longyue Wang, Michael R. Lyu, Shuming Shi 0001
EMNLP/IJCNLP (1)6
2019 Semi-supervised Text Style Transfer: Cross Projection in Latent Space
abstract
Mingyue Shang, Piji Li, Zhenxin Fu, Lidong Bing, Dongyan Zhao, Shuming Shi, Rui Yan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mingyue Shang, Piji Li, Zhenxin Fu, Lidong Bing, Dongyan Zhao 0001, Shuming Shi 0001, Rui Yan 0001
EMNLP/IJCNLP (1)6
2019 One Model to Learn Both: Zero Pronoun Prediction and Translation
abstract
Longyue Wang, Zhaopeng Tu, Xing Wang, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Longyue Wang, Zhaopeng Tu, Xing Wang 0007, Shuming Shi 0001
EMNLP/IJCNLP (1)4
2019 Self-Attention with Structural Position Representations
abstract
Xing Wang, Zhaopeng Tu, Longyue Wang, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xing Wang 0007, Zhaopeng Tu, Longyue Wang, Shuming Shi 0001
EMNLP/IJCNLP (1)4
2018 Improving Sequence-to-Sequence Constituency Parsing
abstract
Sequence-to-sequence constituency parsing casts the tree structured prediction problem as a general sequential problem by top-down tree linearization,and thus it is very easy to train in parallel with distributed facilities. Despite its success, it relies on a probabilistic attention mechanism for a general purpose, which can not guarantee the selected context to be informative in the specific parsing scenario. Previous work introduced a deterministic attention to select the informative context for sequence-to-sequence parsing, but it is based on the bottom-up linearization even if it was observed that top-down linearization is better than bottom-up linearization for standard sequence-to-sequence constituency parsing. In this paper, we thereby extend the deterministic attention to directly conduct on the top-down tree linearization. Intensive experiments show that our parser delivers substantial improvements over the bottom-up linearization in accuracy, and it achieves 92.3 Fscore on the Penn English Treebank section 23 and 85.4 Fscore on the Penn Chinese Treebank test dataset, without reranking or semi-supervised training.
Lemao Liu, Muhua Zhu, Shuming Shi 0001
AAAI3
2018 Translating Pro-Drop Languages With Reconstruction Models
abstract
Pronouns are frequently omitted in pro-drop languages, such as Chinese, generally leading to significant challenges with respect to the production of complete translations. To date, very little attention has been paid to the dropped pronoun (DP) problem within neural machine translation (NMT). In this work, we propose a novel reconstruction-based approach to alleviating DP translation problems for NMT models. Firstly, DPs within all source sentences are automatically annotated with parallel information extracted from the bilingual training corpus. Next, the annotated source sentence is reconstructed from hidden representations in the NMT model. With auxiliary training objectives, in the terms of reconstruction scores, the parameters associated with the NMT model are guided to produce enhanced hidden representations that are encouraged as much as possible to embed annotated DP information. Experimental results on both Chinese-English and Japanese-English dialogue translation tasks show that the proposed approach significantly and consistently improves translation performance over a strong NMT baseline, which is directly built on the training data annotated with DPs.
Longyue Wang, Zhaopeng Tu, Shuming Shi 0001, Tong Zhang 0001, Yvette Graham, Qun Liu 0001
AAAI3
2018 hyperdoc2vec: Distributed Representations of Hypertext Documents
abstract
Hypertext documents, such as web pages and academic papers, are of great importance in delivering information in our daily life.Although being effective on plain documents, conventional text embedding methods suffer from information loss if directly adapted to hyper-documents.In this paper, we propose a general embedding approach for hyper-documents, namely, hyperdoc2vec, along with four criteria characterizing necessary information that hyper-document embedding models should preserve.Systematic comparisons are conducted between hyperdoc2vec and several competitors on two tasks, i.e., paper classification and citation recommendation, in the academic paper domain.Analyses and experiments both validate the superiority of hyperdoc2vec to other models w.r.t. the four criteria.
Jialong Han, Yan Song 0003, Wayne Xin Zhao, Shuming Shi 0001, Haisong Zhang
ACL (1)4
2018 Exploiting Deep Representations for Neural Machine Translation
abstract
Advanced neural machine translation (NMT) models generally implement encoder and decoder as multiple layers, which allows systems to model complex functions and capture complicated linguistic structures.However, only the top layers of encoder and decoder are leveraged in the subsequent process, which misses the opportunity to exploit the useful information embedded in other layers.In this work, we propose to simultaneously expose all of these signals with layer aggregation and multi-layer attention mechanisms.In addition, we introduce an auxiliary regularization term to encourage different layers to capture diverse information.Experimental results on widely-used WMT14 English⇒German and WMT17 Chinese⇒English translation data demonstrate the effectiveness and universality of the proposed approach.
Zi-Yi Dou, Zhaopeng Tu, Xing Wang 0007, Shuming Shi 0001, Tong Zhang 0001
EMNLP4
2018 Generating Classical Chinese Poems via Conditional Variational Autoencoder and Adversarial Training
abstract
It is a challenging task to automatically compose poems with not only fluent expressions but also aesthetic wording.Although much attention has been paid to this task and promising progress is made, there exist notable gaps between automatically generated ones with those created by humans, especially on the aspects of term novelty and thematic consistency.Towards filling the gap, in this paper, we propose a conditional variational autoencoder with adversarial training for classical Chinese poem generation, where the autoencoder part generates poems with novel terms and a discriminator is applied to adversarially learn their thematic consistency with their titles.Experimental results on a large poetry corpus confirm the validity and effectiveness of our model, where its automatic and human evaluation scores outperform existing models.
Juntao Li 0005, Yan Song 0003, Haisong Zhang, Dongmin Chen, Shuming Shi 0001, Dongyan Zhao 0001, Rui Yan 0001
EMNLP5
2018 QuaSE: Sequence Editing under Quantifiable Guidance
abstract
We propose the task of Quantifiable Sequence Editing (QuaSE): editing an input sequence to generate an output sequence that satisfies a given numerical outcome value measuring a certain property of the sequence, with the requirement of keeping the main content of the input sequence.For example, an input sequence could be a word sequence, such as review sentence and advertisement text.For a review sentence, the outcome could be the review rating; for an advertisement, the outcome could be the click-through rate.The major challenge in performing QuaSE is how to perceive the outcome-related wordings, and only edit them to change the outcome.In this paper, the proposed framework contains two latent factors, namely, outcome factor and content factor, disentangled from the input sentence to allow convenient editing to change the outcome and keep the content.Our framework explores the pseudo-parallel sentences by modeling their content similarity and outcome differences to enable a better disentanglement of the latent factors, which allows generating an output to better satisfy the desired outcome and keep the content.The dual reconstruction structure further enhances the capability of generating expected output by exploiting the couplings of latent factors of pseudo-parallel sentences.For evaluation, we prepared a dataset of Yelp review sentences with the ratings as outcome.Extensive experimental results are reported and discussed to elaborate the peculiarities of our framework.1
Lidong Bing, Piji Li, Shuming Shi 0001, Wai Lam, Tong Zhang 0001
EMNLP4
2018 Towards Less Generic Responses in Neural Conversation Models: A Statistical Re-weighting Method
abstract
Sequence-to-sequence neural generation models have achieved promising performance on short text conversation tasks.However, they tend to generate generic/dull responses, leading to unsatisfying dialogue experience.We observe that in conversation tasks, each query could have multiple responses, which forms a 1-to-n or m-to-n relationship in the view of the total corpus.The objective function used in standard sequence-to-sequence models will be dominated by loss terms with generic patterns.Inspired by this observation, we introduce a statistical re-weighting method that assigns different weights for the multiple responses of the same query, and trains the standard neural generation model with the weights.Experimental results on a large Chinese dialogue corpus show that our method improves the acceptance rate of generated responses compared with several baseline models and significantly reduces the number of generated generic responses.
Wei Bi, Xiaojiang Liu, Jian Yao 0002, Shuming Shi 0001
EMNLP6
2018 Complementary Learning of Word Embeddings
abstract
Continuous bag-of-words (CB) and skip-gram (SG) models are popular approaches to training word embeddings. Conventionally they are two standing-alone techniques used individually. However, with the same goal of building embeddings by leveraging surrounding words, they are in fact a pair of complementary tasks where the output of one model can be used as input of the other, and vice versa. In this paper, we propose complementary learning of word embeddings based on the CB and SG model. Specifically, one round of learning first integrates the predicted output of a SG model with existing context, then forms an enlarged context as input to the CB model. Final models are obtained through several rounds of parameter updating. Experimental results indicate that our approach can effectively improve the quality of initial embeddings, in terms of intrinsic and extrinsic evaluations.
Yan Song 0003, Shuming Shi 0001
IJCAI2
2018 Joint Learning Embeddings for Chinese Words and their Components via Ladder Structured Networks
abstract
The components, such as characters and radicals, of a Chinese word are important sources to help in capturing semantic information of the word. In this paper, we propose a novel framework, namely, ladder structured networks (LSN), which contains three layers representing word, character and radical and learns their embeddings synchronously. LSN captures not only the relations among words, but also the relations among their component characters and radicals, as well as the relations across layers. Each layer in LSN is pluggable so that any particular type of unit (word, character, radical) can be removed and the LSN is thus adjusted for particular types of inputs. In evaluating our framework, we use word similarity as the intrinsic evaluation and part-of-speech tagging and document classification as extrinsic evaluations. Experimental results confirm the validity of our approach and show superiority of our approach over previous work.
Yan Song 0003, Shuming Shi 0001, Jing Li 0049
IJCAI2
2018 Target Foresight Based Attention for Neural Machine Translation
abstract
Xintong Li, Lemao Liu, Zhaopeng Tu, Shuming Shi, Max Meng. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Lemao Liu, Zhaopeng Tu, Shuming Shi 0001, Max Meng
NAACL-HLT4
2018 Learning to Remember Translation History with a Continuous Cache
abstract
Existing neural machine translation (NMT) models generally translate sentences in isolation, missing the opportunity to take advantage of document-level information. In this work, we propose to augment NMT models with a very light-weight cache-like memory network, which stores recent hidden representations as translation history. The probability distribution over generated words is updated online depending on the translation history retrieved from the memory, endowing NMT models with the capability to dynamically adapt over time. Experiments on multiple domains with different topics and styles show the effectiveness of the proposed approach with negligible impact on the computational cost.
Zhaopeng Tu, Yang Liu 0005, Shuming Shi 0001, Tong Zhang 0001
Trans. Assoc. Comput. Linguistics3
2017 Learning Fine-Grained Expressions to Solve Math Word Problems
abstract
This paper presents a novel templatebased method to solve math word problems.This method learns the mappings between math concept phrases in math word problems and their math expressions from training data.For each equation template, we automatically construct a rich template sketch by aggregating information from various problems with the same template.Our approach is implemented in a two-stage system.It first retrieves a few relevant equation system templates and aligns numbers in math word problems to those templates for candidate equation generation.It then does a fine-grained inference to obtain the final answer.Experiment results show that our method achieves an accuracy of 28.4% on the linear Dolphin18K benchmark, which is 10% (54% relative) higher than previous stateof-the-art systems while achieving an accuracy increase of 12% (59% relative) on the TS6 benchmark subset.
Danqing Huang, Shuming Shi 0001, Chin-Yew Lin, Jian Yin 0001
EMNLP2
2017 Deep Neural Solver for Math Word Problems
abstract
This paper presents a deep neural solver to automatically solve math word problems.In contrast to previous statistical learning approaches, we directly translate math word problems to equation templates using a recurrent neural network (RNN) model, without sophisticated feature engineering.We further design a hybrid model that combines the RNN model and a similarity-based retrieval model to achieve additional performance improvement.Experiments conducted on a large dataset show that the RNN model and the hybrid model significantly outperform stateof-the-art statistical learning methods for math word problem solving. 1 We plan to make the dataset publicly available when the paper is published
Yan Wang 0060, Xiaojiang Liu, Shuming Shi 0001
EMNLP3
2016 How well do Computers Solve Math Word Problems? Large-Scale Dataset Construction and Evaluation
abstract
Recently a few systems for automatically solving math word problems have reported promising results. However, the datasets used for evaluation have limitations in both scale and diversity. In this paper, we build a large-scale dataset which is more than 9 times the size of previous ones, and contains many more problem types. Problems in the dataset are semi-automatically obtained from community question-answering (CQA) web pages. A ranking SVM model is trained to automatically extract problem answers from the answer text provided by CQA users, which significantly reduces human annotation cost. Experiments conducted on the new dataset lead to interesting and surprising results.
Danqing Huang, Shuming Shi 0001, Chin-Yew Lin, Jian Yin 0001, Wei-Ying Ma
ACL (1)2
2015 Automatically Solving Number Word Problems by Semantic Parsing and Reasoning
abstract
This paper presents a semantic parsing and reasoning approach to automatically solving math word problems.A new meaning representation language is designed to bridge natural language text and math expressions.A CFG parser is implemented based on 9,600 semi-automatically created grammar rules.We conduct experiments on a test set of over 1,500 number word problems (i.e., verbally expressed number problems) and yield 95.4% precision and 60.2% recall.
Shuming Shi 0001, Yuehui Wang, Chin-Yew Lin, Xiaojiang Liu, Yong Rui
EMNLP1
2014 Improving Context and Category Matching for Entity Search
abstract
Entity search is to retrieve a ranked list of named entities of target types to a given query. In this paper, we propose an approach of entity search by formalizing both context matching and category matching. In addition, we propose a result re-ranking strategy that can be easily adapted to achieve a hybrid of two context matching strategies. Experiments on the INEX 2009 entity ranking task show that the proposed approach achieves a significant improvement of the entity search performance (xinfAP from 0.27 to 0.39) over the existing solutions.
Yueguo Chen, Lexi Gao, Shuming Shi 0001, Xiaoyong Du 0001, Ji-Rong Wen
AAAI3
2014 Unsupervised Template Mining for Semantic Category Understanding
abstract
We propose an unsupervised approach to constructing templates from a large collection of semantic category names, and use the templates as the semantic representation of categories.The main challenge is that many terms have multiple meanings, resulting in a lot of wrong templates.Statistical data and semantic knowledge are extracted from a web corpus to improve template generation.A nonlinear scoring function is proposed and demonstrated to be effective.Experiments show that our approach achieves significantly better results than baseline methods.As an immediate application, we apply the extracted templates to the cleaning of a category collection and see promising results (precision improved from 81% to 89%).
Lei Shi 0015, Shuming Shi 0001, Chin-Yew Lin, Yidong Shen, Yong Rui
EMNLP2
2012 Ensemble Semantics for Large-scale Unsupervised Relation Extraction
Bonan Min, Shuming Shi 0001, Ralph Grishman, Chin-Yew Lin
EMNLP-CoNLL2
2012 Towards Large-Scale Unsupervised Relation Extraction from the Web
abstract
The Web brings an open-ended set of semantic relations. Discovering the significant types is very challenging. Unsupervised algorithms have been developed to extract relations from a corpus without knowing the relation types in advance, but most rely on tagging arguments of predefined types. One recently reported system is able to jointly extract relations and their argument semantic classes, taking a set of relation instances extracted by an open IE (Information Extraction) algorithm as input. However, it cannot handle polysemy of relation phrases and fails to group many similar (“synonymous”) relation instances because of the sparseness of features. In this paper, the authors present a novel unsupervised algorithm that provides a more general treatment of the polysemy and synonymy problems. The algorithm incorporates various knowledge sources which they will show to be very effective for unsupervised relation extraction. Moreover, it explicitly disambiguates polysemous relation phrases and groups synonymous ones. While maintaining approximately the same precision, the algorithm achieves significant improvement on recall compared to the previous method. It is also very efficient. Experiments on a real-world dataset show that it can handle 14.7 million relation instances and extract a very large set of relations from the Web.
Bonan Min, Shuming Shi 0001, Ralph Grishman, Chin-Yew Lin
Int. J. Semantic Web Inf. Syst.2
2011 Nonlinear Evidence Fusion and Propagation for Hyponymy Relation Mining
Fan Zhang 0092, Shuming Shi 0001, Jing Liu 0022, Shu-Qi Sun, Chin-Yew Lin
ACL2
2010 Efficient term proximity search with term-pair indexes
abstract
There has been a large amount of research on early termination techniques in web search and information retrieval. Such techniques return the top-k documents without scanning and evaluating the full inverted lists of the query terms. Thus, they can greatly improve query processing efficiency. However, only a limited amount of efficient top-k processing work considers the impact of term proximity, i.e., the distance between term occurrences in a document, which has recently been integrated into a number of retrieval models to improve effectiveness.
Shuming Shi 0001, Fan Zhang 0092, Torsten Suel, Ji-Rong Wen
CIKM2
2010 Corpus-based Semantic Class Mining: Distributional vs. Pattern-Based Approaches
Shuming Shi 0001, Huibin Zhang, Xiaojie Yuan, Ji-Rong Wen
COLING1
2010 Revisiting globally sorted indexes for efficient document retrieval
abstract
There has been a large amount of research on efficient document retrieval in both IR and web search areas. One important technique to improve retrieval efficiency is early termination, which speeds up query processing by avoiding scanning the entire inverted lists. Most early termination techniques first build new inverted indexes by sorting the inverted lists in the order of either the term-dependent information, e.g., term frequencies or term IR scores, or the term-independent information, e.g., static rank of the document; and then apply appropriate retrieval strategies on the resulting indexes. Although the methods based only on the static rank have been shown to be ineffective for the early termination, there are still many advantages of using the methods based on term-independent information. In this paper, we propose new techniques to organize inverted indexes based on the term-independent information beyond static rank and study the new retrieval strategies on the resulting indexes. We perform a detailed experimental evaluation on our new techniques and compare them with the existing approaches. Our results on the TREC GOV and GOV2 data sets show that our techniques can improve query efficiency significantly.
Fan Zhang 0092, Shuming Shi 0001, Ji-Rong Wen
WSDM2
2009 Employing Topic Models for Pattern-based Semantic Class Discovery
Huibin Zhang, Mingjie Zhu, Shuming Shi 0001, Ji-Rong Wen
ACL/IJCNLP3
2009 Nonlinear static-rank computation
abstract
Mainstream link-based static-rank algorithms (e.g. PageRank and its variants) express the importance of a page as the linear combination of its in-links and compute page importance scores by solving a linear system in an iterative way. Such linear algorithms, however, may give apparently unreasonable static-rank results for some link structures. In this paper, we examine the static-rank computation problem from the viewpoint of evidence combination and build a probabilistic model for it. Based on the model, we argue that a nonlinear formula should be adopted, due to the correlation or dependence between links. We focus on examining some simple formulas which only consider the correlation between links in the same domain. Experiments conducted on 100 million web pages (with multiple static-rank quality evaluation metrics) show that higher quality static-rank could be yielded by the new nonlinear algorithms. The convergence of the new algorithms is also proved in this paper by nonlinear functional analysis.
Shuming Shi 0001, Yunxiao Ma, Ji-Rong Wen
CIKM1
2009 Effective top-k computation with term-proximity support
Mingjie Zhu, Shuming Shi 0001, Mingjing Li, Ji-Rong Wen
Inf. Process. Manag.2
2008 Pattern-based semantic class discovery with multi-membership support
abstract
A semantic class is a collection of items (words or phrases) sharing common semantic properties. This paper proposes an approach to constructing one or multiple semantic classes for an input item. Two challenges are addressed: multi-membership, and noise-tolerance.
Shuming Shi 0001, Ji-Rong Wen
CIKM1
2008 Can phrase indexing help to process non-phrase queries?
abstract
Modern web search engines, while indexing billions of web pages, are expected to process queries and return results in a very short time. Many approaches have been proposed for efficiently computing top-k query results, but most of them ignore one key factor in the ranking functions of commercial search engines - term-proximity, which is the metric of the distance between query terms in a document. When term-proximity is included in ranking functions, most of the existing top-k algorithms will become inefficient. To address this problem, in this paper we propose to build a compact phrase index to speed up the search process when incorporating the term-proximity factor. The compact phrase index can help more accurately estimate the score upper bounds of unknown documents. The size of the phrase index is controlled by including a small portion of phrases which are possibly helpful for improving search performance. Phrase index has been used to process phrase queries in existing work. It is, however, to the best of our knowledge, the first time that phrase index is used to improve the performance of generic queries. Experimental results show that, compared with the state-of-the-art top-k computation approaches, our approach can reduce average query processing time to 1/5 for typical setttings.
Mingjie Zhu, Shuming Shi 0001, Nenghai Yu, Ji-Rong Wen
CIKM2
2008 Improving relevance judgment of web search results with image excerpts
abstract
Current web search engines return result pages containing mostly text summary even though the matched web pages may contain informative pictures. A text excerpt (i.e. snippet) is generated by selecting keywords around the matched query terms for each returned page to provide context for user's relevance judgment. However, in many scenarios, we found that the pictures in web pages, if selected properly, could be added into search result pages and provide richer contextual description because a picture is worth a thousand words. Such new summary is named as image excerpts. By well designed user study, we demonstrate image excerpts can help users make much quicker relevance judgment of search results for a wide range of query types. To implement this idea, we propose a practicable approach to automatically generate image excerpts in the result pages by considering the dominance of each picture in each web page and the relevance of the picture to the query. We also outline an efficient way to incorporate image excerpts in web search engines. Web search engines can adopt our approach by slightly modifying their index and inserting a few low cost operations in their workflow. Our experiments on a large web dataset indicate the performance of the proposed approach is very promising.
Zhiwei Li 0006, Shuming Shi 0001, Lei Zhang 0001
WWW2
2007 Effective top-k computation in retrieving structured documents with term-proximity support
abstract
Modern web search engines are expected to return top-k results efficiently given a query. Although many dynamic index pruning strategies have been proposed for efficient top-k computation, most of them are prone to ignore some especially important factors in ranking functions, e.g. term proximity (the distance relationship between query terms in a document). The inclusion of term proximity breaks the monotonicity of ranking functions and therefore leads to additional challenges for efficient query processing. This paper studies the performance of some existing top-k computation approaches using term-proximity-enabled ranking functions. Our investigation demonstrates that, when term proximity is incorporated into ranking functions, most existing index structures and top-k strategies become quite inefficient. According to our analysis and experimental results, we propose two index structures and their corresponding index pruning strategies: Structured and Hybrid, which performs much better on the new settings. Moreover, the efficiency of index building and maintenance would not be affected too much with the two approaches.
Mingjie Zhu, Shuming Shi 0001, Mingjing Li, Ji-Rong Wen
CIKM2
2007 Improve Ranking by Using Image Information
Shuming Shi 0001, Zhiwei Li 0006, Ji-Rong Wen, Wei-Ying Ma
ECIR2
2007 Web object retrieval
abstract
The primary function of current Web search engines is essentially relevance ranking at the document level. However, myriad structured information about real-world objects embedded in static Web pages and online Web databases. In this paper, we propose a paradigm shift to enable searching at the object level. In traditional information retrieval models, documents are taken as the retrieval units and the content of a document is considered reliable. However, this reliability assumption is no longer valid in the object retrieval context when multiple copies of information about the same object typically exist. These copies may be inconsistent because of diversity of Web site qualities and the limited performance of current information extraction techniques. In this paper, we propose several language models for Web object retrieval. We test these models on our academic search engine called Libra and compare their performances. 1.
Zaiqing Nie, Yunxiao Ma, Shuming Shi 0001, Ji-Rong Wen, Wei-Ying Ma
WWW3
2007 Web page title extraction and its application
Yewei Xue, Yunhua Hu, Guomao Xin, Ruihua Song, Shuming Shi 0001, Yunbo Cao, Chin-Yew Lin, Hang Li 0001
Inf. Process. Manag.5
2006 Pseudo-anchor text extraction for searching vertical objects
abstract
This paper examines the problem of utilizing pseudo-anchor text to help ranking Web objects in vertical search. We adopt a machine learning based approach to extract pseudo-anchor text for a vertical object from its candidate anchor blocks. Experiments in academic search domain indicate that our approach is able to dramatically improve search performance.
Shuming Shi 0001, Mingjie Zhu, Zaiqing Nie, Ji-Rong Wen
CIKM1
2006 Exploring URL Hit Priors for Web Search
Ruihua Song, Guomao Xin, Shuming Shi 0001, Ji-Rong Wen, Wei-Ying Ma
ECIR3
2005 Title extraction from bodies of HTML documents and its application to web page retrieval
abstract
This paper is concerned with automatic extraction of titles from the bodies of HTML documents. Titles of HTML documents should be correctly defined in the title fields; however, in reality HTML titles are often bogus. It is desirable to conduct automatic extraction of titles from the bodies of HTML documents. This is an issue which does not seem to have been investigated previously. In this paper, we take a supervised machine learning approach to address the problem. We propose a specification on HTML titles. We utilize format information such as font size, position, and font weight as features in title extraction. Our method significantly outperforms the baseline method of using the lines in largest font size as title (20.9%-32.6% improvement in F1 score). As application, we consider web page retrieval. We use the TREC Web Track data for evaluation. We propose a new method for HTML documents retrieval using extracted titles. Experimental results indicate that the use of both extracted titles and title fields is almost always better than the use of title fields alone; the use of extracted titles is particularly helpful in the task of named page finding (23.1% -29.0% improvements).
Yunhua Hu, Guomao Xin, Ruihua Song, Shuming Shi 0001, Yunbo Cao, Hang Li 0001
SIGIR5
2005 Gravitation-based model for information retrieval
abstract
This paper proposes GBM (gravitation-based model), a physical model for information retrieval inspired by Newton's theory of gravitation. A mapping is built in this model from concepts of information retrieval (documents, queries, relevance, etc) to those of physics (mass, distance, radius, attractive force, etc). This model actually provides a new perspective on IR problems. A family of effective term weighting functions can be derived from it, including the well-known BM25 formula. This model has some advantages over most existing ones: First, because it is directly based on basic physical laws, the derived formulas and algorithms can have their explicit physical interpretation. Second, the ranking formulas derived from this model satisfy more intuitive heuristics than most of existing ones, thus have the potential to behave empirically better and to be used safely on various settings. Finally, a new approach for structured document retrieval derived from this model is more reasonable and behaves better than existing ones.
Shuming Shi 0001, Ji-Rong Wen, Ruihua Song, Wei-Ying Ma
SIGIR1