VLDB 2026 Research / reviewers in the wild / expert
Rui Wang 0015
dblp:w/RuiWang15
· DBLP profile ↗
81ranked-venue papers
11as first author
44since 2021 · last 2026
0000-0001-8007-2503ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 80 · 11 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation ModelsabstractRui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang, Haitao Mi, Dong Yu, Kam-Fai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rui Wang 0015, Ce Zhang 0009, Jun-Yu Ma, Hongru Wang 0003, Yi Chen 0007, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Kam-Fai Wong |
ACL (1) | 1 |
| 2026 | A Systematic Assessment of Language Models with Linguistic Minimal Pairs in ChineseabstractAbstract We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the Ba construction. We then train from scratch a suite of Chinese language models (LMs) with different tokenizers, parameter sizes, and token volumes, to study the learning curves of LMs on Chinese. To mitigate the biases introduced by unequal lengths of the sentences in a minimal pair, we propose a new metric named sub-linear length normalized log-probabilities (SLLN-LP). Using SLLN-LP as the metric, our results show that Anaphor, Quantifiers, and Ellipsis in Chinese are difficult for LMs even up to 32B parameters, and that SLLN-LP successfully mitigates biases in ZhoBLiMP, JBLiMP and BLiMP. We conclude that future evaluations should be more carefully designed to consider the intricate relations between linking functions, LMs, and targeted minimal pairs. Yikang Liu 0002, Yeting Shen, Hongao Zhu, Lilong Xu, Zhiheng Qian, Siyuan Song, Jialong Tang, Pei Zhang 0011, Baosong Yang, Rui Wang 0015, Hai Hu 0001 |
Trans. Assoc. Comput. Linguistics | 11 |
| 2025 | Enhancing Machine Translation with Self-Supervised Preference DataabstractModel alignment methods like Direct Preference Optimization and Contrastive Preference Optimization have enhanced machine translation performance by leveraging preference data to enable models to reject suboptimal outputs. During preference data construction, previous approaches primarily rely on humans, strong models like GPT4 or model self-sampling. In this study, we first explain the shortcomings of this practice. Then, we propose Self-Supervised Preference Optimization (SSPO), a novel framework which efficiently constructs translation preference data for iterative DPO training. Applying SSPO to 14B parameters large language models (LLMs) achieves comparable or better performance than GPT-4o on FLORES and multi-domain test datasets. We release an augmented MQM dataset in https://github.com/sunny-sjtu/MQM-aug. Haoxiang Sun, Pei Zhang 0011, Baosong Yang, Rui Wang 0015 |
ACL (1) | 5 |
| 2025 | GALLa: Graph Aligned Large Language Models for Improved Source Code UnderstandingabstractProgramming languages possess rich semantic information - such as data flow - that is represented by graphs and not available from the surface form of source code. Recent code language models have scaled to billions of parameters, but model source code solely as text tokens while ignoring any other structural information. Conversely, models that do encode structural information of code make modifications to the Transformer architecture, limiting their scale and compatibility with pretrained LLMs. In this work, we take the best of both worlds with GALLa - Graph Aligned Large Language Models. GALLa utilizes graph neural networks and cross-modal alignment technologies to inject the structural information of code into LLMs as an auxiliary task during finetuning. This framework is both model-agnostic and task-agnostic, as it can be applied to any code LLM for any code downstream task, and requires the structural graph data only at training time from a corpus unrelated to the finetuning data, while incurring no cost at inference time over the baseline LLM. Experiments on five code tasks with six different baseline LLMs ranging in size from 350M to 14B validate the effectiveness of GALLa, demonstrating consistent improvement over the baseline, even for powerful models such as LLaMA3 and Qwen2.5-Coder. Ziyin Zhang, Hang Yu 0002, Sage Lee, Peng Di, Rui Wang 0015 |
ACL (1) | 6 |
| 2025 | RetroDiff: Retrosynthesis as Multi-stage Distribution InterpolationabstractRetrosynthesis poses a key challenge in biopharmaceuticals, aiding chemists in finding appropriate reactant molecules for given product molecules. With reactants and products represented as 2D graphs, retrosynthesis constitutes a conditional graph-to-graph (G2G) generative task. Inspired by advancements in discrete diffusion models for graph generation, we aim to design a diffusion-based method to address this problem. However, integrating a diffusion-based G2G framework while retaining essential chemical reaction template information presents a notable challenge. Our key innovation involves a multi-stage diffusion process. We decompose the retrosynthesis procedure to first sample external groups from the dummy distribution given products, then generate external bonds to connect products and generated groups. Interestingly, this generation process mirrors the reverse of the widely adapted semi-template retrosynthesis workflow, i.e., from reaction center identification to synthon completion. Based on these designs, we introduce Retrosynthesis Diffusion (RetroDiff), a novel diffusion-based method for the retrosynthesis task. Experimental results demonstrate that RetroDiff surpasses all semi-template methods in accuracy, and outperforms template-based and template-free methods in large-scale scenarios and molecular validity, respectively. Yiming Wang 0011, Yuxuan Song 0002, Minkai Xu, Rui Wang 0015, Hao Zhou 0012, Wei-Ying Ma |
AISTATS | 5 |
| 2025 | Cognitive Insights into Document Comprehension: The Role of Reading Order and Visual Attention in Human and Large Language Models
Qingxuan Wang, Hao Wang 0097, Huiran Zhang, Chenhui Chu, Rui Wang 0015, Pinpin Zhu |
CogSci | 5 |
| 2025 | Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of TranslationeseabstractYikang Liu, Wanyang Zhang, Yiming Wang, Jialong Tang, Pei Zhang, Baosong Yang, Fei Huang, Rui Wang, Hai Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yikang Liu 0002, Wanyang Zhang, Yiming Wang 0011, Jialong Tang, Pei Zhang 0011, Baosong Yang, Fei Huang 0002, Rui Wang 0015, Hai Hu 0001 |
EMNLP | 8 |
| 2025 | Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form GenerationabstractConventional speculative decoding (SD) methods utilize a predefined length policy for proposing drafts, which implies the premise that the target model smoothly accepts the proposed draft tokens.However, reality deviates from this assumption: the oracle draft length varies significantly, and the fixed-length policy hardly satisfies such a requirement.Moreover, such discrepancy is further exacerbated in scenarios involving complex reasoning and long-form generation, particularly under testtime scaling for reasoning-specialized models.Through both theoretical and empirical estimation, we establish that the discrepancy between the draft and target models can be approximated by the draft model's prediction entropy: a high entropy indicates a low acceptance rate of draft tokens, and vice versa.Based on this insight, we propose SVIP: Self-Verification Length Policy for Long-Context Speculative Decoding, which is a training-free dynamic length policy for speculative decoding systems that adaptively determines the lengths of draft sequences by referring to the draft entropy.Experimental results on mainstream SD benchmarks as well as reasoning-heavy benchmarks demonstrate the superior performance of SVIP, achieving up to 17% speedup on MT-Bench at 8K context compared with fixed draft lengths, and 22% speedup for QwQ in long-form reasoning. Ziyin Zhang, Zhiwei He 0002, Rui Wang 0015, Zhaopeng Tu |
EMNLP | 6 |
| 2025 | Boosting Large Language Model for Speech Synthesis: An Empirical StudyabstractLarge language models (LLMs) have made significant advancements in natural language processing and are concurrently extending the language ability to other modalities, such as speech and vision. Nevertheless, most of the previous work focuses on prompting LLMs with perception abilities like auditory comprehension, and the effective approach for augmenting LLMs with speech synthesis capabilities remains ambiguous. In this paper, we conduct a comprehensive empirical exploration of boosting LLMs with the ability to generate speech, by combining pre-trained LLM LLaMA/OPT and text-to-speech synthesis model VALL-E. We compare three integration methods between LLMs and speech synthesis models, including directly fine-tuned LLMs, superposed layers of LLMs and VALL-E, and coupled LLMs and VALL-E using LLMs as a powerful text encoder. Experimental results show that, using LoRA method to fine-tune LLMs directly to boost the speech synthesis capability does not work well, and superposed LLMs and VALL-E can improve the quality of generated speech both in speaker similarity and word error rate (WER). Among these three methods, coupled methods leveraging LLMs as the text encoder can achieve the best performance, making it outperform original speech synthesis models with a consistently better speaker similarity and a significant (10.9%) WER reduction. Hongkun Hao, Shujie Liu 0001, Jinyu Li 0001, Shujie Hu, Rui Wang 0015, Furu Wei |
ICASSP | 6 |
| 2025 | Latent Space Chain-of-Embedding Enables Output-free LLM Self-EvaluationabstractLLM self-evaluation relies on the LLM's own ability to estimate response correctness, which can greatly improve its deployment reliability.
In this research track, we propose the Chain-of-Embedding (CoE) in the latent space to enable LLMs to perform output-free self-evaluation. CoE consists of all progressive hidden states produced during the inference time, which can be treated as the latent thinking path of LLMs. We find that when LLMs respond correctly and incorrectly, their CoE features differ, these discrepancies assist us in estimating LLM response correctness. Experiments in four diverse domains and seven LLMs fully demonstrate the effectiveness of our method. Meanwhile, its label-free design intent without any training and millisecond-level computational cost ensure real-time feedback in large-scale scenarios.
More importantly, we provide interesting insights into LLM response correctness from the perspective of hidden state changes inside LLMs. Yiming Wang 0011, Pei Zhang 0011, Baosong Yang, Derek F. Wong, Rui Wang 0015 |
ICLR | 5 |
| 2025 | RaSA: Rank-Sharing Low-Rank AdaptationabstractLow-rank adaptation (LoRA) has been prominently employed for parameter-efficient fine-tuning of large language models (LLMs). However, the limited expressive capacity of LoRA, stemming from the low-rank constraint, has been recognized as a bottleneck, particularly in rigorous tasks like code generation and mathematical reasoning. To address this limitation, we introduce Rank-Sharing Low-Rank Adaptation (RaSA), an innovative extension that enhances the expressive capacity of LoRA by leveraging partial rank sharing across layers. By forming a shared rank pool and applying layer-specific weighting, RaSA effectively increases the number of ranks without augmenting parameter overhead. Our theoretically grounded and empirically validated approach demonstrates that RaSA not only maintains the core advantages of LoRA but also significantly boosts performance in challenging code and math tasks. Code, data and scripts are available at: https://github.com/zwhe99/RaSA. Zhiwei He 0002, Zhaopeng Tu, Xing Wang 0007, Wenxiang Jiao, Zhuosheng Zhang 0001, Rui Wang 0015 |
ICLR | 10 |
| 2025 | Do Large Language Models Truly Understand Geometric Structures?abstractGeometric ability is a significant challenge for large language models (LLMs) due to the need for advanced spatial comprehension and abstract thinking. Existing datasets primarily evaluate LLMs on their final answers, but they cannot truly measure their true understanding of geometric structures, as LLMs can arrive at correct answers by coincidence. To fill this gap, we introduce the GeomRel dataset, designed to evaluate LLMs’ understanding of geometric structures by isolating the core step of geometric relationship identification in problem-solving. Using this benchmark, we conduct thorough evaluations of diverse LLMs and identify key limitations in understanding geometric structures. We further propose the Geometry Chain-of-Thought (GeoCoT) method, which enhances LLMs’ ability to identify geometric relationships, resulting in significant performance improvements. Yiming Wang 0011, Rui Wang 0015 |
ICLR | 4 |
| 2025 | Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned ModelabstractAligning language models (LMs) with human preferences has become a key area of research, enabling these models to meet diverse user needs better. Inspired by weak-to-strong generalization, where a strong LM fine-tuned on labels generated by a weaker model can consistently outperform its weak supervisor, we extend this idea to model alignment. In this work, we observe that the alignment behavior in weaker models can be effectively transferred to stronger models and even exhibit an amplification effect. Based on this insight, we propose a method called Weak-to-Strong Preference Optimization (WSPO), which achieves strong model alignment by learning the distribution differences before and after the alignment of the weak model. Experiments demonstrate that WSPO delivers outstanding performance, improving the win rate of Qwen2-7B-Instruct on Arena-Hard from 39.70 to 49.60, achieving a remarkable 47.04 length-controlled win rate on AlpacaEval 2, and scoring 7.33 on MT-bench. Our results suggest that using the weak model to elicit a strong model with a high alignment ability is feasible. The code is available at https://github.com/zwhong714/weak-to-strong-preference-optimization. Zhiwei He 0002, Rui Wang 0015 |
ICLR | 5 |
| 2025 | Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning ModelsabstractThe remarkable performance of long reasoning models can be attributed to their ability to emulate human-like long-time thinking during inference. These models employ extended chain-of-thought (CoT) processes, exploring multiple strategies to enhance problem-solving capabilities. However, a critical question remains: How to intelligently and efficiently scale computational resources during testing. This paper presents the first comprehensive study on the prevalent issue of overthinking in these models, where long reasoning models generate redundant solutions that contribute minimally to accuracy and diversity, thereby wasting computational resources on simple problems with minimal benefit. We introduce novel efficiency metrics from both outcome and process perspectives to evaluate the rational use of computational resources by long reasoning models. Using a self-training paradigm, we propose strategies to mitigate overthinking, simplifying reasoning processes without compromising accuracy. Experimental results show that our approach successfully reduces computational overhead while preserving model performance across a range of testsets with varying difficulty levels, such as GSM8K, MATH500, GPQA, and AIME. Our code is open-source and available at https://github.com/galaxyChen/overthinking. Zhiwei He 0002, Jianhui Pang, Dian Yu 0001, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
ICML | 11 |
| 2025 | Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering TasksabstractRecent advances in Large Language Models (LLMs) have shown promise in function-level code generation, yet repository-level software engineering tasks remain challenging. Current solutions predominantly rely on proprietary LLM agents, which introduce unpredictability and limit accessibility, raising concerns about data privacy and model customization. This paper investigates whether open-source LLMs can effectively address repository-level tasks without requiring agent-based approaches. We demonstrate this is possible by enabling LLMs to comprehend functions and files within codebases through their semantic information and structural dependencies. To this end, we introduce Code Graph Models (CGMs), which integrate repository code graph structures into the LLM's attention mechanism and map node attributes to the LLM's input space using a specialized adapter. When combined with an agentless graph RAG framework, our approach achieves a 43.00% resolution rate on the SWE-bench Lite benchmark using the open-source Qwen2.5-72B model. This performance ranks first among open weight models, second among methods with open-source systems, and eighth overall, surpassing the previous best open-source model-based method by 12.33%. Hongyuan Tao, Ying Zhang 0090, Zhenhao Tang, Hongen Peng, Xukun Zhu, Bingchang Liu, Yingguang Yang, Ziyin Zhang, Zhaogui Xu, Haipeng Zhang 0004, Linchao Zhu, Rui Wang 0015, Hang Yu 0002, Peng Di |
NeurIPS | 12 |
| 2025 | Thoughts Are All Over the Place: On the Underthinking of Long Reasoning ModelsabstractLong reasoning models (LRMs) such as OpenAI's o1 and DeepSeek's R1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term underthinking, where LRMs frequently switch between different reasoning thoughts without sufficiently exploring promising paths to reach a correct solution. This behavior leads to inadequate depth of reasoning and decreased performance, particularly on challenging mathematical problems. To systematically analyze this issue, we conduct experiments on three challenging test sets and two representative open-source LRMs, revealing that frequent thought switching correlates with incorrect responses. We introduce a novel metric to quantify underthinking by measuring token efficiency in incorrect answers. To address underthinking, we propose a decoding strategy with thought switching penalty (Tip) that discourages premature transitions between thoughts, encouraging deeper exploration of each reasoning path. Experimental results demonstrate that our approach improves accuracy across challenging datasets without requiring model fine-tuning. Our findings contribute to understanding reasoning inefficiencies in LRMs and offer a practical solution to enhance their problem-solving capabilities. Our code is open-source and available at https://github.com/wangyuenlp/underthinking. Yue Wang 0039, Qiuzhi Liu, Zhiwei He 0002, Linfeng Song, Dian Yu 0001, Juntao Li 0005, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
NeurIPS | 11 |
| 2025 | Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early DecodingabstractTest-time scaling enhances large language model performance by allocating additional compute resources during decoding. Best-of-$N$ (BoN) sampling serves as a common sampling-based scaling technique, broadening the search space in parallel to find better solutions from the model distribution. However, its cost–performance trade-off is still underexplored. Two main challenges limit the efficiency of BoN sampling:
(1) Generating $N$ full samples consumes substantial GPU memory, reducing inference capacity under limited resources.
(2) Reward models add extra memory and latency overhead, and training strong reward models introduces potential training data costs.
Although some studies have explored efficiency improvements, none have addressed both challenges at once.
To address this gap, we propose **Self-Truncation Best-of-$N$ (ST-BoN)**, a decoding method that avoids fully generating all $N$ samples and eliminates the need for reward models. It leverages early sampling consistency in the model’s internal states to identify the most promising path and truncate suboptimal ones.
In terms of cost, ST-BoN reduces dynamic GPU memory usage by over 80% and inference latency by 50%.
In terms of cost–performance trade-off, ST-BoN achieves the same performance as Full-BoN while saving computational cost by 70%–80%, and under the same cost, it can improve accuracy by 3–4 points. Yiming Wang 0011, Pei Zhang 0011, Baosong Yang, Zhuosheng Zhang 0001, Fei Huang 0002, Rui Wang 0015 |
NeurIPS | 7 |
| 2025 | PolyMath: Evaluating Mathematical Reasoning in Multilingual ContextsabstractIn this paper, we introduce PolyMath, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, and high-quality translation, making it a highly discriminative multilingual mathematical benchmark in the era of reasoning LLMs.We conduct a comprehensive evaluation for advanced LLMs and find that even Qwen-3-235B-A22B-Thinking and Gemini-2.5-pro, achieve only 54.6 and 52.2 benchmark scores, with about 40% accuracy under the highest level.From a language perspective, our benchmark reveals several key challenges of LLMs in multilingual reasoning:(1) Reasoning performance varies widely across languages for current LLMs;(2) Input-output language consistency is low in reasoning LLMs and may be correlated with performance;(3) The thinking length differs significantly by language for current LLMs.Additionally, we demonstrate that controlling the output language in the instructions has the potential to affect reasoning performance, especially for some low-resource languages, suggesting a promising direction for improving multilingual capabilities in LLMs. Yiming Wang 0011, Pei Zhang 0011, Jialong Tang, Baosong Yang, Rui Wang 0015, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang 0002, Junyang Lin, Fei Huang 0005, Jingren Zhou 0001 |
NeurIPS | 6 |
| 2024 | Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language ModelsabstractZhiwei He, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang, Zhaopeng Tu, Zhuosheng Zhang, Rui Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhiwei He 0002, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang 0007, Zhaopeng Tu, Zhuosheng Zhang 0001, Rui Wang 0015 |
ACL (1) | 8 |
| 2024 | MELA: Multilingual Evaluation of Linguistic AcceptabilityabstractIn this work, we present the largest benchmark to date on linguistic acceptability: Multilingual Evaluation of Linguistic Acceptability-MELA, with 46K samples covering 10 languages from a diverse set of language families.We establish LLM baselines on this benchmark, and investigate cross-lingual transfer in acceptability judgements with XLM-R.In pursuit of multilingual interpretability, we conduct probing experiments with fine-tuned XLM-R to explore the process of syntax capability acquisition.Our results show that GPT-4o exhibits a strong multilingual ability, outperforming fine-tuned XLM-R, while open-source multilingual models lag behind by a noticeable gap.Cross-lingual transfer experiments show that transfer in acceptability judgment is non-trivial: 500 Icelandic fine-tuning examples lead to 23 MCC performance in a completely unrelated language-Chinese.Results of our probing experiments indicate that training on MELA improves the performance of XLM-R on syntaxrelated tasks. Ziyin Zhang, Yikang Liu 0002, Weifang Huang, Junyu Mao, Rui Wang 0015, Hai Hu 0001 |
ACL (1) | 5 |
| 2024 | Towards Human-Like Machine Comprehension: Few-Shot Relational Learning in Visually-Rich DocumentsabstractKey-value relations are prevalent in Visually-Rich Documents (VRDs), often depicted in distinct spatial regions accompanied by specific color and font styles. These non-textual cues serve as important indicators that greatly enhance human comprehension and acquisition of such relation triplets. However, current document AI approaches often fail to consider this valuable prior information related to visual and spatial features, resulting in suboptimal performance, particularly when dealing with limited examples. To address this limitation, our research focuses on few-shot relational learning, specifically targeting the extraction of key-value relation triplets in VRDs. Given the absence of a suitable dataset for this task, we introduce two new few-shot benchmarks built upon existing supervised benchmark datasets. Furthermore, we propose a variational approach that incorporates relational 2D-spatial priors and prototypical rectification techniques. This approach aims to generate relation representations that are more aware of the spatial context and unseen relation in a manner similar to human perception. Experimental results demonstrate the effectiveness of our proposed method by showcasing its ability to outperform existing methods. This study also opens up new possibilities for practical applications. Hao Wang 0097, Chenhui Chu, Rui Wang 0015, Pinpin Zhu |
LREC/COLING | 4 |
| 2024 | Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebateabstractModern large language models (LLMs) like ChatGPT have shown remarkable performance on general language tasks but still struggle on complex reasoning tasks, which drives the research on cognitive behaviors of LLMs to explore human-like problem-solving strategies.Along this direction, one representative strategy is self-reflection, which asks an LLM to refine the solution with the feedback generated by itself iteratively.However, our study shows that such reflection-style methods suffer from the Degeneration-of-Thought (DoT) problem: once the LLM has established confidence in its solutions, it is unable to generate novel thoughts later through reflection even if its initial stance is incorrect.To address the DoT problem, we propose a Multi-Agent Debate (MAD) framework, in which multiple agents express their arguments in the state of "tit for tat" and a judge manages the debate process to obtain a final solution.Clearly, our MAD framework encourages divergent thinking in LLMs which would be helpful for tasks that require deep levels of contemplation.Experiment results on two challenging datasets, commonsense machine translation and counterintuitive arithmetic reasoning, demonstrate the effectiveness of our MAD framework.Extensive analyses suggest that the adaptive break of debate and the modest level of "tit for tat" state are required for MAD to obtain good performance.Moreover, we find that LLMs might not be a fair judge if different LLMs are used for agents.Code is available at https://github. com/Skytliang/Multi-Agents-Debate. Zhiwei He 0002, Wenxiang Jiao, Xing Wang 0007, Yan Wang 0060, Rui Wang 0015, Yujiu Yang 0001, Shuming Shi 0001, Zhaopeng Tu |
EMNLP | 6 |
| 2024 | Improving Open-Ended Text Generation via Adaptive DecodingabstractCurrent language models decode text token by token according to probabilistic distribution, and determining the appropriate candidates for the next token is crucial to ensure generation quality. This study introduces adaptive decoding, a mechanism that dynamically empowers language models to ascertain a sensible candidate set during generation. Specifically, we introduce an entropy-based metric called confidence and conceptualize determining the optimal candidate set as a confidence-increasing process. The rationality of including a token in the candidate set is assessed by leveraging the increment of confidence. Experimental results reveal that our method balances diversity and coherence well. The human evaluation shows that our method can generate human-preferred text. Additionally, our method can potentially improve the reasoning ability of language models. Hongkun Hao, Zhiwei He 0002, Yiming Ai, Rui Wang 0015 |
ICML | 5 |
| 2024 | Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward ModelabstractZhiwei He, Xing Wang, Wenxiang Jiao, Zhuosheng Zhang, Rui Wang, Shuming Shi, Zhaopeng Tu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhiwei He 0002, Xing Wang 0007, Wenxiang Jiao, Zhuosheng Zhang 0001, Rui Wang 0015, Shuming Shi 0001, Zhaopeng Tu |
NAACL-HLT | 5 |
| 2024 | Embedding Trajectory for Out-of-Distribution Detection in Mathematical ReasoningabstractReal-world data deviating from the independent and identically distributed (\textit{i.i.d.}) assumption of in-distribution training data poses security threats to deep networks, thus advancing out-of-distribution (OOD) detection algorithms. Detection methods in generative language models (GLMs) mainly focus on uncertainty estimation and embedding distance measurement, with the latter proven to be most effective in traditional linguistic tasks like summarization and translation. However, another complex generative scenario mathematical reasoning poses significant challenges to embedding-based methods due to its high-density feature of output spaces, but this feature causes larger discrepancies in the embedding shift trajectory between different samples in latent spaces. Hence, we propose a trajectory-based method TV score, which uses trajectory volatility for OOD detection in mathematical reasoning. Experiments show that our method outperforms all traditional algorithms on GLMs under mathematical reasoning scenarios and can be extended to more applications with high-density features in output spaces, such as multiple-choice questions. Yiming Wang 0011, Pei Zhang 0011, Baosong Yang, Derek F. Wong, Zhuosheng Zhang 0001, Rui Wang 0015 |
NeurIPS | 6 |
| 2024 | Exploring Human-Like Translation Strategy with Large Language ModelsabstractAbstract Large language models (LLMs) have demonstrated impressive capabilities in general scenarios, exhibiting a level of aptitude that approaches, in some aspects even surpasses, human-level intelligence. Among their numerous skills, the translation abilities of LLMs have received considerable attention. Compared to typical machine translation that focuses solely on source-to-target mapping, LLM-based translation can potentially mimic the human translation process, which might take preparatory steps to ensure high-quality translation. This work explores this possibility by proposing the MAPS framework, which stands for Multi-Aspect Prompting and Selection. Specifically, we enable LLMs first to analyze the given source sentence and induce three aspects of translation-related knowledge (keywords, topics, and relevant demonstrations) to guide the final translation process. Moreover, we employ a selection mechanism based on quality estimation to filter out noisy and unhelpful knowledge. Both automatic (3 LLMs × 11 directions × 2 automatic metrics) and human evaluation (preference study and MQM) demonstrate the effectiveness of MAPS. Further analysis shows that by mimicking the human translation process, MAPS reduces various translation errors such as hallucination, ambiguity, mistranslation, awkward style, untranslated text, and omission. Source code is available at https://github.com/zwhe99/MAPS-mt. Zhiwei He 0002, Wenxiang Jiao, Zhuosheng Zhang 0001, Yujiu Yang 0001, Rui Wang 0015, Zhaopeng Tu, Shuming Shi 0001, Xing Wang 0007 |
Trans. Assoc. Comput. Linguistics | 6 |
| 2023 | Element-aware Summarization with Large Language Models: Expert-aligned Evaluation and Chain-of-Thought MethodabstractAutomatic summarization generates concise summaries that contain key ideas of source documents.As the most mainstream datasets for the news sub-domain, CNN/DailyMail and BBC XSum have been widely used for performance benchmarking.However, the reference summaries of those datasets turn out to be noisy, mainly in terms of factual hallucination and information redundancy.To address this challenge, we first annotate new expertwriting Element-aware test sets following the "Lasswell Communication Model" proposed by Lasswell (1948), allowing reference summaries to focus on more fine-grained news elements objectively and comprehensively.Utilizing the new test sets, we observe the surprising zero-shot summary ability of LLMs, which addresses the issue of the inconsistent results between human preference and automatic evaluation metrics of LLMs' zero-shot summaries in prior work.Further, we propose a Summary Chain-of-Thought (SumCoT) technique to elicit LLMs to generate summaries step by step, which helps them integrate more finegrained details of source documents into the final summaries that correlate with the human writing mindset.Experimental results show our method outperforms state-of-the-art fine-tuned PLMs and zero-shot LLMs by +4.33/+4.77 in ROUGE-L on the two datasets, respectively.Dataset and code are publicly available at https://github.com/Alsace08/SumCoT. Yiming Wang 0011, Zhuosheng Zhang 0001, Rui Wang 0015 |
ACL (1) | 3 |
| 2023 | Rethinking Word-Level Auto-Completion in Computer-Aided TranslationabstractWord-Level Auto-Completion (WLAC) plays a crucial role in Computer-Assisted Translation.It aims at providing word-level autocompletion suggestions for human translators.While previous studies have primarily focused on designing complex model architectures, this paper takes a different perspective by rethinking the fundamental question: what kind of words are good auto-completions?We introduce a measurable criterion to answer this question and discover that existing WLAC models often fail to meet this criterion.Building upon this observation, we propose an effective approach to enhance WLAC performance by promoting adherence to the criterion.Notably, the proposed approach is general and can be applied to various encoder-based architectures.Through extensive experiments, we demonstrate that our approach outperforms the top-performing system submitted to the WLAC shared tasks in WMT2022, while utilizing significantly smaller model sizes ¶ . Lemao Liu, Guoping Huang, Zhirui Zhang, Shuming Shi 0001, Rui Wang 0015 |
EMNLP | 7 |
| 2023 | Nearest Neighbor Machine Translation is Meta-Optimizer on Output Projection LayerabstractNearest Neighbor Machine Translation (kNN-MT) has achieved great success in domain adaptation tasks by integrating pre-trained Neural Machine Translation (NMT) models with domain-specific token-level retrieval.However, the reasons underlying its success have not been thoroughly investigated.In this paper, we comprehensively analyze kNN-MT through theoretical and empirical studies.Initially, we provide new insights into the working mechanism of kNN-MT as an efficient technique to implicitly execute gradient descent on the output projection layer of NMT, indicating that it is a specific case of model fine-tuning.Subsequently, we conduct multi-domain experiments and word-level analysis to examine the differences in performance between kNN-MT and entire-model fine-tuning.Our findings suggest that: (i) Incorporating kNN-MT with adapters yields comparable translation performance to fine-tuning on in-domain test sets, while achieving better performance on out-of-domain test sets; (ii) Fine-tuning significantly outperforms kNN-MT on the recall of in-domain low-frequency words, but this gap could be bridged by optimizing the context representations with additional adapter layers. 1 Zhirui Zhang, Yichao Du, Lemao Liu, Rui Wang 0015 |
EMNLP | 5 |
| 2023 | Vision-Enhanced Semantic Entity Recognition in Document Images via Visually-Asymmetric Consistency LearningabstractExtracting meaningful entities belonging to predefined categories from Visually-rich Formlike Documents (VFDs) is a challenging task.Visual and layout features such as font, background, color, and bounding box location and size provide important cues for identifying entities of the same type.However, existing models commonly train a visual encoder with weak cross-modal supervision signals, resulting in a limited capacity to capture these nontextual features and suboptimal performance.In this paper, we propose a novel Visually-Asymmetric coNsistenCy Learning (VANCL) approach that addresses the above limitation by enhancing the model's ability to capture finegrained visual and layout features through the incorporation of color priors.Experimental results on benchmark datasets show that our approach substantially outperforms the strong LayoutLM series baseline, demonstrating the effectiveness of our approach.Additionally, we investigate the effects of different color schemes on our approach, providing insights for optimizing model performance.We believe our work will inspire future research on multimodal information extraction. Hao Wang 0097, Xiahua Chen, Rui Wang 0015, Chenhui Chu |
EMNLP | 3 |
| 2023 | Penalty Decoding: Well Suppress the Self-Reinforcement Effect in Open-Ended Text GenerationabstractThe decoding algorithm is critical for openended text generation, transforming latent representations into coherent and meaningful outputs.This paper investigates the selfreinforcement effect in text generation and the effectiveness of a repetition penalty to mitigate it.However, determining the optimal repetition penalty value is challenging.To tackle this, we propose a forgetting mechanism that disregards distant tokens, reducing the burden of penalty selection.In addition, we introduce a length penalty to address overly short sentences caused by excessive penalties.Our penalty decoding approach incorporating three strategies helps resolve issues with sampling methods deviating from factual information.Experimental results demonstrate the efficacy of our approach in generating high-quality sentences resembling human output. 1 Hongkun Hao, Rui Wang 0015 |
EMNLP | 3 |
| 2023 | Imitation Attacks Can Steal More Than You Think from Machine Translation Systems
Tianxiang Hu, Pei Zhang 0011, Baosong Yang, Rui Wang 0015 |
NLPCC (1) | 5 |
| 2023 | Universal Multimodal Representation for Language UnderstandingabstractRepresentation learning is the foundation of natural language processing (NLP). This work presents new methods to employ visual information as assistant signals to general NLP tasks. For each sentence, we first retrieve a flexible number of images either from a light topic-image lookup table extracted over the existing sentence-image pairs or a shared cross-modal embedding space that is pre-trained on out-of-shelf text-image pairs. Then, the text and images are encoded by a Transformer encoder and convolutional neural network, respectively. The two sequences of representations are further fused by an attention layer for the interaction of the two modalities. In this study, the retrieval process is controllable and flexible. The universal visual representation overcomes the lack of large-scale bilingual sentence-image pairs. Our method can be easily applied to text-only tasks without manually annotated multimodal parallel corpora. We apply the proposed method to a wide range of natural language generation and understanding tasks, including neural machine translation, natural language inference, and semantic similarity. Experimental results show that our method is generally effective for different tasks and languages. Analysis indicates that the visual signals enrich textual representations of content words, provide fine-grained grounding information about the relationship between concepts and events, and potentially conduce to disambiguation. Zhuosheng Zhang 0001, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Zuchao Li, Hai Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Bridging the Data Gap between Training and Inference for Unsupervised Neural Machine TranslationabstractBack-translation is a critical component of Unsupervised Neural Machine Translation (UNMT), which generates pseudo parallel data from target monolingual data.A UNMT model is trained on the pseudo parallel data with translated source, and translates natural source sentences in inference.The source discrepancy between training and inference hinders the translation performance of UNMT models.By carefully designing experiments, we identify two representative characteristics of the data gap in source: (1) style gap (i.e., translated vs. natural text style) that leads to poor generalization capability; (2) content gap that induces the model to produce hallucination content biased towards the target language.To narrow the data gap, we propose an online self-training approach, which simultaneously uses the pseudo parallel data {natural source, translated target} to mimic the inference scenario.Experimental results on several widelyused language pairs show that our approach outperforms two strong baselines (XLM and MASS) by remedying the style and content gaps. 1 Model En-Fr En-De En-Ro Avg.⇒ ⇐ ⇒ ⇐ ⇒ ⇐ Full Test Set SNMT 38.4 33.6 29.5 33.9 33.7 32.5 33.6 XLM 37.4 34.5 27.2 34.3 34.6 32.7 33.5 MASS 37.8 34.9 27.1 35.2 35.1 33.4 33.9 Model En-Fr En-De En-Ro Avg.⇒ ⇐ ⇒ ⇐ ⇒ ⇐ Full Test Set SNMT 37.3 33.4 29.7 33.8 33.8 32.4 33.4 XLM 36.3 34.3 27.4 34.1 34.8 32.4 33.2 MASS 36.6 34.7 27.3 35.1 35.2 33.0 33.7 Zhiwei He 0002, Xing Wang 0007, Rui Wang 0015, Shuming Shi 0001, Zhaopeng Tu |
ACL (1) | 3 |
| 2022 | Effective Graph Context Representation for Document-level Machine TranslationabstractDocument-level neural machine translation (DocNMT) universally encodes several local sentences or the entire document. Thus, DocNMT does not consider the relevance of document-level contextual information, for example, some context (i.e., content words, logical order, and co-occurrence relation) is more effective than another auxiliary context (i.e., functional and auxiliary words). To address this issue, we first utilize the word frequency information to recognize content words in the input document, and then use heuristical relations to summarize content words and sentences as a graph structure without relying on external syntactic knowledge. Furthermore, we apply graph attention networks to this graph structure to learn its feature representation, which allows DocNMT to more effectively capture the document-level context. Experimental results on several widely-used document-level benchmarks demonstrated the effectiveness of the proposed approach. Kehai Chen, Muyun Yang, Masao Utiyama, Eiichiro Sumita, Rui Wang 0015, Min Zhang 0005 |
IJCAI | 5 |
| 2022 | Text Compression-Aided Transformer EncodingabstractText encoding is one of the most important steps in Natural Language Processing (NLP). It has been done well by the self-attention mechanism in the current state-of-the-art Transformer encoder, which has brought about significant improvements in the performance of many NLP tasks. Though the Transformer encoder may effectively capture general information in its resulting representations, the backbone information, meaning the gist of the input text, is not specifically focused on. In this paper, we propose explicit and implicit text compression approaches to enhance the Transformer encoding and evaluate models using this approach on several typical downstream tasks that rely on the encoding heavily. Our explicit text compression approaches use dedicated models to compress text, while our implicit text compression approach simply adds an additional module to the main model to handle text compression. We propose three ways of integration, namely backbone source-side fusion, target-side fusion, and both-side fusion, to integrate the backbone information into Transformer-based models for various downstream tasks. Our evaluation on benchmark datasets shows that the proposed explicit and implicit text compression approaches improve results in comparison to strong baselines. We therefore conclude, when comparing the encodings to the baseline models, text compression helps the encoders to learn better language representations. Zuchao Li, Zhuosheng Zhang 0001, Hai Zhao 0001, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | SG-Net: Syntax Guided Transformer for Language RepresentationabstractUnderstanding human language is one of the key themes of artificial intelligence. For language representation, the capacity of effectively modeling the linguistic knowledge from the detail-riddled and lengthy texts and getting ride of the noises is essential to improve its performance. Traditional attentive models attend to all words without explicit constraint, which results in inaccurate concentration on some dispensable words. In this work, we propose using syntax to guide the text modeling by incorporating explicit syntactic constraints into attention mechanisms for better linguistically motivated word representations. In detail, for self-attention network (SAN) sponsored Transformer-based encoder, we introduce syntactic dependency of interest (SDOI) design into the SAN to form an SDOI-SAN with syntax-guided self-attention. Syntax-guided network (SG-Net) is then composed of this extra SDOI-SAN and the SAN from the original Transformer encoder through a dual contextual architecture for better linguistics inspired representation. The proposed SG-Net is applied to typical Transformer encoders. Extensive experiments on popular benchmark tasks, including machine reading comprehension, natural language inference, and neural machine translation show the effectiveness of the proposed SG-Net design. Zhuosheng Zhang 0001, Yuwei Wu 0003, Junru Zhou, Sufeng Duan, Hai Zhao 0001, Rui Wang 0015 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Tri-training for Dependency Parsing Domain AdaptationabstractIn recent years, the research on dependency parsing focuses on improving the accuracy of the domain-specific (in-domain) test datasets and has made remarkable progress. However, there are innumerable scenarios in the real world that are not covered by the dataset, namely, the out-of-domain dataset. As a result, parsers that perform well on the in-domain data usually suffer from significant performance degradation on the out-of-domain data. Therefore, to adapt the existing in-domain parsers with high performance to a new domain scenario, cross-domain transfer learning methods are essential to solve the domain problem in parsing. This paper examines two scenarios for cross-domain transfer learning: semi-supervised and unsupervised cross-domain transfer learning. Specifically, we adopt a pre-trained language model BERT for training on the source domain (in-domain) data at the subword level and introduce self-training methods varied from tri-training for these two scenarios. The evaluation results on the NLPCC-2019 shared task and universal dependency parsing task indicate the effectiveness of the adopted approaches on cross-domain transfer learning and show the potential of self-learning to cross-lingual transfer learning. Zuchao Li, Hai Zhao 0001, Bao-Liang Lu, Rui Wang 0015 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2022 | Integrating Prior Translation Knowledge Into Neural Machine TranslationabstractNeural machine translation (NMT), which is an encoder-decoder joint neural language model with an attention mechanism, has achieved impressive results on various machine translation tasks in the past several years. However, the language model attribute of NMT tends to produce fluent yet sometimes unfaithful translations, which hinders the improvement of translation capacity. In response to this problem, we propose a simple and efficient method to integrate prior translation knowledge into NMT in a universal manner that is compatible with neural networks. Meanwhile, it enables NMT to consider the crossing language translation knowledge from the source-side of the training pipeline of NMT, thereby making full use of the prior translation knowledge to enhance the performance of NMT. The experimental results on two large-scale benchmark translation tasks demonstrated that our approach achieved a significant improvement over a strong baseline. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Self-Training for Unsupervised Neural Machine Translation in Unbalanced Training Data ScenariosabstractHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
NAACL-HLT | 2 |
| 2021 | Context-aware positional representation for self-attention networks
Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
Neurocomputing | 2 |
| 2021 | Unsupervised Neural Machine Translation for Similar and Distant Language Pairs: An Empirical StudyabstractUnsupervised neural machine translation (UNMT) has achieved remarkable results for several language pairs, such as French–English and German–English. Most previous studies have focused on modeling UNMT systems; few studies have investigated the effect of UNMT on specific languages. In this article, we first empirically investigate UNMT for four diverse language pairs (French/German/Chinese/Japanese–English). We confirm that the performance of UNMT in translation tasks for similar language pairs (French/German–English) is dramatically better than for distant language pairs (Chinese/Japanese–English). We empirically show that the lack of shared words and different word orderings are the main reasons that lead UNMT to underperform in Chinese/Japanese–English. Based on these findings, we propose several methods, including artificial shared words and pre-ordering, to improve the performance of UNMT for distant language pairs. Moreover, we propose a simple general method to improve translation performance for all these four language pairs. The existing UNMT model can generate a translation of a reasonable quality after a few training epochs owing to a denoising mechanism and shared latent representations. However, learning shared latent representations restricts the performance of translation in both directions, particularly for distant language pairs, while denoising dramatically delays convergence by continuously modifying the training data. To avoid these problems, we propose a simple, yet effective and efficient, approach that (like UNMT) relies solely on monolingual corpora: pseudo-data-based unsupervised neural machine translation. Experimental results for these four language pairs show that our proposed methods significantly outperform UNMT baselines. Haipeng Sun, Rui Wang 0015, Masao Utiyama, Benjamin Marie, Kehai Chen, Eiichiro Sumita, Tiejun Zhao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2021 | Modeling Future Cost for Neural Machine TranslationabstractExisting neural machine translation (NMT) systems utilize sequence-to-sequence neural networks to generate target translation word by word, and then make the generated word at each time-step and the counterpart in the references as consistent as possible. However, the trained translation model tends to focus on ensuring the accuracy of the generated target word at the current time-step and does not consider its future cost which means the expected cost of generating the subsequent target translation (i.e., the next target word). To respond to this issue, in this article, we propose a simple and effective method to model the future cost of each target word for NMT systems. In detail, a future cost representation is learned based on the current generated target word and its contextual information to compute an additional loss to guide the training of the NMT model. Furthermore, the learned future cost representation at the current time-step is used to help the generation of the next target word in the decoding. Experimental results on three widely-used translation datasets, including the WMT14 English-to-German, WMT14 English-to-French, and WMT17 Chinese-to-English, show that the proposed approach achieves significant improvements over strong Transformer-based NMT baseline. Chaoqun Duan, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Conghui Zhu, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Detecting Source Contextual Barriers for Understanding Neural Machine TranslationabstractIn machine translation evaluation, the traditional wisdom measures model's generalization ability in an average sense, for example by using corpus BLEU. However, the statistics of corpus BLEU cannot provide comprehensive understanding and fine-grained analysis on model's generalization ability. As a remedy, this paper attempts to understand NMT at fine-grained level, by detecting contextual barriers within an unseen input sentence that \textit{cause} the degradation in model's translation quality. It proposes a principled definition of source contextual barriers as well as its modified version which is tractable in computation and operates at word-level. Based on the modified one, three simple methods are proposed for barrier detection by search-aware risk estimation through counterfactual generation. Extensive analyses are conducted on those detected contextual barrier words on both Zh$\Leftrightarrow$En NIST benchmarks. Potential usages motivated from barrier words are also discussed. Lemao Liu, Conghui Zhu, Rui Wang 0015, Tiejun Zhao, Shuming Shi 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | SG-Net: Syntax-Guided Machine Reading ComprehensionabstractFor machine reading comprehension, the capacity of effectively modeling the linguistic knowledge from the detail-riddled and lengthy passages and getting ride of the noises is essential to improve its performance. Traditional attentive models attend to all words without explicit constraint, which results in inaccurate concentration on some dispensable words. In this work, we propose using syntax to guide the text modeling by incorporating explicit syntactic constraints into attention mechanism for better linguistically motivated word representations. In detail, for self-attention network (SAN) sponsored Transformer-based encoder, we introduce syntactic dependency of interest (SDOI) design into the SAN to form an SDOI-SAN with syntax-guided self-attention. Syntax-guided network (SG-Net) is then composed of this extra SDOI-SAN and the SAN from the original Transformer encoder through a dual contextual architecture for better linguistics inspired representation. To verify its effectiveness, the proposed SG-Net is applied to typical pre-trained language model BERT which is right based on a Transformer encoder. Extensive experiments on popular benchmarks including SQuAD 2.0 and RACE show that the proposed SG-Net design helps achieve substantial performance improvement over strong baselines. Zhuosheng Zhang 0001, Yuwei Wu 0003, Junru Zhou, Sufeng Duan, Hai Zhao 0001, Rui Wang 0015 |
AAAI | 6 |
| 2020 | Explicit Sentence Compression for Neural Machine TranslationabstractState-of-the-art Transformer-based neural machine translation (NMT) systems still follow a standard encoder-decoder framework, in which source sentence representation can be well done by an encoder with self-attention mechanism. Though Transformer-based encoder may effectively capture general information in its resulting source sentence representation, the backbone information, which stands for the gist of a sentence, is not specifically focused on. In this paper, we propose an explicit sentence compression method to enhance the source sentence representation for NMT. In practice, an explicit sentence compression goal used to learn the backbone information in a sentence. We propose three ways, including backbone source-side fusion, target-side fusion, and both-side fusion, to integrate the compressed sentence into NMT. Our empirical tests on the WMT English-to-French and English-to-German translation tasks show that the proposed sentence compression method significantly improves the translation performances over strong baselines. Zuchao Li, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Zhuosheng Zhang 0001, Hai Zhao 0001 |
AAAI | 2 |
| 2020 | Content Word Aware Neural Machine TranslationabstractNeural machine translation (NMT) encodes the source sentence in a universal way to generate the target sentence word-byword.However, NMT does not consider the importance of word in the sentence meaning, for example, some words (i.e., content words) express more important meaning than others (i.e., function words).To address this limitation, we first utilize word frequency information to distinguish between content and function words in a sentence, and then design a content word-aware NMT to improve translation performance.Empirical results on the WMT14 English-to-German, WMT14 English-to-French, and WMT17 Chineseto-English translation tasks show that the proposed methods can significantly improve the performance of Transformer-based NMT. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
ACL | 2 |
| 2020 | Regularized Context Gates on Transformer for Machine TranslationabstractContext gates are effective to control the contributions from the source and target contexts in the recurrent neural network (RNN) based neural machine translation (NMT).However, it is challenging to extend them into the advanced Transformer architecture, which is more complicated than RNN.This paper first provides a method to identify source and target contexts and then introduce a gate mechanism to control the source and target contributions in Transformer.In addition, to further reduce the bias problem in the gate mechanism, this paper proposes a regularization method to guide the learning of the gates with supervision automatically generated using pointwise mutual information.Extensive experiments on 4 translation datasets demonstrate that the proposed model obtains an averaged gain of 1.0 BLEU score over a strong Transformer baseline. Lemao Liu, Rui Wang 0015, Guoping Huang, Max Meng |
ACL | 3 |
| 2020 | Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationabstractUnsupervised neural machine translation (UNMT) has recently achieved remarkable results for several language pairs. However, it can only translate between a single language pair and cannot produce translation results for multiple language pairs at the same time. That is, research on multilingual UNMT has been limited. In this paper, we empirically introduce a simple method to translate between thirteen languages using a single encoder and a single decoder, making use of multilingual data to improve UNMT for all language pairs. On the basis of the empirical findings, we propose two knowledge distillation methods to further enhance multilingual UNMT performance. Our experiments on a dataset with English translated to and from twelve other languages (including three language families and six language branches) show remarkable results, surpassing strong unsupervised individual baselines while achieving promising performance between non-English language pairs in zero-shot translation scenarios and alleviating poor performance in low-resource language pairs. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
ACL | 2 |
| 2020 | Robust Unsupervised Neural Machine Translation with Adversarial Denoising TrainingabstractUnsupervised neural machine translation (UNMT) has recently attracted great interest in the machine translation community.The main advantage of the UNMT lies in its easy collection of required large training text sentences while with only a slightly worse performance than supervised neural machine translation which requires expensive annotated translation pairs on some translation tasks.In most studies, the UMNT is trained with clean data without considering its robustness to the noisy data.However, in real-world scenarios, there usually exists noise in the collected input sentences which degrades the performance of the translation system since the UNMT is sensitive to the small perturbations of the input sentences.In this paper, we first time explicitly take the noisy data into consideration to improve the robustness of the UNMT based systems.First of all, we clearly defined two types of noises in training sentences, i.e., word noise and word order noise, and empirically investigate its effect in the UNMT, then we propose adversarial training methods with denoising process in the UNMT.Experimental results on several language pairs show that our proposed methods substantially improved the robustness of the conventional UNMT systems in noisy scenarios. Haipeng Sun, Rui Wang 0015, Kehai Chen, Xugang Lu, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
COLING | 2 |
| 2020 | Neural Machine Translation with Universal Visual Representation
Zhuosheng Zhang 0001, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Zuchao Li, Hai Zhao 0001 |
ICLR | 3 |
| 2020 | Data-dependent Gaussian Prior Objective for Language Generation
Zuchao Li, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Zhuosheng Zhang 0001, Hai Zhao 0001 |
ICLR | 2 |
| 2020 | Towards More Diverse Input Representation for Neural Machine TranslationabstractSource input information plays a very important role in the Transformer-based translation system. In practice, word embedding and positional embedding of each word are added as the input representation. Then self-attention networks are used to encode the global dependencies in the input representation to generate a source representation. However, this processing on the source representation only adopts a single source feature and excludes richer and more diverse features such as recurrence features, local features, and syntactic features, which results in tedious representation and thereby hinders the further translation performance improvement. In this paper, we introduce a simple and efficient method to encode more diverse source features into the input representation simultaneously, and thereby learning an effective source representation by self-attention networks. In particular, the proposed grouped strategy is only applied to the input representation layer, to keep the diversity of translation information and the efficiency of the self-attention networks at the same time. Experimental results show that our approach improves the translation performance over the state-of-the-art baselines of Transformer in regard to WMT14 English-to-German and NIST Chinese-to-English machine translation tasks. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao, Muyun Yang, Hai Zhao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Memory Network for Linguistic Structure ParsingabstractMemory-based learning can be characterized as a lazy learning method in machine learning terminology because it delays the processing of input by storing the input until needed. Linguistic structure parsing, which has been in a performance improvement bottleneck since the latest series of works was presented, determines the syntactic or semantic structure of a sentence. In this article, we construct a memory component and use it to augment a linguistic structure parser which allows the parser to directly extract patterns from the known training treebank to form memory. The experimental results show that existing state-of-the-art parsers reach new heights of performance on the main benchmarks for dependency parsing and semantic role labeling with this memory network. Zuchao Li, Chaoyu Guan, Hai Zhao 0001, Rui Wang 0015, Kevin Parnow, Zhuosheng Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Unsupervised Neural Machine Translation With Cross-Lingual Language Representation AgreementabstractUnsupervised cross-lingual language representation initialization methods such as unsupervised bilingual word embedding (UBWE) pre-training and cross-lingual masked language model (CMLM) pre-training, together with mechanisms such as denoising and back-translation, have advanced unsupervised neural machine translation (UNMT), which has achieved impressive results on several language pairs, particularly French-English and German-English. Typically, UBWE focuses on initializing the word embedding layer in the encoder and decoder of UNMT, whereas the CMLM focuses on initializing the entire encoder and decoder of UNMT. However, UBWE/CMLM training and UNMT training are independent, which makes it difficult to assess how the quality of UBWE/CMLM affects the performance of UNMT during UNMT training. In this paper, we first empirically explore relationships between UNMT and UBWE/CMLM. The empirical results demonstrate that the performance of UBWE and CMLM has a significant influence on the performance of UNMT. Motivated by this, we propose a novel UNMT structure with cross-lingual language representation agreement to capture the interaction between UBWE/CMLM and UNMT during UNMT training. Experimental results on several language pairs demonstrate that the proposed UNMT models improve significantly over the corresponding state-of-the-art UNMT baselines. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | A Novel Sentence-Level Agreement Architecture for Neural Machine TranslationabstractIn neural machine translation (NMT), there is a natural correspondence between source and target sentences. The traditional NMT method does not explicitly model the translation agreement on sentence-level. In this article, we propose a comprehensive and novel sentence-level agreement architecture to alleviate this problem. It directly minimizes the difference between the representations of the source-side and target-side sentence on sentence-level. First, we compare a variety of sentence representation strategies and propose a “Gated Sum” sentence representation to achieve better sentence semantic information. Then, rather than a single-layer sentence-level agreement architecture, we further propose a multi-layer sentence agreement architecture to make the source and target semantic spaces closer layer by layer. The proposed agreement module can be integrated into NMT as an additional training objective function, and can also be used to enhance the representation of the source-side sentences. Experiments on the NIST Chinese-to-English and the WMT English-to-German translation tasks show that the proposed agreement architecture achieves significant improvements over state-of-the-art baselines, demonstrating the effectiveness and necessity of exploiting sentence-level agreement for NMT. Rui Wang 0015, Kehai Chen, Xing Wang 0007, Tiejun Zhao, Min Zhang 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Neural Machine Translation with Reordering EmbeddingsabstractThe reordering model plays an important role in phrase-based statistical machine translation.However, there are few works that exploit the reordering information in neural machine translation.In this paper, we propose a reordering mechanism to learn the reordering embedding of a word based on its contextual information.These reordering embeddings are stacked together with self-attention networks to learn sentence representation for machine translation.The reordering mechanism can be easily integrated into both the encoder and the decoder in the Transformer translation system.Experimental results on WMT'14 English-to-German, NIST Chinese-to-English, and WAT ASPEC Japanese-to-English translation tasks demonstrate that the proposed methods can significantly improve the performance of the Transformer translation system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
ACL (1) | 2 |
| 2019 | Unsupervised Bilingual Word Embedding Agreement for Unsupervised Neural Machine TranslationabstractUnsupervised bilingual word embedding (UBWE), together with other technologies such as back-translation and denoising, has helped unsupervised neural machine translation (UNMT) achieve remarkable results in several language pairs.In previous methods, UBWE is first trained using nonparallel monolingual corpora and then this pre-trained UBWE is used to initialize the word embedding in the encoder and decoder of UNMT.That is, the training of UBWE and UNMT are separate.In this paper, we first empirically investigate the relationship between UBWE and UNMT.The empirical findings show that the performance of UNMT is significantly affected by the performance of UBWE.Thus, we propose two methods that train UNMT with UBWE agreement.Empirical results on several language pairs show that the proposed methods significantly outperform conventional UNMT. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
ACL (1) | 2 |
| 2019 | Lattice-Based Transformer Encoder for Neural Machine TranslationabstractNeural machine translation (NMT) takes deterministic sequences for source representations.However, either wordlevel or subword-level segmentations have multiple choices to split a source sequence with different word segmentors or different subword vocabulary sizes.We hypothesize that the diversity in segmentations may affect the NMT performance.To integrate different segmentations with the state-of-the-art NMT model, Transformer, we propose lattice-based encoders to explore effective word or subword representation in an automatic way during training.We propose two methods: 1) lattice positional encoding and 2) lattice-aware self-attention.These two methods can be used together and show complementary to each other to further improve translation performance.Experiment results show superiorities of lattice-based encoders in word-level and subword-level representations over conventional Transformer encoder. Fengshun Xiao, Jiangtong Li, Hai Zhao 0001, Rui Wang 0015, Kehai Chen |
ACL (1) | 4 |
| 2019 | Sentence-Level Agreement for Neural Machine TranslationabstractThe training objective of neural machine translation (NMT) is to minimize the loss between the words in the translated sentences and those in the references. In NMT, there is a natural correspondence between the source sentence and the target sentence. However, this relationship has only been represented using the entire neural network and the training objective is computed in word-level. In this paper, we propose a sentence-level agreement module to directly minimize the difference between the representation of source and target sentence. The proposed agreement module can be integrated into NMT as an additional training objective function and can also be used to enhance the representation of the source sentences. Empirical results on the NIST Chinese-to-English and WMT English-to-German tasks show the proposed agreement module can significantly improve the NMT performance. Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Min Zhang 0005, Tiejun Zhao |
ACL (1) | 2 |
| 2019 | Recurrent Positional Embedding for Neural Machine TranslationabstractKehai Chen, Rui Wang, Masao Utiyama, Eiichiro Sumita. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Cross-Domain Transfer Learning for Dependency Parsing
Zuchao Li, Junru Zhou, Hai Zhao 0001, Rui Wang 0015 |
NLPCC (2) | 4 |
| 2019 | Neural Machine Translation With Sentence-Level Topic ContextabstractTraditional neural machine translation (NMT) methods use the word-level context to predict target language translation while neglecting the sentence-level context, which has been shown to be beneficial for translation prediction in statistical machine translation. This paper represents the sentence-level context as latent topic representations by using a convolution neural network, and designs a topic attention to integrate source sentence-level topic context information into both attention-based and Transformer-based NMT. In particular, our method can improve the performance of NMT by modeling source topics and translations jointly. Experiments on the large-scale LDC Chinese-to-English translation tasks and WMT'14 English-to-German translation tasks show that the proposed approach can achieve significant improvements compared with baseline systems. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Syntax-Directed Attention for Neural Machine TranslationabstractAttention mechanism, including global attention and local attention, plays a key role in neural machine translation (NMT). Global attention attends to all source words for word prediction. In comparison, local attention selectively looks at fixed-window source words. However, alignment weights for the current target word often decrease to the left and right by linear distance centering on the aligned source position and neglect syntax distance constraints. In this paper, we extend the local attention with syntax-distance constraint, which focuses on syntactically related source words with the predicted target word to learning a more effective context vector for predicting translation. Moreover, we further propose a double context NMT architecture, which consists of a global context vector and a syntax-directed context vector from the global attention, to provide more translation performance for NMT from source representation. The experiments on the large-scale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves a substantial and significant improvement over the baseline system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
AAAI | 2 |
| 2018 | A Survey of Domain Adaptation for Neural Machine TranslationabstractNeural machine translation (NMT) is a deep learning based approach for machine translation, which yields the state-of-the-art translation performance in scenarios where large-scale parallel corpora are available. Although the high-quality and domain-specific translation is crucial in the real world, domain-specific corpora are usually scarce or nonexistent, and thus vanilla NMT performs poorly in such scenarios. Domain adaptation that leverages both out-of-domain parallel corpora as well as monolingual corpora for in-domain translation, is very important for domain-specific translation. In this paper, we give a comprehensive survey of the state-of-the-art domain adaptation techniques for NMT. Chenhui Chu, Rui Wang 0015 |
COLING | 2 |
| 2018 | Exploring Recombination for Efficient Decoding of Neural Machine TranslationabstractIn Neural Machine Translation (NMT), the decoder can capture the features of the entire prediction history with neural connections and representations.This means that partial hypotheses with different prefixes will be regarded differently no matter how similar they are.However, this might be inefficient since some partial hypotheses can contain only local differences that will not influence future predictions.In this work, we introduce recombination in NMT decoding based on the concept of the "equivalence" of partial hypotheses.Heuristically, we use a simple n-gram suffix based equivalence function and adapt it into beam search decoding.Through experiments on large-scale Chinese-to-English and English-to-Germen translation tasks, we show that the proposed method can obtain similar translation quality with a smaller beam size, making NMT decoding more efficient. Zhisong Zhang, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001 |
EMNLP | 2 |
| 2018 | Graph-Based Bilingual Word Embedding for Statistical Machine TranslationabstractBilingual word embedding has been shown to be helpful for Statistical Machine Translation (SMT). However, most existing methods suffer from two obvious drawbacks. First, they only focus on simple contexts such as an entire document or a fixed-sized sliding window to build word embedding and ignore latent useful information from the selected context. Second, the word sense but not the word should be the minimal semantic unit; however, most existing methods still use word representation. To overcome these drawbacks, this article presents a novel Graph-Based Bilingual Word Embedding (GBWE) method that projects bilingual word senses into a multidimensional semantic space. First, a bilingual word co-occurrence graph is constructed using the co-occurrence and pointwise mutual information between the words. Then, maximum complete subgraphs (cliques), which play the role of a minimal unit for bilingual sense representation, are dynamically extracted according to the contextual information. Consequently, correspondence analysis, principal component analyses, and neural networks are used to summarize the clique-word matrix into lower dimensions to build the embedding model. Without contextual information, the proposed GBWE can be applied to lexical translation. In addition, given contextual information, GBWE is able to give a dynamic solution for bilingual word representations, which can be applied to phrase translation and generation. Empirical results show that GBWE can enhance the performance of lexical translation, as well as Chinese/French-to-English and Chinese-to-Japanese phrase-based SMT tasks (IWSLT, NTCIR, NIST, and WAT). Rui Wang 0015, Hai Zhao 0001, Sabine Ploux, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2018 | A Neural Approach to Source Dependence Based Context Model for Statistical Machine TranslationabstractIn statistical machine translation, translation prediction considers not only the aligned source word itself but also its source contextual information. Learning context representation is a promising method for improving translation results, particularly through neural networks. Most of the existing methods process context words sequentially and neglect source long-distance dependencies. In this paper, we propose a novel neural approach to source dependence-based context representation for translation prediction. The proposed model is capable of not only encoding source long-distance dependencies but also capturing functional similarities to better predict translations (i.e., word form translations and ambiguous word translations). To verify our method, the proposed mode is incorporated into phrase-based and hierarchical phrase-based translation models, respectively. Experiments on large-scale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves significant improvement over the baseline systems and outperforms several existing context-enhanced methods. Kehai Chen, Tiejun Zhao, Muyun Yang, Lemao Liu, Akihiro Tamura, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2018 | Sentence Selection and Weighting for Neural Machine Translation Domain AdaptationabstractNeural machine translation (NMT) has been prominent in many machine translation tasks. However, in some domain-specific tasks, only the corpora from similar domains can improve translation performance. If out-of-domain corpora are directly added into the in-domain corpus, the translation performance may even degrade. Therefore, domain adaptation techniques are essential to solve the NMT domain problem. Most existing methods for domain adaptation are designed for the conventional phrase-based machine translation. For NMT domain adaptation, there have been only a few studies on topics such as fine tuning, domain tags, and domain features. In this paper, we have four goals for sentence level NMT domain adaptation. First, the NMT's internal sentence embedding is exploited and the sentence embedding similarity is used to select out-of-domain sentences that are close to the in-domain corpus. Second, we propose three sentence weighting methods, i.e., sentence weighting, domain weighting, and batch weighting, to balance the data distribution during NMT training. Third, in addition, we propose dynamic training methods to adjust the sentence selection and weighting during NMT training. Fourth, to solve the multidomain problem in a real-world NMT scenario where the domain distributions of training and testing data often mismatch, we proposed a multidomain sentence weighting method to balance the domain distributions of training data and match the domain distributions of training and testing data. The proposed methods are evaluated in international workshop on spoken language translation (IWSLT) English-to-French/German tasks and a multidomain English-to-French task. Empirical results show that the sentence selection and weighting methods can significantly improve the NMT performance, outperforming the existing baselines. Rui Wang 0015, Masao Utiyama, Andrew M. Finch, Lemao Liu, Kehai Chen, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Neural Machine Translation with Source Dependency RepresentationabstractSource dependency information has been successfully introduced into statistical machine translation.However, there are only a few preliminary attempts for Neural Machine Translation (NMT), such as concatenating representations of source word and its dependency label together.In this paper, we propose a novel attentional NMT with source dependency representation to improve translation performance of NMT, especially on long sentences.Empirical results on NIST Chinese-to-English translation task show that our method achieves 1.6 BLEU improvements on average over a strong NMT system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Lemao Liu, Akihiro Tamura, Eiichiro Sumita, Tiejun Zhao |
EMNLP | 2 |
| 2017 | Instance Weighting for Neural Machine Translation Domain AdaptationabstractInstance weighting has been widely applied to phrase-based machine translation domain adaptation.However, it is challenging to be applied to Neural Machine Translation (NMT) directly, because NMT is not a linear model.In this paper, two instance weighting technologies, i.e., sentence weighting and domain weighting with a dynamic weight learning strategy, are proposed for NMT domain adaptation.Empirical results on the IWSLT English-German/French tasks show that the proposed methods can substantially improve NMT performance by up to 2.7-6.7 BLEU points, outperforming the existing baselines by up to 1.6-3.6BLEU points. Rui Wang 0015, Masao Utiyama, Lemao Liu, Kehai Chen, Eiichiro Sumita |
EMNLP | 1 |
| 2017 | Context-Aware Smoothing for Neural Machine TranslationabstractIn Neural Machine Translation (NMT), each word is represented as a low-dimension, real-value vector for encoding its syntax and semantic information. This means that even if the word is in a different sentence context, it is represented as the fixed vector to learn source representation. Moreover, a large number of Out-Of-Vocabulary (OOV) words, which have different syntax and semantic information, are represented as the same vector representation of “unk”. To alleviate this problem, we propose a novel context-aware smoothing method to dynamically learn a sentence-specific vector for each word (including OOV words) depending on its local context words in a sentence. The learned context-aware representation is integrated into the NMT to improve the translation performance. Empirical results on NIST Chinese-to-English translation task show that the proposed approach achieves 1.78 BLEU improvements on average over a strong attentional NMT, and outperforms some existing systems. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IJCNLP(1) | 2 |
| 2016 | Connecting Phrase based Statistical Machine Translation AdaptationabstractAlthough more additional corpora are now available for Statistical Machine Translation (SMT), only the ones which belong to the same or similar domains of the original corpus can indeed enhance SMT performance directly. A series of SMT adaptation methods have been proposed to select these similar-domain data, and most of them focus on sentence selection. In comparison, phrase is a smaller and more fine grained unit for data selection, therefore we propose a straightforward and efficient connecting phrase based adaptation method, which is applied to both bilingual phrase pair and monolingual n-gram adaptation. The proposed method is evaluated on IWSLT/NIST data sets, and the results show that phrase based SMT performances are significantly improved (up to +1.6 in comparison with phrase based SMT baseline system and +0.9 in comparison with existing methods). Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
COLING | 1 |
| 2016 | A Bilingual Graph-Based Semantic Model for Statistical Machine Translation
Rui Wang 0015, Hai Zhao 0001, Sabine Ploux, Bao-Liang Lu, Masao Utiyama |
IJCAI | 1 |
| 2016 | Converting Continuous-Space Language Models into N-gram Language Models with Efficient Bilingual Pruning for Statistical Machine TranslationabstractThe Language Model (LM) is an essential component of Statistical Machine Translation (SMT). In this article, we focus on developing efficient methods for LM construction. Our main contribution is that we propose a Natural N -grams based Converting (NNGC) method for transforming a Continuous-Space Language Model (CSLM) to a Back-off N -gram Language Model (BNLM). Furthermore, a Bilingual LM Pruning (BLMP) approach is developed for enhancing LMs in SMT decoding and speeding up CSLM converting. The proposed pruning and converting methods can convert a large LM efficiently by working jointly. That is, a LM can be effectively pruned before it is converted from CSLM without sacrificing performance, and further improved if an additional corpus contains out-of-domain information. For different SMT tasks, our experimental results indicate that the proposed NNGC and BLMP methods outperform the existing counterpart approaches significantly in BLEU and computational cost. Rui Wang 0015, Masao Utiyama, Isao Goto, Eiichiro Sumita, Hai Zhao 0001, Bao-Liang Lu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2015 | Neural Network Language Model for Chinese Pinyin Input Method Engine
Shenyuan Chen, Hai Zhao 0001, Rui Wang 0015 |
PACLIC | 3 |
| 2015 | A Machine Learning Method to Distinguish Machine Translation from Human Translation
Rui Wang 0015, Hai Zhao 0001 |
PACLIC | 2 |
| 2015 | English to Chinese Translation: How Chinese Character Matters
Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu |
PACLIC | 1 |
| 2015 | Bilingual Continuous-Space Language Model Growing for Statistical Machine TranslationabstractLarger n-gram language models (LMs) perform better in statistical machine translation (SMT). However, the existing approaches have two main drawbacks for constructing larger LMs: 1) it is not convenient to obtain larger corpora in the same domain as the bilingual parallel corpora in SMT; 2) most of the previous studies focus on monolingual information from the target corpora only, and redundant n-grams have not been fully utilized in SMT. Nowadays, continuous-space language model (CSLM), especially neural network language model (NNLM), has been shown great improvement in the estimation accuracies of the probabilities for predicting the target words. However, most of these CSLM and NNLM approaches still consider monolingual information only or require additional corpus. In this paper, we propose a novel neural network based bilingual LM growing method. Compared to the existing approaches, the proposed method enables us to use bilingual parallel corpus for LM growing in SMT. The results show that our new method outperforms the existing approaches on both SMT performance and computational efficiency significantly. Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Neural Network Based Bilingual Language Model Growing for Statistical Machine TranslationabstractSince larger n-gram Language Model (LM) usually performs better in Statistical Machine Translation (SMT), how to construct efficient large LM is an important topic in SMT.However, most of the existing LM growing methods need an extra monolingual corpus, where additional LM adaption technology is necessary.In this paper, we propose a novel neural network based bilingual LM growing method, only using the bilingual parallel corpus in SMT.The results show that our method can improve both the perplexity score for LM evaluation and BLEU score for SMT, and significantly outperforms the existing LM growing methods without extra corpus. Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
EMNLP | 1 |
| 2013 | Converting Continuous-Space Language Models into N-Gram Language Models for Statistical Machine TranslationabstractNeural network language models, or continuous-space language models (CSLMs), have been shown to improve the performance of statistical machine translation (SMT) when they are used for reranking n-best translations.However, CSLMs have not been used in the first pass decoding of SMT, because using CSLMs in decoding takes a lot of time.In contrast, we propose a method for converting CSLMs into back-off n-gram language models (BNLMs) so that we can use converted CSLMs in decoding.We show that they outperform the original BNLMs and are comparable with the traditional use of CSLMs in reranking. Rui Wang 0015, Masao Utiyama, Isao Goto, Eiichiro Sumita, Hai Zhao 0001, Bao-Liang Lu |
EMNLP | 1 |