EDBT 2026 Demo / reviewers in the wild / expert
Xing Wang 0007
dblp:02/3674-7
· DBLP profile ↗
37ranked-venue papers
7as first author
16since 2021 · last 2025
0000-0002-0737-9653ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Competing Large Language Models in Multi-Agent Gaming EnvironmentsabstractDecision-making is a complex process requiring diverse abilities, making it an excellent framework for evaluating Large Language Models (LLMs). Researchers have examined LLMs' decision-making through the lens of Game Theory. However, existing evaluation mainly focus on two-player scenarios where an LLM competes against another. Additionally, previous benchmarks suffer from test set leakage due to their static design. We introduce GAMA($\gamma$)-Bench, a new framework for evaluating LLMs' Gaming Ability in Multi-Agent environments. It includes eight classical game theory scenarios and a dynamic scoring scheme specially designed to quantitatively assess LLMs' performance. $\gamma$-Bench allows flexible game settings and adapts the scoring system to different game parameters, enabling comprehensive evaluation of robustness, generalizability, and strategies for improvement. Our results indicate that GPT-3.5 demonstrates strong robustness but limited generalizability, which can be enhanced using methods like Chain-of-Thought. We also evaluate 13 LLMs from 6 model families, including GPT-3.5, GPT-4, Gemini, LLaMA-3.1, Mixtral, and Qwen-2. Gemini-1.5-Pro outperforms others, scoring of $69.8$ out of $100$, followed by LLaMA-3.1-70B ($65.9$) and Mixtral-8x22B ($62.4$). Our code and experimental results are publicly available at https://github.com/CUHK-ARISE/GAMABench. Jen-tse Huang 0001, Eric John Li, Man Ho Lam, Wenxuan Wang 0001, Youliang Yuan, Wenxiang Jiao, Xing Wang 0007, Zhaopeng Tu, Michael R. Lyu |
ICLR | 8 |
| 2025 | RaSA: Rank-Sharing Low-Rank AdaptationabstractLow-rank adaptation (LoRA) has been prominently employed for parameter-efficient fine-tuning of large language models (LLMs). However, the limited expressive capacity of LoRA, stemming from the low-rank constraint, has been recognized as a bottleneck, particularly in rigorous tasks like code generation and mathematical reasoning. To address this limitation, we introduce Rank-Sharing Low-Rank Adaptation (RaSA), an innovative extension that enhances the expressive capacity of LoRA by leveraging partial rank sharing across layers. By forming a shared rank pool and applying layer-specific weighting, RaSA effectively increases the number of ranks without augmenting parameter overhead. Our theoretically grounded and empirically validated approach demonstrates that RaSA not only maintains the core advantages of LoRA but also significantly boosts performance in challenging code and math tasks. Code, data and scripts are available at: https://github.com/zwhe99/RaSA. Zhiwei He 0002, Zhaopeng Tu, Xing Wang 0007, Wenxiang Jiao, Zhuosheng Zhang 0001, Rui Wang 0015 |
ICLR | 3 |
| 2025 | Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning CapabilityabstractMathematical reasoning tasks pose significant challenges for large language models (LLMs) because they require precise logical deduction and sequence analysis. In this work, we introduce the concept of critical tokens – elements within reasoning trajectories that significantly influence incorrect outcomes. We present a novel framework for identifying these tokens through rollout sampling and demonstrate their substantial divergence from traditional error tokens. Through extensive experiments on datasets such as GSM8K and MATH500, we show that identifying and replacing critical tokens significantly improves model accuracy. We propose an efficient methodology for pinpointing these tokens in large-scale datasets using contrastive estimation and extend this framework to enhance model training processes with direct preference optimization (DPO). Experimental results on GSM8K and MATH500 benchmarks with the widely used models Llama-3 (8B and 70B) and Deepseek-math (7B) demonstrate the effectiveness of the proposed approach, cDPO. Our results underscore the potential of leveraging critical tokens to reduce errors in reasoning tasks, advancing the development of AI systems capable of robust logical deduction. Zicheng Lin, Qiuzhi Liu, Xing Wang 0007, Ruilin Luo, Chufan Shi, Siheng Li, Yujiu Yang 0001, Zhaopeng Tu |
ICML | 5 |
| 2025 | DrugAssist: a large language model for molecule optimizationabstractRecently, the impressive performance of large language models (LLMs) on a wide range of tasks has attracted an increasing number of attempts to apply LLMs in drug discovery. However, molecule optimization, a critical task in the drug discovery pipeline, is currently an area that has seen little involvement from LLMs. Most of existing approaches focus solely on capturing the underlying patterns in chemical structures provided by the data, without taking advantage of expert feedback. These non-interactive approaches overlook the fact that the drug discovery process is actually one that requires the integration of expert experience and iterative refinement. To address this gap, we propose DrugAssist, an interactive molecule optimization model which performs optimization through human-machine dialogue by leveraging LLM's strong interactivity and generalizability. DrugAssist has achieved leading results in both single and multiple property optimization, simultaneously showcasing immense potential in transferability and iterative optimization. In addition, we publicly release a large instruction-based dataset called 'MolOpt-Instructions' for fine-tuning language models on molecule optimization tasks. We have made our code and data publicly available at https://github.com/blazerye/DrugAssist, which we hope to pave the way for future research in LLMs' application for drug discovery. Geyan Ye, Xibao Cai, Houtim Lai, Xing Wang 0007, Junhong Huang, Longyue Wang, Wei Liu 0005, Xiangxiang Zeng |
Briefings Bioinform. | 4 |
| 2024 | Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language ModelsabstractZhiwei He, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang, Zhaopeng Tu, Zhuosheng Zhang, Rui Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhiwei He 0002, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang 0007, Zhaopeng Tu, Zhuosheng Zhang 0001, Rui Wang 0015 |
ACL (1) | 5 |
| 2024 | Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebateabstractModern large language models (LLMs) like ChatGPT have shown remarkable performance on general language tasks but still struggle on complex reasoning tasks, which drives the research on cognitive behaviors of LLMs to explore human-like problem-solving strategies.Along this direction, one representative strategy is self-reflection, which asks an LLM to refine the solution with the feedback generated by itself iteratively.However, our study shows that such reflection-style methods suffer from the Degeneration-of-Thought (DoT) problem: once the LLM has established confidence in its solutions, it is unable to generate novel thoughts later through reflection even if its initial stance is incorrect.To address the DoT problem, we propose a Multi-Agent Debate (MAD) framework, in which multiple agents express their arguments in the state of "tit for tat" and a judge manages the debate process to obtain a final solution.Clearly, our MAD framework encourages divergent thinking in LLMs which would be helpful for tasks that require deep levels of contemplation.Experiment results on two challenging datasets, commonsense machine translation and counterintuitive arithmetic reasoning, demonstrate the effectiveness of our MAD framework.Extensive analyses suggest that the adaptive break of debate and the modest level of "tit for tat" state are required for MAD to obtain good performance.Moreover, we find that LLMs might not be a fair judge if different LLMs are used for agents.Code is available at https://github. com/Skytliang/Multi-Agents-Debate. Zhiwei He 0002, Wenxiang Jiao, Xing Wang 0007, Yan Wang 0060, Rui Wang 0015, Yujiu Yang 0001, Shuming Shi 0001, Zhaopeng Tu |
EMNLP | 4 |
| 2024 | Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward ModelabstractZhiwei He, Xing Wang, Wenxiang Jiao, Zhuosheng Zhang, Rui Wang, Shuming Shi, Zhaopeng Tu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhiwei He 0002, Xing Wang 0007, Wenxiang Jiao, Zhuosheng Zhang 0001, Rui Wang 0015, Shuming Shi 0001, Zhaopeng Tu |
NAACL-HLT | 2 |
| 2024 | Improving Gloss-free Sign Language Translation by Reducing Representation DensityabstractGloss-free sign language translation (SLT) aims to develop well-performing SLT systems with no requirement for the costly gloss annotations, but currently still lags behind gloss-based approaches significantly. In this paper, we identify **a representation density problem** that could be a bottleneck in restricting the performance of gloss-free SLT. Specifically, the representation density problem describes that the visual representations of semantically distinct sign gestures tend to be closely packed together in feature space, which makes gloss-free methods struggle with distinguishing different sign gestures and suffer from a sharp performance drop. To address the representation density problem, we introduce a simple but effective contrastive learning strategy, namely SignCL, which encourages gloss-free models to learn more discriminative feature representation in a self-supervised manner. Our experiments demonstrate that the proposed SignCL can significantly reduce the representation density and improve performance across various translation frameworks. Specifically, SignCLachieves a significant improvement in BLEU score for the Sign Language Transformer and GFSLT-VLP on the CSL-Daily dataset by 39\% and 46\%, respectively, without any increase of model parameters. Compared to Sign2GPT, a state-of-the-art method based on large-scale pre-trained vision and language models, SignCLachieves better performance with only 35\% of its parameters. We will release our code and model to facilitate further research. Jinhui Ye, Xing Wang 0007, Wenxiang Jiao, Junwei Liang 0001, Hui Xiong 0001 |
NeurIPS | 2 |
| 2024 | Exploring Human-Like Translation Strategy with Large Language ModelsabstractAbstract Large language models (LLMs) have demonstrated impressive capabilities in general scenarios, exhibiting a level of aptitude that approaches, in some aspects even surpasses, human-level intelligence. Among their numerous skills, the translation abilities of LLMs have received considerable attention. Compared to typical machine translation that focuses solely on source-to-target mapping, LLM-based translation can potentially mimic the human translation process, which might take preparatory steps to ensure high-quality translation. This work explores this possibility by proposing the MAPS framework, which stands for Multi-Aspect Prompting and Selection. Specifically, we enable LLMs first to analyze the given source sentence and induce three aspects of translation-related knowledge (keywords, topics, and relevant demonstrations) to guide the final translation process. Moreover, we employ a selection mechanism based on quality estimation to filter out noisy and unhelpful knowledge. Both automatic (3 LLMs × 11 directions × 2 automatic metrics) and human evaluation (preference study and MQM) demonstrate the effectiveness of MAPS. Further analysis shows that by mimicking the human translation process, MAPS reduces various translation errors such as hallucination, ambiguity, mistranslation, awkward style, untranslated text, and omission. Source code is available at https://github.com/zwhe99/MAPS-mt. Zhiwei He 0002, Wenxiang Jiao, Zhuosheng Zhang 0001, Yujiu Yang 0001, Rui Wang 0015, Zhaopeng Tu, Shuming Shi 0001, Xing Wang 0007 |
Trans. Assoc. Comput. Linguistics | 9 |
| 2023 | Scaling Back-Translation with Domain Text Generation for Sign Language Gloss TranslationabstractSign language gloss translation aims to translate the sign glosses into spoken language texts, which is challenging due to the scarcity of labeled gloss-text parallel data.Back-translation (BT), which generates pseudo parallel data by translating in-domain spoken language texts into sign glosses, has been applied to alleviate the data scarcity problem.However, the lack of large-scale high-quality in-domain spoken language text data limits the effect of BT.In this paper, to overcome the limitation, we propose a Prompt based domain text Generation (PGEN) approach to produce the large-scale in-domain spoken language text data.Specifically, PGEN randomly concatenates sentences from the original in-domain spoken language text data as prompts to induce a pre-trained language model (i.e., GPT-2) to generate spoken language texts in similar style.Experimental results on three benchmarks of sign language gloss translation in varied languages demonstrate that BT with spoken language texts generated by PGEN significantly outperforms the compared methods.In addition, as the scale of spoken language texts generated by PGEN increases, the BT technique can achieve further improvements, demonstrating the effectiveness of our approach.We release the code and data for facilitating future research in this field 1 . Jinhui Ye, Wenxiang Jiao, Xing Wang 0007, Zhaopeng Tu |
EACL | 3 |
| 2022 | Bridging the Data Gap between Training and Inference for Unsupervised Neural Machine TranslationabstractBack-translation is a critical component of Unsupervised Neural Machine Translation (UNMT), which generates pseudo parallel data from target monolingual data.A UNMT model is trained on the pseudo parallel data with translated source, and translates natural source sentences in inference.The source discrepancy between training and inference hinders the translation performance of UNMT models.By carefully designing experiments, we identify two representative characteristics of the data gap in source: (1) style gap (i.e., translated vs. natural text style) that leads to poor generalization capability; (2) content gap that induces the model to produce hallucination content biased towards the target language.To narrow the data gap, we propose an online self-training approach, which simultaneously uses the pseudo parallel data {natural source, translated target} to mimic the inference scenario.Experimental results on several widelyused language pairs show that our approach outperforms two strong baselines (XLM and MASS) by remedying the style and content gaps. 1 Model En-Fr En-De En-Ro Avg.⇒ ⇐ ⇒ ⇐ ⇒ ⇐ Full Test Set SNMT 38.4 33.6 29.5 33.9 33.7 32.5 33.6 XLM 37.4 34.5 27.2 34.3 34.6 32.7 33.5 MASS 37.8 34.9 27.1 35.2 35.1 33.4 33.9 Model En-Fr En-De En-Ro Avg.⇒ ⇐ ⇒ ⇐ ⇒ ⇐ Full Test Set SNMT 37.3 33.4 29.7 33.8 33.8 32.4 33.4 XLM 36.3 34.3 27.4 34.1 34.8 32.4 33.2 MASS 36.6 34.7 27.3 35.1 35.2 33.0 33.7 Zhiwei He 0002, Xing Wang 0007, Rui Wang 0015, Shuming Shi 0001, Zhaopeng Tu |
ACL (1) | 2 |
| 2022 | Understanding and Improving Sequence-to-Sequence Pretraining for Neural Machine TranslationabstractWenxuan Wang, Wenxiang Jiao, Yongchang Hao, Xing Wang, Shuming Shi, Zhaopeng Tu, Michael Lyu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Wenxuan Wang 0001, Wenxiang Jiao, Yongchang Hao, Xing Wang 0007, Shuming Shi 0001, Zhaopeng Tu, Michael R. Lyu |
ACL (1) | 4 |
| 2022 | Exploiting Inactive Examples for Natural Language Generation With Data RejuvenationabstractRecent years have witnessed the success of natural language generation (NLG) accomplished by deep neural networks, which require a large amount of training data for optimization. With the constant increase of data scale, the complex patterns and potential noises make training NLG models difficult. In order to fully utilize large-scale training data, we explore inactive examples in the training data and propose to rejuvenate the inactive examples for improving the performance of NLG models. Specifically, we define inactive examples as those sentence pairs that contribute less to the performance of NLG models, and show that their existence is independent of model variants but mainly determined by the data distribution. We further introducedata rejuvenationto improve the training of NLG models by re-labeling the inactive examples. The rejuvenated examples and active examples are combined to train a final NLG model. We evaluate our approach by experiments on machine translation (MT) and text summarization (TS) tasks, and achieve significant improvements of performance. Extensive analyses reveal that inactive examples are more difficult to learn than active ones and rejuvenation can reduce the learning difficulty, which stabilizes and accelerates the training process of NLG models and results in models with better generalization capability. Wenxiang Jiao, Xing Wang 0007, Shilin He, Zhaopeng Tu, Irwin King, Michael R. Lyu |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Self-Training Sampling with Monolingual Data Uncertainty for Neural Machine TranslationabstractWenxiang Jiao, Xing Wang, Zhaopeng Tu, Shuming Shi, Michael Lyu, Irwin King. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wenxiang Jiao, Xing Wang 0007, Zhaopeng Tu, Shuming Shi 0001, Michael R. Lyu, Irwin King |
ACL/IJCNLP (1) | 2 |
| 2021 | Multi-Task Learning with Shared Encoder for Non-Autoregressive Machine TranslationabstractYongchang Hao, Shilin He, Wenxiang Jiao, Zhaopeng Tu, Michael Lyu, Xing Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yongchang Hao, Shilin He, Wenxiang Jiao, Zhaopeng Tu, Michael R. Lyu, Xing Wang 0007 |
NAACL-HLT | 6 |
| 2021 | On the diversity of multi-head attention
Jian Li 0054, Xing Wang 0007, Zhaopeng Tu, Michael R. Lyu |
Neurocomputing | 2 |
| 2020 | Neuron Interaction Based Representation Composition for Neural Machine TranslationabstractRecent NLP studies reveal that substantial linguistic information can be attributed to single neurons, i.e., individual dimensions of the representation vectors. We hypothesize that modeling strong interactions among neurons helps to better capture complex information by composing the linguistic properties embedded in individual neurons. Starting from this intuition, we propose a novel approach to compose representations learned by different components in neural machine translation (e.g., multi-layer networks or multi-head attention), based on modeling strong interactions among neurons in the representation vectors. Specifically, we leverage bilinear pooling to model pairwise multiplicative interactions among individual neurons, and a low-rank approximation to make the model computationally feasible. We further propose extended bilinear pooling to incorporate first-order representations. Experiments on WMT14 English⇒German and English⇒French translation tasks show that our model consistently improves performances over the SOTA Transformer baseline. Further analyses demonstrate that our approach indeed captures more syntactic and semantic information as expected. Jian Li 0054, Xing Wang 0007, Baosong Yang, Shuming Shi 0001, Michael R. Lyu, Zhaopeng Tu |
AAAI | 2 |
| 2020 | How Does Selective Mechanism Improve Self-Attention Networks?abstractSelf-attention networks (SANs) with selective mechanism has produced substantial improvements in various NLP tasks by concentrating on a subset of input words.However, the underlying reasons for their strong performance have not been well explained.In this paper, we bridge the gap by assessing the strengths of selective SANs (SSANs), which are implemented with a flexible and universal Gumbel-Softmax.Experimental results on several representative NLP tasks, including natural language inference, semantic role labelling, and machine translation, show that SSANs consistently outperform the standard SANs.Through well-designed probing experiments, we empirically validate that the improvement of SSANs can be attributed in part to mitigating two commonly-cited weaknesses of SANs: word order encoding and structure modeling.Specifically, the selective mechanism improves SANs by paying more attention to content words that contribute to the meaning of the sentence.The code and data are released at https://github.com/xwgeng/SSAN. Xinwei Geng, Longyue Wang, Xing Wang 0007, Bing Qin 0001, Ting Liu 0001, Zhaopeng Tu |
ACL | 3 |
| 2020 | Data Rejuvenation: Exploiting Inactive Training Examples for Neural Machine TranslationabstractLarge-scale training datasets lie at the core of the recent success of neural machine translation (NMT) models.However, the complex patterns and potential noises in the large-scale data make training NMT models difficult.In this work, we explore to identify the inactive training examples which contribute less to the model performance, and show that the existence of inactive examples depends on the data distribution.We further introduce data rejuvenation to improve the training of NMT models on large-scale datasets by exploiting inactive examples.The proposed framework consists of three phases.First, we train an identification model on the original training data, and use it to distinguish inactive examples and active examples by their sentence-level output probabilities.Then, we train a rejuvenation model on the active examples, which is used to re-label the inactive examples with forwardtranslation.Finally, the rejuvenated examples and the active examples are combined to train the final NMT model.Experimental results on WMT14 English-German and English-French datasets show that the proposed data rejuvenation consistently and significantly improves performance for several strong NMT models.Extensive analyses reveal that our approach stabilizes and accelerates the training process of NMT models, resulting in final models with better generalization capability. 1 Wenxiang Jiao, Xing Wang 0007, Shilin He, Irwin King, Michael R. Lyu, Zhaopeng Tu |
EMNLP (1) | 2 |
| 2020 | Incorporating Phrase-Level Agreement into Neural Machine Translation
Xing Wang 0007, Min Zhang 0005, Tiejun Zhao |
NLPCC (1) | 2 |
| 2020 | Exploiting deep representations for natural language processing
Zi-Yi Dou, Xing Wang 0007, Shuming Shi 0001, Zhaopeng Tu |
Neurocomputing | 2 |
| 2020 | A Novel Sentence-Level Agreement Architecture for Neural Machine TranslationabstractIn neural machine translation (NMT), there is a natural correspondence between source and target sentences. The traditional NMT method does not explicitly model the translation agreement on sentence-level. In this article, we propose a comprehensive and novel sentence-level agreement architecture to alleviate this problem. It directly minimizes the difference between the representations of the source-side and target-side sentence on sentence-level. First, we compare a variety of sentence representation strategies and propose a “Gated Sum” sentence representation to achieve better sentence semantic information. Then, rather than a single-layer sentence-level agreement architecture, we further propose a multi-layer sentence agreement architecture to make the source and target semantic spaces closer layer by layer. The proposed agreement module can be integrated into NMT as an additional training objective function, and can also be used to enhance the representation of the source-side sentences. Experiments on the NIST Chinese-to-English and the WMT English-to-German translation tasks show that the proposed agreement architecture achieves significant improvements over state-of-the-art baselines, demonstrating the effectiveness and necessity of exploiting sentence-level agreement for NMT. Rui Wang 0015, Kehai Chen, Xing Wang 0007, Tiejun Zhao, Min Zhang 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Dynamic Layer Aggregation for Neural Machine Translation with Routing-by-AgreementabstractWith the promising progress of deep neural networks, layer aggregation has been used to fuse information across layers in various fields, such as computer vision and machine translation. However, most of the previous methods combine layers in a static fashion in that their aggregation strategy is independent of specific hidden states. Inspired by recent progress on capsule networks, in this paper we propose to use routing-by-agreement strategies to aggregate layers dynamically. Specifically, the algorithm learns the probability of a part (individual layer representations) assigned to a whole (aggregated representations) in an iterative way and combines parts accordingly. We implement our algorithm on top of the state-of-the-art neural machine translation model TRANSFORMER and conduct experiments on the widely-used WMT14 sh⇒German and WMT17 Chinese⇒English translation datasets. Experimental results across language pairs show that the proposed approach consistently outperforms the strong baseline model and a representative static aggregation model. Zi-Yi Dou, Zhaopeng Tu, Xing Wang 0007, Longyue Wang, Shuming Shi 0001, Tong Zhang 0001 |
AAAI | 3 |
| 2019 | Context-Aware Self-Attention NetworksabstractSelf-attention model has shown its flexibility in parallel computation and the effectiveness on modeling both long- and short-term dependencies. However, it calculates the dependencies between representations without considering the contextual information, which has proven useful for modeling dependencies among neural representations in various natural language tasks. In this work, we focus on improving self-attention networks through capturing the richness of context. To maintain the simplicity and flexibility of the self-attention networks, we propose to contextualize the transformations of the query and key layers, which are used to calculate the relevance between elements. Specifically, we leverage the internal representations that embed both global and deep contexts, thus avoid relying on external resources. Experimental results on WMT14 English⇒German and WMT17 Chinese⇒English translation tasks demonstrate the effectiveness and universality of the proposed methods. Furthermore, we conducted extensive analyses to quantify how the context vectors participate in the self-attention model. Baosong Yang, Jian Li 0054, Derek F. Wong, Lidia S. Chao, Xing Wang 0007, Zhaopeng Tu |
AAAI | 5 |
| 2019 | Exploiting Sentential Context for Neural Machine TranslationabstractIn this work, we present novel approaches to exploit sentential context for neural machine translation (NMT).Specifically, we first show that a shallow sentential context extracted from the top encoder layer only, can improve translation performance via contextualizing the encoding representations of individual words.Next, we introduce a deep sentential context, which aggregates the sentential context representations from all the internal layers of the encoder to form a more comprehensive context representation.Experimental results on the WMT14 English⇒German and English⇒French benchmarks show that our model consistently improves performance over the strong TRANSFORMER model (Vaswani et al., 2017), demonstrating the necessity and effectiveness of exploiting sentential context for NMT. Xing Wang 0007, Zhaopeng Tu, Longyue Wang, Shuming Shi 0001 |
ACL (1) | 1 |
| 2019 | Multi-Granularity Self-Attention for Neural Machine TranslationabstractJie Hao, Xing Wang, Shuming Shi, Jinfeng Zhang, Zhaopeng Tu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xing Wang 0007, Shuming Shi 0001, Zhaopeng Tu |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Towards Better Modeling Hierarchical Structure for Self-Attention with Ordered NeuronsabstractJie Hao, Xing Wang, Shuming Shi, Jinfeng Zhang, Zhaopeng Tu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xing Wang 0007, Shuming Shi 0001, Zhaopeng Tu |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Towards Understanding Neural Machine Translation with Word ImportanceabstractShilin He, Zhaopeng Tu, Xing Wang, Longyue Wang, Michael Lyu, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shilin He, Zhaopeng Tu, Xing Wang 0007, Longyue Wang, Michael R. Lyu, Shuming Shi 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | One Model to Learn Both: Zero Pronoun Prediction and TranslationabstractLongyue Wang, Zhaopeng Tu, Xing Wang, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Longyue Wang, Zhaopeng Tu, Xing Wang 0007, Shuming Shi 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Self-Attention with Structural Position RepresentationsabstractXing Wang, Zhaopeng Tu, Longyue Wang, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xing Wang 0007, Zhaopeng Tu, Longyue Wang, Shuming Shi 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Exploiting Deep Representations for Neural Machine TranslationabstractAdvanced neural machine translation (NMT) models generally implement encoder and decoder as multiple layers, which allows systems to model complex functions and capture complicated linguistic structures.However, only the top layers of encoder and decoder are leveraged in the subsequent process, which misses the opportunity to exploit the useful information embedded in other layers.In this work, we propose to simultaneously expose all of these signals with layer aggregation and multi-layer attention mechanisms.In addition, we introduce an auxiliary regularization term to encourage different layers to capture diverse information.Experimental results on widely-used WMT14 English⇒German and WMT17 Chinese⇒English translation data demonstrate the effectiveness and universality of the proposed approach. Zi-Yi Dou, Zhaopeng Tu, Xing Wang 0007, Shuming Shi 0001, Tong Zhang 0001 |
EMNLP | 3 |
| 2018 | Incorporating Statistical Machine Translation Word Knowledge Into Neural Machine TranslationabstractNeural machine translation (NMT) has gained more and more attention in recent years, mainly due to its simplicity yet state-of-the-art performance. However, previous research has shown that NMT suffers from several limitations: source coverage guidance, translation of rare words, and the limited vocabulary, while statistical machine translation (SMT) has complementary properties that correspond well to these limitations. It is straightforward to improve the translation performance by combining the advantages of two kinds of models. This paper proposes a general framework for incorporating the SMT word knowledge into NMT to alleviate above word-level limitations. In our framework, the NMT decoder makes more accurate word prediction by referring to the SMT word recommendations in both training and testing phases. Specifically, the SMT model offers informative word recommendations based on the NMT decoding information. Then, we use the SMT word predictions as prior knowledge to adjust the NMT word generation probability, which unitizes a neural network based classifier to digest the discrete word knowledge. In this paper, we use two model variants to implement the framework, one with a gating mechanism and the other with a direct competition mechanism. Experimental results on Chinese-to-English and English-to-German translation tasks show that the proposed framework can take advantage of the SMT word knowledge and consistently achieve significant improvements over NMT and SMT baseline systems. Xing Wang 0007, Zhaopeng Tu, Min Zhang 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Neural Machine Translation Advised by Statistical Machine TranslationabstractNeural Machine Translation (NMT) is a new approach to machine translation that has made great progress in recent years. However, recent studies show that NMT generally produces fluent but inadequate translations (Tu et al. 2016b; 2016a; He et al. 2016; Tu et al. 2017). This is in contrast to conventional Statistical Machine Translation (SMT), which usually yields adequate but non-fluent translations. It is natural, therefore, to leverage the advantages of both models for better translations, and in this work we propose to incorporate SMT model into NMT framework. More specifically, at each decoding step, SMT offers additional recommendations of generated words based on the decoding information from NMT (e.g., the generated partial translation and attention history). Then we employ an auxiliary classifier to score the SMT recommendations and a gating function to combine the SMT recommendations with NMT generations, both of which are jointly trained within the NMT architecture in an end-to-end manner. Experimental results on Chinese-English translation show that the proposed approach achieves significant and consistent improvements over state-of-the-art NMT and SMT systems on multiple NIST test sets. Xing Wang 0007, Zhengdong Lu, Zhaopeng Tu, Hang Li 0001, Deyi Xiong, Min Zhang 0005 |
AAAI | 1 |
| 2017 | Translating Phrases in Neural Machine TranslationabstractPhrases play an important role in natural language understanding and machine translation (Sag et al., 2002;Villavicencio et al., 2005).However, it is difficult to integrate them into current neural machine translation (NMT) which reads and generates sentences word by word.In this work, we propose a method to translate phrases in NMT by integrating a phrase memory storing target phrases from a phrase-based statistical machine translation (SMT) system into the encoder-decoder architecture of NMT.At each decoding step, the phrase memory is first re-written by the SMT model, which dynamically generates relevant target phrases with contextual information provided by the NMT model.Then the proposed model reads the phrase memory to make probability estimations for all phrases in the phrase memory.If phrase generation is carried on, the NMT decoder selects an appropriate phrase from the memory to perform phrase translation and updates its decoding state by consuming the words in the selected phrase.Otherwise, the NMT decoder generates a word from the vocabulary as the general NMT decoder does.Experiment results on the Chinese→English translation show that the proposed model achieves significant improvements over the baseline on various test sets. Xing Wang 0007, Zhaopeng Tu, Deyi Xiong, Min Zhang 0005 |
EMNLP | 1 |
| 2015 | Learning Semantic Representations for Nonterminals in Hierarchical Phrase-Based TranslationabstractIn hierarchical phrase-based translation, coarse-grained nonterminal Xs may generate inappropriate translations due to the lack of sufficient information for phrasal substitution.In this paper we propose a framework to refine nonterminals in hierarchical translation rules with real-valued semantic representations.The semantic representations are learned via a weighted mean value and a minimum distance method using phrase vector representations obtained from large scale monolingual corpus.Based on the learned semantic vectors, we build a semantic nonterminal refinement model to measure semantic similarities between phrasal substitutions and nonterminal Xs in translation rules.Experiment results on Chinese-English translation show that the proposed model significantly improves translation quality on NIST test sets. Xing Wang 0007, Deyi Xiong, Min Zhang 0005 |
EMNLP | 1 |
| 2015 | Topic-Based Coherence Modeling for Statistical Machine TranslationabstractCoherence that ties sentences of a text into a meaningfully connected structure is of great importance to text generation and translation. In this paper, we propose topic-based coherence models to produce coherence for document translation, in terms of the continuity of sentence topics in a text. We automatically extract a coherence chain for each source text to be translated. Based on the extracted source coherence chain, we adopt a maximum entropy classifier to predict the target coherence chain that defines a linear topic structure for the target document. We build two topic-based coherence models on the predicted target coherence chain: 1) a word level coherence model that helps the decoder select coherent word translations and 2) a phrase level coherence model that guides the decoder to select coherent phrase translations. We integrate the two models into a state-of-the-art phrase-based machine translation system. Experiments on large-scale training data show that our coherence models achieve substantial improvements over both the baseline and models that are built on either document topics or sentence topics obtained under the assumption of direct topic correspondence between the source and target side. Additionally, further evaluations on translation outputs suggest that target translations generated by our coherence models are more coherent and similar to reference translations than those generated by the baseline. Deyi Xiong, Min Zhang 0005, Xing Wang 0007 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | A Topic-Based Reordering Model for Statistical Machine Translation
Xing Wang 0007, Deyi Xiong, Min Zhang 0005, Yu Hong 0001, Jianmin Yao 0001 |
NLPCC | 1 |