Dongdong Zhang 0001

dblp:02/621-1 · DBLP profile ↗
← Back
56ranked-venue papers
3as first author
25since 2021 · last 2026
0000-0003-0833-7903ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 3 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 From Word to World: Can Large Language Models be Implicit Text-based World Models?
abstract
Yixia Li, Hongru Wang, Jiahao Qiu, Zhenfei Yin, Dongdong Zhang, Cheng Qian, Zeping Li, Xiaoteng Ma, Guanhua Chen, Heng Ji. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yixia Li, Hongru Wang 0003, Jiahao Qiu, Zhenfei Yin, Dongdong Zhang 0001, Cheng Qian 0008, Zeping Li, Xiaoteng Ma, Guanhua Chen 0001, Heng Ji 0001
ACL (1)5
2025 Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective
abstract
Yiyao Yu, Yuxiang Zhang, Dongdong Zhang, Xiao Liang, Hengyuan Zhang, Xingxing Zhang, Mahmoud Khademi, Hany Hassan Awadalla, Junjie Wang, Yujiu Yang, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yiyao Yu, Dongdong Zhang 0001, Xingxing Zhang 0002, Mahmoud Khademi, Hany Hassan, Junjie Wang 0011, Yujiu Yang 0001, Furu Wei
ACL (1)3
2025 ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive Framework
abstract
Hengyuan Zhang, Chenming Shang, Sizhe Wang, Dongdong Zhang, Yiyao Yu, Feng Yao, Renliang Sun, Yujiu Yang, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chenming Shang, Dongdong Zhang 0001, Yiyao Yu, Renliang Sun, Yujiu Yang 0001, Furu Wei
ACL (1)4
2024 Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models
abstract
Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, Ji-Rong Wen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wenyang Luo, Haoyang Huang, Dongdong Zhang 0001, Xiaolei Wang 0005, Wayne Xin Zhao, Furu Wei, Ji-Rong Wen
ACL (1)4
2024 Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
abstract
Tianyi Tang, Hongyuan Lu, Yuchen Jiang, Haoyang Huang, Dongdong Zhang, Xin Zhao, Tom Kocmi, Furu Wei. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hongyuan Lu, Haoyang Huang, Dongdong Zhang 0001, Wayne Xin Zhao, Tom Kocmi, Furu Wei
NAACL-HLT5
2024 DeepNet: Scaling Transformers to 1,000 Layers
abstract
In this paper, we propose a simple yet effective method to stabilize extremely deep Transformers. Specifically, we introduce a new normalization function (DeepNorm) to modify the residual connection in Transformer, accompanying with theoretically derived initialization. In-depth theoretical analysis shows that model updates can be bounded in a stable way. The proposed method combines the best of two worlds, i.e., good performance of Post-LN and stable training of Pre-LN, makingDeepNorma preferred alternative. We successfully scale Transformers up to 1,000 layers (i.e., 2,500 attention and feed-forward network sublayers) without difficulty, which is one order of magnitude deeper than previous deep Transformers. Extensive experiments demonstrate thatDeepNethas superior performance across various benchmarks, including machine translation, language modeling (i.e., BERT, GPT) and vision pre-training (i.e., BEiT). Remarkably, on a multilingual benchmark with 7,482 translation directions, our 200-layer model with 3.2B parameters significantly outperforms the 48-layer state-of-the-art model with 12B parameters by 5 BLEU points, which indicates a promising scaling direction. Our code is available athttps://aka.ms/torchscale.
Hongyu Wang 0009, Shuming Ma, Li Dong 0004, Shaohan Huang, Dongdong Zhang 0001, Furu Wei
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Discourse-Centric Evaluation of Document-level Machine Translation with a New Densely Annotated Parallel Corpus of Novels
abstract
Yuchen Eleanor Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Mrinmaya Sachan, Ryan Cotterell. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yuchen Eleanor Jiang, Tianyu Liu 0004, Shuming Ma, Dongdong Zhang 0001, Mrinmaya Sachan, Ryan Cotterell
ACL (1)4
2023 GanLM: Encoder-Decoder Pre-training with an Auxiliary Discriminator
abstract
Jian Yang, Shuming Ma, Li Dong, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang, Liqun Yang, Furu Wei, Zhoujun Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jian Yang 0030, Shuming Ma, Li Dong 0004, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang 0001, Liqun Yang, Furu Wei, Zhoujun Li 0001
ACL (1)7
2023 HanoiT: Enhancing Context-aware Translation via Selective Context
Jian Yang 0030, Yuwei Yin, Shuming Ma, Liqun Yang, Hongcheng Guo, Haoyang Huang, Dongdong Zhang 0001, Yutao Zeng, Zhoujun Li 0001, Furu Wei
DASFAA (3)7
2023 Are More Layers Beneficial to Graph Transformers?
Haiteng Zhao, Shuming Ma, Dongdong Zhang 0001, Zhi-Hong Deng 0001, Furu Wei
ICLR3
2023 On the Pareto Front of Multilingual Neural Machine Translation
abstract
In this work, we study how the performance of a given direction changes with its sampling ratio in Multilingual Neural Machine Translation (MNMT). By training over 200 multilingual models with various model sizes, data sizes, and language directions, we find it interesting that the performance of certain translation direction does not always improve with the increase of its weight in the multi-task optimization objective. Accordingly, scalarization method leads to a multitask trade-off front that deviates from the traditional Pareto front when there exists data imbalance in the training corpus, which poses a great challenge to improve the overall performance of all directions. Based on our observations, we propose the Double Power Law to predict the unique performance trade-off front in MNMT, which is robust across various languages, data adequacy, and the number of tasks. Finally, we formulate the sample ratio selection problem in MNMT as an optimization problem based on the Double Power Law. Extensive experiments show that it achieves better performance than temperature searching and gradient manipulation methods with only 1/5 to 1/2 of the total training budget. We release the code at https://github.com/pkunlp-icler/ParetoMNMT for reproduction.
Liang Chen 0024, Shuming Ma, Dongdong Zhang 0001, Furu Wei, Baobao Chang
NeurIPS3
2023 Syntax-Aware Data Augmentation for Neural Machine Translation
abstract
Data augmentation is an effective method for the performance enhancement of neural machine translation (NMT) by generating additional bilingual data. In this paper, we propose a novel data augmentation strategy for neural machine translation. Unlike existing data augmentation methods that simply modify words with the same probability across different sentences, we introduce a sentence-specific probability approach for word selection based on the syntactic roles of words in the sentence. Our motivation is to consider a linguistics-motivated method to obtain more ingenious language generation rather than relying on computation-motivated approaches only. We argue that high-quality aligned bilingual data is crucial for NMT, and only computation-motivated data augmentation is insufficient to provide good enough extra enhancement data. Our approach leverages dependency parse trees of input sentences to determine the selection probability of each word in the sentence using three different functions to calculate probabilities for words with different depths. Besides, our method also revises the probability for words considering the sentence length. We evaluate our methods on multiple translation tasks. The experimental results demonstrate that our proposed data augmentation method does effectively boost existing sentence-independent methods for significant improvement of performance on translation tasks. Furthermore, an ablation study shows that our method does select fewer essential words and preserves the syntactic structure.
Sufeng Duan, Hai Zhao 0001, Dongdong Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 GTrans: Grouping and Fusing Transformer Layers for Neural Machine Translation
abstract
Transformer structure, stacked by a sequence of encoder and decoder network layers, achieves significant development in neural machine translation. However, vanilla Transformer mainly exploits the top-layer representation, assuming the lower layers provide trivial or redundant information and thus ignoring the bottom-layer feature that is potentially valuable. In this work, we propose theGroup-Transformer model (GTrans) that flexibly divides multi-layer representations of both encoder and decoder into different groups and then fuses these group features to generate target words. To corroborate the effectiveness of the proposed method, extensive experiments and analytic experiments are conducted on three bilingual translation benchmarks and three multilingual translation tasks, including the IWLST-14, IWLST-17, LDC, WMT-14, WMT-21 and OPUS-100 benchmark. Experimental and analytical results demonstrate that our model outperforms its Transformer counterparts by a consistent gain. Furthermore, it can be successfully scaled up to 60 encoder layers and 36 decoder layers.
Jian Yang 0030, Yuwei Yin, Liqun Yang, Shuming Ma, Haoyang Huang, Dongdong Zhang 0001, Furu Wei, Zhoujun Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2022 Towards Making the Most of Cross-Lingual Transfer for Zero-Shot Neural Machine Translation
abstract
Guanhua Chen, Shuming Ma, Yun Chen, Dongdong Zhang, Jia Pan, Wenping Wang, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Guanhua Chen 0001, Shuming Ma, Yun Chen 0007, Dongdong Zhang 0001, Jia Pan 0001, Wenping Wang 0001, Furu Wei
ACL (1)4
2022 PAEG: Phrase-level Adversarial Example Generation for Neural Machine Translation
abstract
While end-to-end neural machine translation (NMT) has achieved impressive progress, noisy input usually leads models to become fragile and unstable. Generating adversarial examples as the augmented data has been proved to be useful to alleviate this problem. Existing methods for adversarial example generation (AEG) are word-level or character-level, which ignore the ubiquitous phrase structure. In this paper, we propose a Phrase-level Adversarial Example Generation (PAEG) framework to enhance the robustness of the translation model. Our method further improves the gradient-based word-level AEG method by adopting a phrase-level substitution strategy. We verify our method on three benchmarks, including LDC Chinese-English, IWSLT14 German-English, and WMT14 English-German tasks. Experimental results demonstrate that our approach significantly improves translation performance and robustness to noise compared to previous strong baselines.
Juncheng Wan, Jian Yang 0030, Shuming Ma, Dongdong Zhang 0001, Weinan Zhang 0001, Yong Yu 0001, Zhoujun Li 0001
COLING4
2022 LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine Translation
abstract
Multimodal Machine Translation (MMT) focuses on enhancing text-only translation with visual features, which has attracted considerable attention from both natural language processing and computer vision communities.Recent advances still struggle to train a separate model for each language pair, which is costly and unaffordable when the number of languages increases in the real world.In other words, the multilingual multimodal machine translation (Multilingual MMT) task has not been investigated, which aims to handle the aforementioned issues by providing a shared semantic space for multiple languages.Besides, the image modality has no language boundaries, which is superior to bridging the semantic gap between languages.To this end, we first propose the Multilingual MMT task by establishing two new Multilingual MMT benchmark datasets covering seven languages.Then, an effective baseline LVP-M 3 using visual prompts is proposed to support translations between different languages, which includes three stages (token encoding, language-aware visual prompt generation, and language translation).Extensive experimental results on our constructed benchmark datasets demonstrate the effectiveness of LVP-M 3 method for Multilingual MMT.* First two authors contributed equally.
Hongcheng Guo, Haoyang Huang, Jian Yang 0030, Zhoujun Li 0001, Dongdong Zhang 0001
EMNLP6
2022 Zero-shot Cross-lingual Transfer of Prompt-based Tuning with a Unified Multilingual Prompt
abstract
Prompt-based tuning has been proven effective for pretrained language models (PLMs).While most of the existing work focuses on the monolingual prompts, we study the multilingual prompts for multilingual PLMs, especially in the zero-shot cross-lingual setting.To alleviate the effort of designing different prompts for multiple languages, we propose a novel model that uses a unified prompt for all languages, called UniPrompt.Different from the discrete prompts and soft prompts, the unified prompt is model-based and languageagnostic.Specifically, the unified prompt is initialized by a multilingual PLM to produce language-independent representation, after which is fused with the text input.During inference, the prompts can be pre-computed so that no extra computation cost is needed.To collocate with the unified prompt, we propose a new initialization method for the target label word to further improve the model's transferability across languages.Extensive experiments show that our proposed methods can significantly outperform the strong baselines across different languages.We release data and code to facilitate future research 1 .
Lianzhe Huang, Shuming Ma, Dongdong Zhang 0001, Furu Wei, Houfeng Wang
EMNLP3
2022 High-resource Language-specific Training for Multilingual Neural Machine Translation
abstract
Multilingual neural machine translation (MNMT) trained in multiple language pairs has attracted considerable attention due to fewer model parameters and lower training costs by sharing knowledge among multiple languages. Nonetheless, multilingual training is plagued by language interference degeneration in shared parameters because of the negative interference among different translation directions, especially on high-resource languages. In this paper, we propose the multilingual translation model with the high-resource language-specific training (HLT-MT) to alleviate the negative interference, which adopts the two-stage training with the language-specific selection mechanism. Specifically, we first train the multilingual model only with the high-resource pairs and select the language-specific modules at the top of the decoder to enhance the translation quality of high-resource directions. Next, the model is further trained on all available corpora to transfer knowledge from high-resource languages (HRLs) to low-resource languages (LRLs). Experimental results show that HLT-MT outperforms various strong baselines on WMT-10 and OPUS-100 benchmarks. Furthermore, the analytic experiments validate the effectiveness of our method in mitigating the negative interference in multilingual training.
Jian Yang 0030, Yuwei Yin, Shuming Ma, Dongdong Zhang 0001, Zhoujun Li 0001, Furu Wei
IJCAI4
2022 UM4: Unified Multilingual Multiple Teacher-Student Model for Zero-Resource Neural Machine Translation
abstract
Most translation tasks among languages belong to the zero-resource translation problem where parallel corpora are unavailable. Multilingual neural machine translation (MNMT) enables one-pass translation using shared semantic space for all languages compared to the two-pass pivot translation but often underperforms the pivot-based method. In this paper, we propose a novel method, named as Unified Multilingual Multiple teacher-student Model for NMT (UM4). Our method unifies source-teacher, target-teacher, and pivot-teacher models to guide the student model for the zero-resource translation. The source teacher and target teacher force the student to learn the direct source-target translation by the distilled knowledge on both source and target sides. The monolingual corpus is further leveraged by the pivot-teacher model to enhance the student model. Experimental results demonstrate that our model of 72 directions significantly outperforms previous methods on the WMT benchmark.
Jian Yang 0030, Yuwei Yin, Shuming Ma, Dongdong Zhang 0001, Shuangzhi Wu, Hongcheng Guo, Zhoujun Li 0001, Furu Wei
IJCAI4
2022 BlonDe: An Automatic Evaluation Metric for Document-level Machine Translation
abstract
Yuchen Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Jian Yang, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, Ming Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Tianyu Liu 0004, Shuming Ma, Dongdong Zhang 0001, Jian Yang 0030, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, Ming Zhou 0001
NAACL-HLT4
2021 M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-Training
abstract
We present M3P, a Multitask Multilingual Multimodal Pre-trained model that combines multilingual pre-training and multimodal pre-training into a unified framework via multitask pre-training. Our goal is to learn universal representations that can map objects occurred in different modalities or texts expressed in different languages into a common semantic space. In addition, to explicitly encourage fine-grained alignment between images and non-English languages, we also propose Multimodal Code-switched Training (MCT) to combine monolingual pre-training and multimodal pre-training via a code-switch strategy. Experiments are performed on the multilingual image retrieval task across two benchmark datasets, including MSCOCO and Multi30K. M3P can achieve comparable results for English and new state-of-the-art results for non-English languages.
Minheng Ni, Haoyang Huang, Edward Dong Bo Cui, Taroon Bharti, Dongdong Zhang 0001, Nan Duan 0001
CVPR7
2021 Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained Encoders
abstract
Previous work mainly focuses on improving cross-lingual transfer for NLU tasks with a multilingual pretrained encoder (MPE), or improving the performance on supervised machine translation with BERT.However, it is under-explored that whether the MPE can help to facilitate the cross-lingual transferability of NMT model.In this paper, we focus on a zero-shot cross-lingual transfer task in NMT.In this task, the NMT model is trained with parallel dataset of only one language pair and an off-the-shelf MPE, then it is directly tested on zero-shot language pairs.We propose SixT, a simple yet effective model for this task.SixT leverages the MPE with a two-stage training schedule and gets further improvement with a position disentangled encoder and a capacity-enhanced decoder.Using this method, SixT significantly outperforms mBART, a pretrained multilingual encoderdecoder model explicitly designed for NMT, with an average improvement of 7.1 BLEU on zero-shot any-to-English test sets across 14 source languages.Furthermore, with much less training computation cost and training data, our model achieves better performance on 15 any-to-English test sets than CRISS and m2m-100, two strong multilingual NMT baselines.
Guanhua Chen 0001, Shuming Ma, Yun Chen 0007, Li Dong 0004, Dongdong Zhang 0001, Jia Pan 0001, Wenping Wang 0001, Furu Wei
EMNLP (1)5
2021 Smart-Start Decoding for Neural Machine Translation
abstract
Jian Yang, Shuming Ma, Dongdong Zhang, Juncheng Wan, Zhoujun Li, Ming Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Jian Yang 0030, Shuming Ma, Dongdong Zhang 0001, Juncheng Wan, Zhoujun Li 0001, Ming Zhou 0001
NAACL-HLT3
2021 XGPT: Cross-modal Generative Pre-Training for Image Captioning
Qiaolin Xia, Haoyang Huang, Nan Duan 0001, Dongdong Zhang 0001, Lei Ji 0001, Zhifang Sui, Edward Dong Bo Cui, Taroon Bharti, Ming Zhou 0001
NLPCC (1)4
2021 Learning to Select Relevant Knowledge for Neural Machine Translation
Jian Yang 0030, Juncheng Wan, Shuming Ma, Haoyang Huang, Dongdong Zhang 0001, Yong Yu 0001, Zhoujun Li 0001, Furu Wei
NLPCC (1)5
2020 Alternating Language Modeling for Cross-Lingual Pre-Training
abstract
Language model pre-training has achieved success in many natural language processing tasks. Existing methods for cross-lingual pre-training adopt Translation Language Model to predict masked words with the concatenation of the source sentence and its target equivalent. In this work, we introduce a novel cross-lingual pre-training method, called Alternating Language Modeling (ALM). It code-switches sentences of different languages rather than simple concatenation, hoping to capture the rich cross-lingual context of words and phrases. More specifically, we randomly substitute source phrases with target translations to create code-switched sentences. Then, we use these code-switched data to train ALM model to learn to predict words of different languages. We evaluate our pre-training ALM on the downstream tasks of machine translation and cross-lingual classification. Experiments show that ALM can outperform the previous pre-training methods on three benchmarks.1
Jian Yang 0030, Shuming Ma, Dongdong Zhang 0001, Shuangzhi Wu, Zhoujun Li 0001, Ming Zhou 0001
AAAI3
2020 A Simple and Effective Unified Encoder for Document-Level Machine Translation
abstract
Most of the existing models for documentlevel machine translation adopt dual-encoder structures.The representation of the source sentences and the document-level contexts 1 are modeled with two separate encoders.Although these models can make use of the document-level contexts, they do not fully model the interaction between the contexts and the source sentences, and can not directly adapt to the recent pre-training models (e.g., BERT) which encodes multiple sentences with a single encoder.In this work, we propose a simple and effective unified encoder that can outperform the baseline models of dualencoder models in terms of BLEU and ME-TEOR scores.Moreover, the pre-training models can further boost the performance of our proposed model.
Shuming Ma, Dongdong Zhang 0001, Ming Zhou 0001
ACL2
2020 Improving Neural Machine Translation with Soft Template Prediction
abstract
Although neural machine translation (NMT) has achieved significant progress in recent years, most previous NMT models only depend on the source text to generate translation.Inspired by the success of template-based and syntax-based approaches in other fields, we propose to use extracted templates from tree structures as soft target templates to guide the translation procedure.In order to learn the syntactic structure of the target sentences, we adopt the constituency-based parse tree to generate candidate templates.We incorporate the template information into the encoder-decoder framework to jointly utilize the templates and source text.Experiments show that our model significantly outperforms the baseline models on four benchmarks and demonstrate the effectiveness of soft target templates.
Jian Yang 0030, Shuming Ma, Dongdong Zhang 0001, Zhoujun Li 0001, Ming Zhou 0001
ACL3
2020 BiGCNN: Bidirectional Gated Convolutional Neural Network for Chinese Named Entity Recognition
Tianyang Zhao 0003, Haoyan Liu 0001, Qianhui Wu, Changzhi Sun, Dongdong Zhang 0001, Zhoujun Li 0001
DASFAA (1)5
2019 Effective Soft-Adaptation for Neural Machine Translation
Shuangzhi Wu, Dongdong Zhang 0001, Ming Zhou 0001
NLPCC (2)2
2019 Learning Unsupervised Word Mapping via Maximum Mean Discrepancy
Fuli Luo, Shuangzhi Wu, Jingjing Xu 0001, Dongdong Zhang 0001
NLPCC (1)5
2018 Improved Neural Machine Translation with Chinese Phonologic Features
Jian Yang 0030, Shuangzhi Wu, Dongdong Zhang 0001, Zhoujun Li 0001, Ming Zhou 0001
NLPCC (1)3
2018 Dependency-to-Dependency Neural Machine Translation
abstract
Recent research has proven that syntactic knowledge is effective to improve the performance of neural machine translation (NMT). Most previous work focuses on leveraging either source or target syntax in the recurrent neural network (RNN) based encoder–decoder model. In this paper, we simultaneously use both source and target dependency tree to improve the NMT model. First, we propose a simple but effective syntax-aware encoder to incorporate source dependency tree into NMT. The new encoder enriches each source state with dependence relations from the tree. Then, we propose a novel sequence-to-dependence framework. In this framework, the target translation and its corresponding dependence tree are jointly constructed and modeled. During decoding, the tree structure is used as context to facilitate word generations. Finally, we extend the sequence-to-dependence framework with the syntax-aware encoder to build a dependence-NMT model and apply the dependence-based framework to the Transformer. Experimental results on several translation tasks show that both source and target dependence structures can improve the translation quality and their effects can be accumulated.
Shuangzhi Wu, Dongdong Zhang 0001, Zhirui Zhang, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Sequence-to-Dependency Neural Machine Translation
abstract
Nowadays a typical Neural Machine Translation (NMT) model generates translations from left to right as a linear sequence, during which latent syntactic structures of the target sentences are not explicitly concerned.Inspired by the success of using syntactic knowledge of target language for improving statistical machine translation, in this paper we propose a novel Sequence-to-Dependency Neural Machine Translation (SD-NMT) method, in which the target word sequence and its corresponding dependency structure are jointly constructed and modeled, and this structure is used as context to facilitate word generations.Experimental results show that the proposed method significantly outperforms state-of-the-art baselines on Chinese-English and Japanese-English translation tasks.
Shuangzhi Wu, Dongdong Zhang 0001, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001
ACL (1)2
2017 Improved Neural Machine Translation with Source Syntax
abstract
Neural Machine Translation (NMT) based on the encoder-decoder architecture has recently achieved the state-of-the-art performance. Researchers have proven that extending word level attention to phrase level attention by incorporating source-side phrase structure can enhance the attention model and achieve promising improvement. However, word dependencies that can be crucial to correctly understand a source sentence are not always in a consecutive fashion (i.e. phrase structure), sometimes they can be in long distance. Phrase structures are not the best way to explicitly model long distance dependencies. In this paper we propose a simple but effective method to incorporate source-side long distance dependencies into NMT. Our method based on dependency trees enriches each source state with global dependency structures, which can better capture the inherent syntactic structure of source sentences. Experiments on Chinese-English and English-Japanese translation tasks show that our proposed method outperforms state-of-the-art SMT and NMT baselines.
Shuangzhi Wu, Ming Zhou 0001, Dongdong Zhang 0001
IJCAI3
2017 Modeling Indicative Context for Statistical Machine Translation
Shuangzhi Wu, Dongdong Zhang 0001, Shujie Liu 0001, Ming Zhou 0001
NLPCC2
2015 Efficient Disfluency Detection with Transition-based Parsing
abstract
Shuangzhi Wu, Dongdong Zhang, Ming Zhou, Tiejun Zhao. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Shuangzhi Wu, Dongdong Zhang 0001, Ming Zhou 0001, Tiejun Zhao
ACL (1)2
2014 Machine Translation with Real-Time Web Search
Lei Cui 0001, Ming Zhou 0001, Dongdong Zhang 0001, Mu Li 0001
AAAI4
2014 Learning Topic Representation for SMT with Neural Networks
abstract
Statistical Machine Translation (SMT) usually utilizes contextual information to disambiguate translation candidates. However, it is often limited to contexts within sentence boundaries, hence broader topical information cannot be leveraged. In this paper, we propose a novel approach to learning topic representation for paral-lel data using a neural network architec-ture, where abundant topical contexts are embedded via topic relevant monolingual data. By associating each translation rule with the topic representation, topic rele-vant rules are selected according to the dis-tributional similarity with the source text during SMT decoding. Experimental re-sults show that our method significantly improves translation accuracy in the NIST Chinese-to-English translation task com-pared to a state-of-the-art baseline. 1
Lei Cui 0001, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Muyun Yang
ACL (1)2
2014 A Lexicalized Reordering Model for Hierarchical Phrase-based Translation
Hailong Cao, Dongdong Zhang 0001, Mu Li 0001, Ming Zhou 0001, Tiejun Zhao
COLING2
2014 Soft Dependency Matching for Hierarchical Phrase-based Machine Translation
Hailong Cao, Dongdong Zhang 0001, Ming Zhou 0001, Tiejun Zhao
COLING2
2013 Punctuation Prediction with Transition-based Parsing
Dongdong Zhang 0001, Shuangzhi Wu, Nan Yang 0002, Mu Li 0001
ACL (1)1
2013 Multi-Domain Adaptation for SMT Using Multi-Task Learning
abstract
Domain adaptation for SMT usually adapts models to an individual specific domain.However, it often lacks some correlation among different domains where common knowledge could be shared to improve the overall translation quality.In this paper, we propose a novel multi-domain adaptation approach for SMT using Multi-Task Learning (MTL), with in-domain models tailored for each specific domain and a general-domain model shared by different domains.The parameters of these models are tuned jointly via MTL so that they can learn general knowledge more accurately and exploit domain knowledge better.Our experiments on a largescale English-to-Chinese translation task validate that the MTL-based adaptation approach significantly and consistently improves the translation quality compared to a non-adapted baseline.Furthermore, it also outperforms the individual adaptation of each specific domain.
Lei Cui 0001, Xilun Chen 0002, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001
EMNLP3
2013 Collective Corpus Weighting and Phrase Scoring for SMT Using Graph-Based Random Walk
Lei Cui 0001, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001
NLPCC2
2012 Hierarchical Chunk-to-String Translation
Yang Feng 0004, Dongdong Zhang 0001, Mu Li 0001, Qun Liu 0001
ACL (1)2
2012 A Ranking-based Approach to Word Reordering for Statistical Machine Translation
Nan Yang 0002, Mu Li 0001, Dongdong Zhang 0001, Nenghai Yu
ACL (1)3
2011 Function Word Generation in Statistical Machine Translation Systems
Lei Cui 0001, Dongdong Zhang 0001, Mu Li 0001, Ming Zhou 0001
MTSummit2
2010 Mixture Model-based Minimum Bayes Risk Decoding using Multiple Machine Translation Systems
Nan Duan 0001, Mu Li 0001, Dongdong Zhang 0001, Ming Zhou 0001
COLING3
2010 Adaptive Development Data Selection for Log-linear Model in Statistical Machine Translation
Mu Li 0001, Yinggong Zhao, Dongdong Zhang 0001, Ming Zhou 0001
COLING3
2009 Collaborative Decoding: Partial Hypothesis Re-ranking Using Translation Consensus between Decoders
Mu Li 0001, Nan Duan 0001, Dongdong Zhang 0001, Chi-Ho Li, Ming Zhou 0001
ACL/IJCNLP3
2009 Better Synchronous Binarization for Machine Translation
Tong Xiao 0001, Mu Li 0001, Dongdong Zhang 0001, Ming Zhou 0001
EMNLP3
2009 Introduction to China's CWMT2008 Machine Translation Evaluation
Hongmei Zhao, Qun Liu 0001, Yajuan Lü, Dongdong Zhang 0001, Mu Li 0001
MTSummit5
2008 Measure Word Generation for English-Chinese SMT Systems
Dongdong Zhang 0001, Mu Li 0001, Nan Duan 0001, Chi-Ho Li, Ming Zhou 0001
ACL1
2008 Diagnostic Evaluation of Machine Translation Systems Using Automatically Constructed Linguistic Check-Points
Ming Zhou 0001, Shujie Liu 0001, Mu Li 0001, Dongdong Zhang 0001, Tiejun Zhao
COLING5
2007 A Probabilistic Approach to Syntax-based Reordering for Statistical Machine Translation
Chi-Ho Li, Dongdong Zhang 0001, Mu Li 0001, Ming Zhou 0001, Yi Guan
ACL3
2007 Phrase Reordering Model Integrating Syntactic Knowledge for SMT
Dongdong Zhang 0001, Mu Li 0001, Chi-Ho Li, Ming Zhou 0001
EMNLP-CoNLL1