Tong Xiao 0001

dblp:05/5091-1 · DBLP profile ↗
← Back
95ranked-venue papers
11as first author
61since 2021 · last 2026
0000-0002-5842-6501ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 87 · 11 first-author · 54 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 3 first-author · 20 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 SageLM: A Multi-aspect and Explainable Large Language Model for Speech Judgement
abstract
Speech-to-Speech (S2S) Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling end-to-end spoken dialogue systems. However, evaluating these models remains a fundamental challenge. We propose SageLM, an end-to-end, multi-aspect, and explainable speech LLM for comprehensive S2S LLMs evaluation. First, unlike cascaded approaches that disregard acoustic features, SageLM jointly assesses both semantic and acoustic dimensions. Second, it leverages rationale-based supervision to enhance explainability and guide model learning, achieving superior alignment with evaluation outcomes compared to rule-based reinforcement learning methods. Third, we introduce SpeechFeedback, a synthetic preference dataset, and employ a two-stage training paradigm to mitigate the scarcity of speech preference data. Trained on both semantic and acoustic dimensions, SageLM achieves an 82.79% agreement rate with human evaluators, outperforming cascaded and SLM-based baselines by at least 7.42% and 26.20%, respectively.
Yuan Ge 0001, Junxiang Zhang, Xiangnan Ma, Chenglong Wang 0002, Kaiyang Ye, Yangfan Du, Linfeng Zhang 0001, Yuxin Huang 0004, Tong Xiao 0001, Zhengtao Yu 0001
AAAI11
2026 WaveEx: Accelerating Flow Matching-based Speech Generation via Wavelet-guided Extrapolation
abstract
Flow matching-based generative models offer a principled approach to modeling continuous-time dynamics in speech generation. However, inference is often computationally expensive due to repeated neural network evaluations required by ODE solvers. We propose WaveEx, a training-free and plug-in acceleration framework which replaces portions of ODE integration with wavelet-guided extrapolation. By leveraging the multi-scale structure of latent trajectories, WaveEx predicts future states directly in the frequency domain without additional model evaluations or architectural changes. WaveEx consistently accelerates inference across diverse speech generation tasks. The gains are especially pronounced in tasks like speech synthesis (up to 5.73× speedup) and music generation (2.75×), where flow matching plays a central role in alignment modeling and dense ODE integration. Even in tasks with simpler input-output mappings such as speech enhancement (4.55×) and voice conversion (2.75×), WaveEx still achieves notable acceleration, demonstrating the robustness and generalizability of the approach. These results highlight wavelet-guided extrapolation as a lightweight and broadly applicable alternative to full ODE solving for flow matching-based speech generation.
Xiyan Gui, Zhengkun Ge, Yuan Ge 0001, Chang Zou, Zhikang Niu, Qixi Zheng, Chen Xu 0008, Xie Chen 0001, Tong Xiao 0001, Linfeng Zhang 0001
AAAI11
2026 Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
abstract
Previous methods evaluate reward models by testing them on a fixed pairwise ranking test set, but they typically do not provide performance information on each preference dimension. In this work, we address the evaluation challenge of reward models by probing preference representations. To confirm the effectiveness of this evaluation method, we construct a Multi-dimensional Reward Model Benchmark (MRMBench), a collection of six probing tasks for different preference dimensions. We design it to favor and encourage reward models that better capture preferences across different dimensions. Furthermore, we introduce an analysis method, inference-time probing, which identifies the dimensions used during the reward prediction and enhances its interpretability. Through extensive experiments, we find that MRMBench strongly correlates with LLM alignment performance, supporting it as a reliable reference for developing advanced reward models. By analyzing the evaluation results on MRMBench, we reveal that reward models struggle to simultaneously capture preferences across multiple dimensions, highlighting the potential of multi-objective optimization in reward modeling. Furthermore, our results demonstrate that the proposed inference-time probing method provides a reliable metric for assessing the confidence of reward predictions, leading to improved alignment of large language models.
Chenglong Wang 0002, Yifu Huo, Yang Gan, Yongyu Mu, Qiaozhi He, Murun Yang, Chunliang Zhang, Tongran Liu, Anxiang Ma, Zhengtao Yu 0001, Tong Xiao 0001
AAAI13
2026 GRAM-R²: Self-Training Generative Foundation Reward Models for Reward Reasoning
abstract
Major progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs to generalist reward models. Despite this trend, developing effective reward models remains a fundamental challenge: the heavy reliance on large-scale labeled preference data. Pre-training on abundant unlabeled data offers a promising direction, but existing approaches fall short in instilling explicit reasoning capabilities into reward models. To bridge this gap, we propose a self-training approach that can leverage unlabeled data to scale up reward reasoning in reward models. Based on this approach, we develop GRAM-R² a generative reward model trained to produce not only preference labels but also accompanying reward rationales. GRAM-R² can serve as a foundation model for reward reasoning and can be applied to a wide range of tasks with minimal or no additional fine-tuning. It can support downstream applications such as policy optimization and task-specific reward tuning. Experiments on response ranking, task adaptation, and reinforcement learning from human feedback demonstrate that GRAM-R² consistently delivers strong performance, outperforming several strong discriminative and generative baselines.
Chenglong Wang 0002, Yongyu Mu, Yifu Huo, Jiali Zeng, Murun Yang, Xiaoyang Hao, Chunliang Zhang, Fandong Meng, Tong Xiao 0001
AAAI13
2026 LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance
abstract
Yuchun Fan, Bei Li, Peiguang Li, Yilin Wang, Yongyu Mu, Jian Yang, Xin Chen, Rongxiang Weng, Jingang Wang, Xunliang Cai, JingBo Zhu, Tong Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuchun Fan, Peiguang Li, Yongyu Mu, Rongxiang Weng, Jingang Wang, Tong Xiao 0001
ACL (1)12
2026 On the Emotion Understanding of Synthesized Speech
abstract
Yuan Ge, Haishu Zhao, AoKai Hao, Junxiang Zhang, Bei Li, Xiaoqian Liu, Chenglong Wang, Jianjin Wang, Bingsen Zhou, Bingyu Liu, JingBo Zhu, Zhengtao Yu, Tong Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuan Ge 0001, Haishu Zhao, Aokai Hao, Junxiang Zhang, Chenglong Wang 0002, Jianjin Wang, Bingsen Zhou, Zhengtao Yu 0001, Tong Xiao 0001
ACL (1)13
2026 Empirical Analysis of Decoding Biases in Masked Diffusion Models
abstract
Pengcheng Huang, Tianming Liu, Zhenghao Liu, Yukun Yan, Shuo Wang, Tong Xiao, Zulong Chen, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pengcheng Huang 0004, Zhenghao Liu 0001, Yukun Yan, Shuo Wang 0013, Tong Xiao 0001, Zulong Chen, Maosong Sun 0001
ACL (1)6
2026 NiuTrans.LMT: Toward Inclusive and Scalable Multilingual Machine Translation with LLMs
abstract
Yingfeng Luo, Ziqiang Xu, Yuxuan Ouyang, MuRun Yang, DingYang Lin, Kaiyan Chang, Tong Zheng, Bei Li, Peinan Feng, Quan Du, Tong Xiao, JingBo Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yingfeng Luo, Yuxuan Ouyang, Murun Yang, Dingyang Lin, Kaiyan Chang 0001, Peinan Feng, Quan Du, Tong Xiao 0001
ACL (1)11
2026 MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
abstract
Junhao Ruan, Abudukeyumu Abudula, Bei Li, Yongjing Yin, Xinyu Liu, Kechen Jiao, Xin Chen, Jingang Wang, Xunliang Cai, Tong Xiao, JingBo Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junhao Ruan, Abudukeyumu Abudula, Yongjing Yin, Kechen Jiao, Jingang Wang, Tong Xiao 0001
ACL (1)10
2026 CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling
abstract
Runsong Zhao, Shilei Liu, Jiwei Tang, Langming Liu, Haibin Chen, Weidong Zhang, Yujin Yuan, Tong Xiao, JingBo Zhu, Wenbo Su, Bo Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Runsong Zhao, Shilei Liu, Jiwei Tang, Langming Liu, Yujin Yuan, Tong Xiao 0001, Wenbo Su, Bo Zheng 0007
ACL (1)8
2026 Cross-layer Attention Sharing for Pre-trained Large Language Models
abstract
Abstract To enhance the efficiency of the attention mechanism within large language models (LLMs), previous works primarily compress the Key-Value cache or group attention heads, while largely overlooking redundancy between layers. Our comprehensive analyses across various LLMs show that highly similar attention patterns persist within most layers. It’s intuitive to reduce the redundancy by sharing attention weights across layers. However, further analysis reveals two challenges: (1) Directly sharing the weight matrix without carefully rearranging the attention heads proves to be ineffective; (2) Shallow layers are vulnerable to small deviations in attention weights. Driven by these insights, we introduce LiSA, a lightweight substitute for self-attention in well-trained LLMs. LiSA employs tiny feed-forward networks to align attention heads between adjacent layers and low-rank matrices to approximate differences in layer-wise attention weights. Evaluations encompassing 13 typical benchmarks demonstrate that LiSA maintains high response quality in terms of accuracy and perplexity while reducing redundant attention calculations within 53% −84% of the total layers. Our implementations of LiSA achieve a 6 × compression of Q and K matrices within the attention mechanism, with maximum throughput improvements 19.5%, 32.3%, and 40.1% for LLaMA3-8B, LLaMA2-7B, and LLaMA2-13B, respectively. Our code is available at https://github.com/takagi97/lisa.
Yongyu Mu, Yuzhang Wu, Yuchun Fan, Chenglong Wang 0002, Jiali Zeng, Qiaozhi He, Murun Yang, Fandong Meng, Jie Zhou 0016, Tong Xiao 0001
Trans. Assoc. Comput. Linguistics11
2025 RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
abstract
Large vision-language models (LVLMs) often fail to align with human preferences, leading to issues like generating misleading content without proper visual context (also known as hallucination). A promising solution to this problem is using human-preference alignment techniques, such as best-of-n sampling and reinforcement learning. However, these techniques face the difficulty arising from the scarcity of visual preference data, which is required to train a visual reward model (VRM). In this work, we continue the line of research. We present a Robust Visual Reward Model (RoVRM) which improves human-preference alignment for LVLMs. RoVRM leverages auxiliary textual preference data through a three-phase progressive training and optimal transport-based preference data selection to effectively mitigate the scarcity of visual preference data. We experiment with RoVRM on the commonly used vision-language tasks based on the LLaVA-1.5-7B and -13B models. Experimental results demonstrate that RoVRM consistently outperforms traditional VRMs. Furthermore, our three-phase progressive training and preference data selection approaches can yield consistent performance gains over ranking-based alignment techniques, such as direct preference optimization.
Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Murun Yang, Qiaozhi He, Tong Xiao 0001, Chunliang Zhang, Tongran Liu
AAAI7
2025 Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention Learning
abstract
Lei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu, Yangfan Ye, Liang Zhao, Weihong Zhong, Baoxin Wang, Dayong Wu, Guoping Hu, Lingpeng Kong, Tong Xiao, Ting Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Lei Huang 0021, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu 0004, Yangfan Ye, Weihong Zhong, Baoxin Wang, Dayong Wu, Lingpeng Kong, Tong Xiao 0001, Ting Liu 0001, Bing Qin 0001
ACL (1)14
2025 Lost in Literalism: How Supervised Training Shapes Translationese in LLMs
abstract
Large language models (LLMs) have achieved remarkable success in machine translation, demonstrating impressive performance across diverse languages. However, translationese—characterized by overly literal and unnatural translations—remains a persistent challenge in LLM-based translation systems. Despite their pre-training on vast corpora of natural utterances, LLMs exhibit translationese errors and generate unexpected unnatural translations, stemming from biases introduced during supervised fine-tuning (SFT). In this work, we systematically evaluate the prevalence of translationese in LLM-generated translations and investigate its roots during supervised training. We introduce methods to mitigate these biases, including polishing golden references and filtering unnatural training instances. Empirical evaluations demonstrate that these approaches significantly reduce translationese while improving translation naturalness, validated by human evaluations and automatic metrics. Our findings highlight the need for training-aware adjustments to optimize LLM translation outputs, paving the way for more fluent and target-language-consistent translations.
Yafu Li, Ronghao Zhang, Zhilin Wang, Leyang Cui, Yongjing Yin, Tong Xiao 0001, Yue Zhang 0004
ACL (1)7
2025 Enhancing Neural Machine Translation Through Target Language Data: A kNN-LM Approach for Domain Adaptation
abstract
Neural machine translation (NMT) has advanced significantly, yet challenges remain in adapting to new domains . In scenarios where bilingual data is limited, this issue is further exacerbated. To address this, we propose kNN-LM-NMT, a method that leverages semantically similar target language sentences in the kNN framework. Our approach generates a probability distribution over these sentences during decoding, and this distribution is then interpolated with the NMT model’s distribution. Additionally, we introduce an n-gram-based approach to focus on similar fragments, enabling the model to avoid the noise introduced by the non-similar parts. To enhance accuracy, we further incorporate cross-lingual retrieval similarity to refine the kNN probability distribution. Extensive experiments on multi-domain datasets demonstrate significant performance improvements in both high-resource and low-resource scenarios. Our approach effectively extracts translation knowledge from limited target domain data, and well benefits from large-scale monolingual data for robust context representation.
Abudurexiti Reheman, Junhao Ruan, Abudukeyumu Abudula, Yingfeng Luo, Tong Xiao 0001
ACL (1)6
2025 SLAM: Towards Efficient Multilingual Reasoning via Selective Language Alignment
abstract
Despite the significant improvements achieved by large language models (LLMs) in English reasoning tasks, these models continue to struggle with multilingual reasoning. Recent studies leverage a full-parameter and two-stage training paradigm to teach models to first understand non-English questions and then reason. However, this method suffers from both substantial computational resource computing and catastrophic forgetting. The fundamental cause is that, with the primary goal of enhancing multilingual comprehension, an excessive number of irrelevant layers and parameters are tuned during the first stage. Given our findings that the representation learning of languages is merely conducted in lower-level layers, we propose an efficient multilingual reasoning alignment approach that precisely identifies and fine-tunes the layers responsible for handling multilingualism. Experimental results show that our method, SLAM, only tunes 6 layers’ feed-forward sub-layers including 6.5-8% of all parameters within 7B and 13B LLMs, achieving superior average performance than all strong baselines across 10 languages. Meanwhile, SLAM only involves one training stage, reducing training time by 4.1-11.9× compared to the two-stage method.
Yuchun Fan, Yongyu Mu, Lei Huang 0021, Junhao Ruan, Tong Xiao 0001, Shujian Huang
COLING7
2025 Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models
abstract
Kaiyan Chang, Yonghao Shi, Chenglong Wang, Hang Zhou, Chi Hu, Xiaoqian Liu, Yingfeng Luo, Yuan Ge, Tong Xiao, JingBo Zhu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Kaiyan Chang 0001, Yonghao Shi, Chenglong Wang 0002, Chi Hu, Yingfeng Luo, Yuan Ge 0001, Tong Xiao 0001
EMNLP9
2025 IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
abstract
Xinyu Liu, Bei Li, Jiahao Liu, Junhao Ruan, Kechen Jiao, Hongyin Tang, Jingang Wang, Tong Xiao, JingBo Zhu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Junhao Ruan, Kechen Jiao, Hongyin Tang, Jingang Wang, Tong Xiao 0001
EMNLP8
2025 Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
abstract
Weiqiao Shan, Yuang Li, Yuhao Zhang, Yingfeng Luo, Chen Xu, Xiaofeng Zhao, Long Meng, Yunfei Lu, Min Zhang, Hao Yang, Tong Xiao, JingBo Zhu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Weiqiao Shan, Yuang Li, Yingfeng Luo, Chen Xu 0008, Long Meng, Yunfei Lu, Min Zhang 0042, Hao Yang 0006, Tong Xiao 0001
EMNLP11
2025 A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
abstract
Simultaneous Speech Translation (SimulST) involves generating target language text while continuously processing streaming speech input, presenting significant real-time challenges. Multi-task learning is often employed to enhance SimulST performance but introduces optimization conflicts between primary and auxiliary tasks, potentially compromising overall efficiency. The existing model-level conflict resolution methods are not well-suited for this task which exacerbates inefficiencies and leads to high GPU memory consumption. To address these challenges, we propose a Modular Gradient Conflict Mitigation (MGCM) strategy that detects conflicts at a finer-grained modular level and resolves them utilizing gradient projection. Experimental results demonstrate that MGCM significantly improves SimulST performance, particularly under medium and high latency conditions, achieving a 0.68 BLEU score gain in offline tasks. Additionally, MGCM reduces GPU memory consumption by over 95% compared to other conflict mitigation methods, establishing it as a robust solution for SimulST tasks.
Yangfan Du, Jianjin Wang, Yuan Ge 0001, Chen Xu 0008, Tong Xiao 0001, Guocheng Chen
ICASSP6
2025 Adaptive Decoding for Efficient Automatic Speech Recognition
abstract
The latency and computation demand of End-to-end (E2E) automatic speech recognition (ASR) models hinder their deployment on lightweight devices. Despite there are many methods proposed for efficiency, the computational burden of the output layer with a large vocabulary is still a major challenge for decoders. In this paper, we propose an adaptive decoding method (ADD) to reduce the latency. Based on the vocal features of words like phonemes or speech units, we cluster the original vocabulary into small sets, allowing the model to inference more efficiently with a smaller search space. Experimental results demonstrate that our method significantly reduces the calculation FLOPs while maintaining performance. We also provide a deeper understanding of speech units from the perspective of phonemes.
Xiangnan Ma, Peizhuo Liu, Kaiqi Kou, Chenghao Gao, Tong Xiao 0001
ICASSP6
2025 Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models
abstract
Previous work on augmenting large multimodal models (LMMs) for text-to-image (T2I) generation has focused on enriching the input space of in-context learning (ICL). This includes providing a few demonstrations and optimizing image descriptions to be more detailed and logical. However, as demand for more complex and flexible image descriptions grows, enhancing comprehension of input text within the ICL paradigm remains a critical yet underexplored area. In this work, we extend this line of research by constructing parallel multilingual prompts aimed at harnessing the multilingual capabilities of LMMs. More specifically, we translate the input text into several languages and provide the models with both the original text and the translations. Experiments on two LMMs across 3 benchmarks show that our method, PMT2I, achieves superior performance in general, compositional, and fine-grained assessments, especially in human preference alignment Additionally, with its advantage of generating more diverse images, PMT2I significantly outperforms baseline prompts when incorporated with reranking methods. Our code and parallel multilingual data can be found at https://github.com/takagi97/PMT2I.
Yongyu Mu, Junxin Wang, Xiaoxuan Zhou, Chenglong Wang 0002, Yingfeng Luo, Qiaozhi He, Tong Xiao 0001, Guocheng Chen
ICASSP8
2025 Optimizing Speech Multi-View Feature Fusion through Conditional Computation
abstract
Recent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech representations. However, our study reveals that while SSL features expedite model convergence, they conflict with traditional spectral features like FBanks in terms of update directions. In response, we propose a novel generalized feature fusion framework grounded in conditional computation, featuring a gradient-sensitive gating network and a multi-stage dropout strategy. This framework mitigates feature conflicts and bolsters model robustness to multi-view input features. By integrating SSL and spectral features, our approach accelerates convergence and maintains performance on par with spectral models across multiple speech translation tasks on the MUSTC dataset.
Weiqiao Shan, Yuchen Han 0001, Yuang Li, Min Zhang 0042, Hao Yang 0006, Tong Xiao 0001
ICASSP9
2025 GRAM: A Generative Foundation Reward Model for Reward Generalization
abstract
In aligning large language models (LLMs), reward models have played an important role, but are standardly trained as discriminative models and rely only on labeled human preference data. In this paper, we explore methods that train reward models using both unlabeled and labeled data. Building on the generative models in LLMs, we develop a generative reward model that is first trained via large-scale unsupervised learning and then fine-tuned via supervised learning. We also show that by using label smoothing, we are in fact optimizing a regularized pairwise ranking loss. This result, in turn, provides a new view of training reward models, which links generative models and discriminative models under the same class of training objectives. The outcome of these techniques is a foundation reward model, which can be applied to a wide range of tasks with little or no further fine-tuning effort. Extensive experiments show that this model generalizes well across several tasks, including response ranking, reinforcement learning from human feedback, and task adaptation with fine-tuning, achieving significant performance improvements over several strong baseline models.
Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Qiaozhi He, Murun Yang, Tong Xiao 0001, Chunliang Zhang, Tongran Liu
ICML8
2025 ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
abstract
Large language models (LLMs) integrated with retrieval-augmented generation (RAG) have improved factuality by grounding outputs in external evidence. However, they remain susceptible to unfaithful generation, where outputs contradict retrieved context despite its relevance and accuracy. Existing approaches aiming to improve faithfulness primarily focus on enhancing the utilization of external context, but often overlook the persistent influence of internal parametric knowledge during generation. In this work, we investigate the internal mechanisms behind unfaithful generation and identify a subset of mid-to-deep feed-forward networks (FFNs) that are disproportionately activated in such cases. Building on this insight, we propose Parametric Knowledge Muting through FFN Suppression (ParamMute), a framework that improves contextual faithfulness by suppressing the activation of unfaithfulness-associated FFNs and calibrating the model toward retrieved knowledge. To evaluate our approach, we introduce CoFaithfulQA, a benchmark specifically designed to evaluate faithfulness in scenarios where internal knowledge conflicts with accurate external evidence. Experimental results show that ParamMute significantly enhances faithfulness across both CoFaithfulQA and the established ConFiQA benchmark, achieving substantial reductions in reliance on parametric memory. These findings underscore the importance of mitigating internal knowledge dominance and provide a new direction for improving LLM trustworthiness in RAG. All codes are available at https://github.com/OpenBMB/ParamMute.
Pengcheng Huang 0004, Zhenghao Liu 0001, Yukun Yan, Xiaoyuan Yi, Zhiyuan Liu 0001, Maosong Sun 0001, Tong Xiao 0001, Ge Yu 0001, Chenyan Xiong
NeurIPS9
2025 MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization
abstract
Recent advances in diffusion language models (DLMs) have presented a promising alternative to traditional autoregressive large language models (LLMs). However, DLMs still lag behind LLMs in reasoning performance, especially as the number of denoising steps decreases. Our analysis reveals that this shortcoming arises primarily from the independent generation of masked tokens across denoising steps, which fails to capture the token correlation. In this paper, we define two types of token correlation: intra-sequence correlation and inter-sequence correlation, and demonstrate that enhancing these correlations improves reasoning performance. To this end, we propose a Multi-Reward Optimization (MRO) approach, which encourages DLMs to consider the token correlation during the denoising process. More specifically, our MRO approach leverages test-time scaling, reject sampling, and reinforcement learning to directly optimize the token correlation with multiple elaborate rewards. Additionally, we introduce group step and importance sampling strategies to mitigate reward variance and enhance sampling efficiency. Through extensive experiments, we demonstrate that MRO not only improves reasoning performance but also achieves significant sampling speedups while maintaining high performance on reasoning benchmarks.
Chenglong Wang 0002, Yang Gan, Chi Hu, Yongyu Mu, Murun Yang, Chunliang Zhang, Tongran Liu, Zhengtao Yu 0001, Tong Xiao 0001
NeurIPS13
2025 StoryBench: A Dataset for Diverse, Explainable, Multi-hop Narrative Text-to-Image Generation
Yuan Ge 0001, Kaiyang Ye, Saihan Chen, Aokai Hao, Xiangnan Ma, Kaiyan Chang 0001, Tong Xiao 0001
NLPCC (2)7
2025 Pseudo-kNN-MT: Enhancing domain adaptability of neural machine translation via target language data
Abudurexiti Reheman, Yingfeng Luo, Junhao Ruan, Tong Xiao 0001
Knowl. Based Syst.5
2024 ESRL: Efficient Sampling-Based Reinforcement Learning for Sequence Generation
abstract
Applying Reinforcement Learning (RL) to sequence generation models enables the direct optimization of long-term rewards (e.g., BLEU and human feedback), but typically requires large-scale sampling over a space of action sequences. This is a computational challenge as presented by the practice of sequence generation problems, such as machine translation, where we often deal with a large action space (e.g., a vocabulary) and a long action sequence (e.g., a translation). In this work, we introduce two-stage sampling and dynamic sampling approaches to improve the sampling efficiency during training sequence generation models via RL. We experiment with our approaches on the traditional sequence generation tasks, including machine translation and abstractive summarization. Furthermore, we evaluate our approaches in RL from human feedback (RLHF) through training a large language model using the reward model. Experimental results show that the efficient sampling-based RL, referred to as ESRL, can outperform all baselines in terms of both training efficiency and memory consumption. Notably, ESRL yields consistent performance gains over the strong REINFORCE, minimum risk training, and proximal policy optimization methods. The code is available at https://github.com/wangclnlp/DeepSpeed-Chat-Extension/examples/esrl.
Chenglong Wang 0002, Yimin Hu, Yifu Huo, Tongran Liu, Tong Xiao 0001
AAAI7
2024 EIT: Enhanced Interactive Transformer
abstract
Two principles: the complementary principle and the consensus principle are widely acknowledged in the literature of multi-view learning.However, the current design of multihead self-attention, an instance of multi-view learning, prioritizes the complementarity while ignoring the consensus.To address this problem, we propose an enhanced multi-head selfattention (EMHA).First, to satisfy the complementary principle, EMHA removes the oneto-one mapping constraint among queries and keys in multiple subspaces and allows each query to attend to multiple keys.On top of that, we develop a method to fully encourage consensus among heads by introducing two interaction models, namely inner-subspace interaction and cross-subspace interaction.Extensive experiments on a wide range of language tasks (e.g., machine translation, abstractive summarization and grammar correction, language modeling), show its superiority, with a very modest increase in model size.Our code would be available at: https://github.com/zhengkid/EI T-Enhanced-Interactive-Transformer.
Huiwen Bao, Tong Xiao 0001
ACL (1)4
2024 RankPrompt: Step-by-Step Comparisons Make Language Models Better Reasoners
abstract
Large Language Models (LLMs) have achieved impressive performance across various reasoning tasks. However, even state-of-the-art LLMs such as ChatGPT are prone to logical errors during their reasoning processes. Existing solutions, such as deploying task-specific verifiers or voting over multiple reasoning paths, either require extensive human annotations or fail in scenarios with inconsistent responses. To address these challenges, we introduce RankPrompt, a new prompting method that enables LLMs to self-rank their responses without additional resources. RankPrompt breaks down the ranking problem into a series of comparisons among diverse responses, leveraging the inherent capabilities of LLMs to generate chains of comparison as contextual exemplars. Our experiments across 11 arithmetic and commonsense reasoning tasks show that RankPrompt significantly enhances the reasoning performance of ChatGPT and GPT-4, with improvements of up to 13%. Moreover, RankPrompt excels in LLM-based automatic evaluations for open-ended tasks, aligning with human judgments 74% of the time in the AlpacaEval dataset. It also exhibits robustness to variations in response order and consistency. Collectively, our results validate RankPrompt as an effective method for eliciting high-quality feedback from language models.
Chi Hu, Yuan Ge 0001, Xiangnan Ma, Qiang Li 0022, Yonghua Yang, Tong Xiao 0001
LREC/COLING7
2024 Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation
abstract
Yuan Ge, Yilun Liu, Chi Hu, Weibin Meng, Shimin Tao, Xiaofeng Zhao, Mahong Xia, Zhang Li, Boxing Chen, Hao Yang, Bei Li, Tong Xiao, JingBo Zhu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yuan Ge 0001, Yilun Liu 0001, Chi Hu, Weibin Meng, Shimin Tao, Mahong Xia, Boxing Chen, Hao Yang 0006, Tong Xiao 0001
EMNLP12
2024 Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-Context Models
abstract
Numerous recent works target to extend effective context length for language models and various methods, tasks and benchmarks exist to measure model's effective memorization length.However, through thorough investigations, we find limitations for currently existing evaluations on model's memorization capability.We provide an extensive survey for limitations in this work and propose a new method called forgetting curve to measure the memorization capability of long-context models.We show that forgetting curve has the advantage of being robust to the tested corpus and the experimental settings, of not relying on prompts and can be applied to any model size.We apply our forgetting curve to a large variety of models involving both transformer and RNN/SSM based architectures.Our measurement provides empirical evidence for the effectiveness of transformer extension techniques while raises questions for the effective length of RNN/SSM based models.We also examine the difference between our measurement and existing benchmarks as well as popular metrics for various models.Our code and results can be found at https://github.com/1azybug/ForgettingCurve.
Runsong Zhao, Pengcheng Huang 0004, Chunyang Xiao, Jingang Wang, Tong Xiao 0001
EMNLP7
2024 Revealing the Parallel Multilingual Learning within Large Language Models
abstract
Yongyu Mu, Peinan Feng, Zhiquan Cao, Yuzhang Wu, Bei Li, Chenglong Wang, Tong Xiao, Kai Song, Tongran Liu, Chunliang Zhang, JingBo Zhu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yongyu Mu, Peinan Feng, Zhiquan Cao, Yuzhang Wu, Chenglong Wang 0002, Tong Xiao 0001, Tongran Liu, Chunliang Zhang
EMNLP7
2024 Bridging the Gaps of Both Modality and Language: Synchronous Bilingual CTC for Speech Translation and Speech Recognition
abstract
In this study, we present synchronous bilingual Connectionist Temporal Classification (CTC), an innovative framework that leverages dual CTC to bridge the gaps of both modality and language in the speech translation (ST) task. Utilizing transcript and translation as concurrent objectives for CTC, our model bridges the gap between audio and text as well as between source and target languages. Building upon the recent advances in CTC application, we develop an enhanced variant, BiL-CTC+, that establishes new state-of-the-art performances on the MuST-C ST benchmarks under resource-constrained scenarios. Intriguingly, our method also yields significant improvements in speech recognition performance, revealing the effect of cross-lingual learning on transcription and demonstrating its broad applicability. The source code is available at https://github.com/xuchennlp/S2T.
Chen Xu 0008, Erfeng He, Qianqian Dong, Tong Xiao 0001, Dapeng Man, Wu Yang 0001
ICASSP6
2024 Soft Alignment of Modality Space for End-to-End Speech Translation
abstract
End-to-end Speech Translation (ST) aims to convert speech into target text within a unified model. The inherent differences between speech and text modalities often impede effective cross-modal and cross-lingual transfer. Existing methods typically employ hard alignment (H-Align) of individual speech and text segments, which can degrade textual representations. To address this, we introduce Soft Alignment (S-Align), using adversarial training to align the representation spaces of both modalities. S-Align creates a modality-invariant space while preserving individual modality quality. Experiments on three languages from the MuST-C dataset show S-Align outperforms H-Align across multiple tasks and offers translation capabilities on par with specialized translation models.
Kaiqi Kou, Chen Xu 0008, Chunliang Zhang, Tong Xiao 0001
ICASSP6
2024 Recent Advances in End-to-End Simultaneous Speech Translation
Yangfan Du, Erfeng He, Yingfeng Luo, Chen Xu 0008, Tong Xiao 0001
IJCAI7
2024 Trans-Rotor: An Active Omnidirectional Aerial-Ground Vehicle With Differential Gear Joint Transformation Mechanism
abstract
Aerial-ground vehicles have shown great potential in various fields due to their superior mobility and outstanding endurance. However, most of morphing aerial-ground vehicles consider little about controllability and traversability in ground mode. We present a novel aerial-ground vehicle called TransRotor. By proposing a differential gear joint, we equip TransRotor with omnidirectional mobility in both air and ground mode. Besides, using a four-wheel-steering model in ground mode provides better traversability and ground flexibility. Moreover, we design mid-mode transformation for Trans-Rotor, which provides smooth and rapid mode switching. In this work, we firstly propose a novel design of an aerial-ground vehicle. Then, we propose a decoupled controller considering the four-wheel-steer model to achieve autonomous navigation of the vehicle. Comprehensive experiments and a benchmark comparison are carried out to validate the outstanding performance of the proposed system, where the system shows ground flexibility and saves energy up to more than 95%.
Xuankang Wu, Haoxiang Sun, Tong Xiao 0001, Yanzhang Pan, Zheng Fang 0001
IROS3
2024 Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient Learning
abstract
Residual networks, as discrete approximations of Ordinary Differential Equations (ODEs), have inspired significant advancements in neural network design, including multistep methods, high-order methods, and multi-particle dynamical systems. The precision of the solution to ODEs significantly affects parameter optimization, thereby impacting model performance. In this work, we present a series of advanced explorations of Transformer architecture design to minimize the error compared to the true ``solution.'' First, we introduce a predictor-corrector learning framework to minimize truncation errors, which consists of a high-order predictor and a multistep corrector. Second, we propose an exponential moving average-based coefficient learning method to strengthen our higher-order predictor. Extensive experiments on large-scale machine translation, abstractive summarization, language modeling, and natural language understanding benchmarks demonstrate the superiority of our approach. On the WMT'14 English-German and English-French tasks, our model achieved BLEU scores of 30.95 and 44.27, respectively. Furthermore, on the OPUS multilingual machine translation task, our model surpasses a robust 3.8B DeepNet by an average of 2.9 SacreBLEU, using only 1/3 parameters. Notably, it also beats LLama models by 5.7 accuracy points on the LM Harness Evaluation.
Rui Wang 0028, Qingyan Guo, Junliang Guo, Xu Tan 0003, Tong Xiao 0001, Jingang Wang
NeurIPS8
2024 Progressive and Consistent Subword Regularization for Neural Machine Translation
Yongqi Gao, Yingfeng Luo, Qinghong Zhang, Huibo Shao, Tong Xiao 0001
NLPCC (3)5
2024 DFS-QA: Dynamic Frame Selection for Better Video Question Answering
Zhibo Ren, Baoyu Hou, Huizhen Wang, Muhua Zhu, Tong Xiao 0001
NLPCC (3)5
2024 Improving End-to-End Speech Translation with Progressive Dual Encoding
Runlai Zhang, Saihan Chen, Yangfan Du, Tong Xiao 0001
NLPCC (3)6
2024 SPMIS: An Investigation of Synthetic Spoken Misinformation Detection
abstract
In recent years, speech generation technology has advanced rapidly, fueled by generative models and large-scale training techniques. While these developments have enabled the production of high-quality synthetic speech, they have also raised concerns about the misuse of this technology, particularly for generating synthetic misinformation. Current research primarily focuses on distinguishing machine-generated speech from human-produced speech, but the more urgent challenge is detecting misinformation within spoken content. This task requires a thorough analysis of factors such as speaker identity, topic, and synthesis. To address this need, we conduct an initial investigation into synthetic spoken misinformation detection by introducing an open-source dataset, SpMis. SpMis includes speech synthesized from over 1,000 speakers across five common topics, utilizing state-of-the-art text-to-speech systems. Although our results show promising detection capabilities, they also reveal substantial challenges for practical implementation, underscoring the importance of ongoing research in this critical area.
Peizhuo Liu, Renqiang He, Haorui He, Huadi Zheng, Jie Shi 0005, Tong Xiao 0001, Zhizheng Wu 0001
SLT8
2023 Prompting Neural Machine Translation with Translation Memories
abstract
Improving machine translation (MT) systems with translation memories (TMs) is of great interest to practitioners in the MT community. However, previous approaches require either a significant update of the model architecture and/or additional training efforts to make the models well-behaved when TMs are taken as additional input. In this paper, we present a simple but effective method to introduce TMs into neural machine translation (NMT) systems. Specifically, we treat TMs as prompts to the NMT model at test time, but leave the training process unchanged. The result is a slight update of an existing NMT system, which can be implemented in a few hours by anyone who is familiar with NMT. Experimental results on several datasets demonstrate that our system significantly outperforms strong baselines.
Abudurexiti Reheman, Yingfeng Luo, Tong Xiao 0001
AAAI5
2023 Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text Data
abstract
We present a method for introducing a text encoder into pre-trained end-to-end speech translation systems. It enhances the ability of adapting one modality (i.e., source-language speech) to another (i.e., source-language text). Thus, the speech translation model can learn from both unlabeled and labeled data, especially when the source-language text data is abundant. Beyond this, we present a denoising method to build a robust text encoder that can deal with both normal and noisy text data. Our system sets new state-of-the-arts on the MuST-C En-De, En-Fr, and LibriSpeech En-Fr tasks.
Chen Xu 0008, Bojie Hu, Chunliang Zhang, Tong Xiao 0001
AAAI5
2023 CTC-based Non-autoregressive Speech Translation
abstract
Chen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun, Yuhao Zhang, Murun Yang, Qianqian Dong, Tom Ko, Mingxuan Wang, Tong Xiao, Anxiang Ma, Jingbo Zhu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Chen Xu 0008, Qingxuan Sun, Murun Yang, Qianqian Dong, Tom Ko, Mingxuan Wang, Tong Xiao 0001, Anxiang Ma
ACL (1)10
2023 Rethinking and Improving Multi-task Learning for End-to-end Speech Translation
abstract
Significant improvements in end-to-end speech translation (ST) have been achieved through the application of multi-task learning.However, the extent to which auxiliary tasks are highly consistent with the ST task, and how much this approach truly helps, have not been thoroughly studied.In this paper, we investigate the consistency between different tasks, considering different times and modules.We find that the textual encoder primarily facilitates cross-modal conversion, but the presence of noise in speech impedes the consistency between text and speech representations.Furthermore, we propose an improved multi-task learning (IMTL) approach for the ST task, which bridges the modal gap by mitigating the difference in length and representation.We conduct experiments on the MuST-C dataset.The results demonstrate that our method attains stateof-the-art results.Moreover, when additional data is used, we achieve the new SOTA result on MuST-C English to Spanish task with 20.8% of the training time required by the current SOTA method.
Chen Xu 0008, Tong Xiao 0001, Chunliang Zhang
EMNLP5
2023 Recent Advances in Direct Speech-to-text Translation
abstract
Recently, speech-to-text translation has attracted more and more attention and many studies have emerged rapidly. In this paper, we present a comprehensive survey on direct speech translation aiming to summarize the current state-of-the-art techniques. First, we categorize the existing research work into three directions based on the main challenges --- modeling burden, data scarcity, and application issues. To tackle the problem of modeling burden, two main structures have been proposed, encoder-decoder framework (Transformer and the variants) and multitask frameworks. For the challenge of data scarcity, recent work resorts to many sophisticated techniques, such as data augmentation, pre-training, knowledge distillation, and multilingual modeling. We analyze and summarize the application issues, which include real-time, segmentation, named entity, gender bias, and code-switching. Finally, we discuss some promising directions for future work.
Chen Xu 0008, Rong Ye, Qianqian Dong, Chengqi Zhao, Tom Ko, Mingxuan Wang, Tong Xiao 0001
IJCAI7
2023 Information Magnitude Based Dynamic Sub-sampling for Speech-to-text
Chenghao Gao, Kaiqi Kou, Chen Xu 0008, Tong Xiao 0001
INTERSPEECH5
2023 Learning Reliable Neural Networks with Distributed Architecture Representations
abstract
Neural architecture search (NAS) has shown the strong performance of learning neural models automatically in recent years. But most NAS systems are unreliable due to the architecture gap brought by discrete representations of atomic architectures. In this article, we improve the performance and robustness of NAS via narrowing the gap between architecture representations. More specifically, we apply a general contraction mapping to model neural networks with distributed representations (Neural Architecture Search with Distributed Architecture Representations (ArchDAR)). Moreover, for a better search result, we present a joint learning approach to integrating distributed representations with advanced architecture search methods. We implement our ArchDAR in a differentiable architecture search model and test learned architectures on the language modeling task. On the Penn Treebank data, it outperforms a strong baseline significantly by 1.8 perplexity scores. Also, the search process with distributed representations is more stable, which yields a faster structural convergence when it works with the differentiable architecture search model.
Yinqiao Li, Runzhe Cao, Qiaozhi He, Tong Xiao 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2022 ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation
abstract
Bei Li, Quan Du, Tao Zhou, Yi Jing, Shuhan Zhou, Xin Zeng, Tong Xiao, JingBo Zhu, Xuebo Liu, Min Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Quan Du, Yi Jing, Shuhan Zhou, Tong Xiao 0001, Xuebo Liu 0002, Min Zhang 0005
ACL (1)7
2022 On Vision Features in Multimodal Machine Translation
abstract
Previous work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is on the quality of vision models.In this work, we investigate the impact of vision models on MMT.Given the fact that Transformer is becoming popular in computer vision, we experiment with various strong models (such as Vision Transformer) and enhanced features (such as object-detection and image captioning).We develop a selective attention model to study the patch-level contribution of an image in MMT.On detailed probing tasks, we find that stronger vision models are helpful for learning translation from the visual modality.Our results also suggest the need of carefully examining MMT models, especially when current benchmarks are small-scale and biased.Our code could be found at https: //github.com/libeineu/fairseq_mmt.
Chuanhao Lv, Zefan Zhou, Tong Xiao 0001, Anxiang Ma
ACL (1)5
2022 Learning Multiscale Transformer Models for Sequence Generation
abstract
Multiscale feature hierarchies have been witnessed the success in the computer vision area. This further motivates researchers to design multiscale Transformer for natural language processing, mostly based on the self-attention mechanism. For example, restricting the receptive field across heads or extracting local fine-grained features via convolutions. However, most of existing works directly modeled local features but ignored the word-boundary information. This results in redundant and ambiguous attention distributions, which lacks of interpretability. In this work, we define those scales in different linguistic units, including sub-words, words and phrases. We built a multiscale Transformer model by establishing relationships among scales based on word-boundary information and phrase-level prior knowledge. The proposed \textbf{U}niversal \textbf{M}ulti\textbf{S}cale \textbf{T}ransformer, namely \textsc{Umst}, was evaluated on two sequence generation tasks. Notably, it yielded consistent performance gains over the strong baseline on several test sets without sacrificing the efficiency.
Yi Jing, Chengbo Jiao, Tong Xiao 0001
ICML5
2022 Coarse-to-Fine Output Predictions for Efficient Decoding in Neural Machine Translation
abstract
Neural Machine Translation (NMT) systems are undesirably slow as the decoder often has to compute probability distributions over large target vocabularies. In this work, we propose a coarse-to-fine approach to reduce the complexity of the decoding process, using only the information of the weight matrix in the Softmax layer. The large target vocabulary is first trimmed to a small candidate set in the coarse-grained phase, and from this candidate set the final top- k results are generated in the fine-grained phase. Tested on an RNN-based NMT system and a Transformer-based NMT system separately, our GPU-friendly method achieved a significant speed-up without harming the translation quality.
Oi Yee Kwong, Yinqiao Li, Tong Xiao 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2021 An Efficient Transformer Decoder with Compressed Sub-layers
abstract
The large attention-based encoder-decoder network (Transformer) has become prevailing recently due to its effectiveness. But the high computation complexity of its decoder raises the inefficiency issue. By examining the mathematic formulation of the decoder, we show that under some mild conditions, the architecture could be simplified by compressing its sub-layers, the basic building block of Transformer, and achieves a higher parallelism. We thereby propose Compressed Attention Network, whose decoder layer consists of only one sub-layer instead of three. Extensive experiments on 14 WMT machine translation tasks show that our model is 1.42x faster with performance on par with a strong baseline. This strong baseline is already 2x faster than the widely used standard baseline without loss in performance.
Yanyang Li, Tong Xiao 0001
AAAI3
2021 Learning Light-Weight Translation Models from Deep Transformer
abstract
Recently, deep models have shown tremendous improvements in neural machine translation (NMT). However, systems of this kind are computationally expensive and memory intensive. In this paper, we take a natural step towards learning strong but light-weight NMT systems. We proposed a novel group-permutation based knowledge distillation approach to compressing the deep Transformer model into a shallow model. The experimental results on several benchmarks validate the effectiveness of our method. Our compressed model is 8 times shallower than the deep model, with almost no loss in BLEU. To further enhance the teacher model, we present a Skipping Sub-Layer method to randomly omit sub-layers to introduce perturbation into training, which achieves a BLEU score of 30.63 on English-German newstest2014. The code is publicly available at https://github.com/libeineu/GPKD.
Quan Du, Tong Xiao 0001, Chunliang Zhang
AAAI5
2021 Weight Distillation: Transferring the Knowledge in Neural Network Parameters
abstract
Ye Lin, Yanyang Li, Ziyang Wang, Bei Li, Quan Du, Tong Xiao, Jingbo Zhu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yanyang Li, Quan Du, Tong Xiao 0001
ACL/IJCNLP (1)6
2021 Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders
abstract
Chen Xu, Bojie Hu, Yanyang Li, Yuhao Zhang, Shen Huang, Qi Ju, Tong Xiao, Jingbo Zhu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chen Xu 0008, Bojie Hu, Yanyang Li, Shen Huang, Qi Ju 0002, Tong Xiao 0001
ACL/IJCNLP (1)7
2021 RankNAS: Efficient Neural Architecture Search by Pairwise Ranking
abstract
This paper addresses the efficiency challenge of Neural Architecture Search (NAS) by formulating the task as a ranking problem.Previous methods require numerous training examples to estimate the accurate performance of architectures, although the actual goal is to find the distinction between "good" and "bad" candidates.Here we do not resort to performance predictors.Instead, we propose a performance ranking method (RankNAS) via pairwise ranking.It enables efficient architecture search using much fewer training examples.Moreover, we develop an architecture selection method to prune the search space and concentrate on more promising candidates.Extensive experiments on machine translation and language modeling tasks show that RankNAS can design high-performance architectures while being orders of magnitude faster than state-ofthe-art NAS systems.
Chi Hu, Chenglong Wang 0002, Xiangnan Ma, Xia Meng, Yinqiao Li, Tong Xiao 0001, Changliang Li
EMNLP (1)6
2021 Non-Autoregressive Translation by Learning Target Categorical Codes
abstract
Yu Bao, Shujian Huang, Tong Xiao, Dongqi Wang, Xinyu Dai, Jiajun Chen. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Shujian Huang, Tong Xiao 0001, Dongqi Wang 0005, Xinyu Dai, Jiajun Chen 0001
NAACL-HLT3
2021 Chinese Poetry Generation with Metrical Constraints
Yingfeng Luo, Changliang Li, Canan Huang, Chen Xu 0008, Binghao Wei, Tong Xiao 0001
NLPCC (1)7
2020 Neural Machine Translation with Joint Representation
abstract
Though early successes of Statistical Machine Translation (SMT) systems are attributed in part to the explicit modelling of the interaction between any two source and target units, e.g., alignment, the recent Neural Machine Translation (NMT) systems resort to the attention which partially encodes the interaction for efficiency. In this paper, we employ Joint Representation that fully accounts for each possible interaction. We sidestep the inefficiency issue by refining representations with the proposed efficient attention operation. The resulting Reformer models offer a new Sequence-to-Sequence modelling paradigm besides the Encoder-Decoder framework and outperform the Transformer baseline in either the small scale IWSLT14 German-English, English-German and IWSLT15 Vietnamese-English or the large scale NIST12 Chinese-English translation tasks by about 1 BLEU point. We also propose a systematic model scaling approach, allowing the Reformer model to beat the state-of-the-art Transformer in IWSLT14 German-English and NIST12 Chinese-English with about 50% fewer parameters. The code is publicly available at https://github.com/lyy1994/reformer.
Yanyang Li, Qiang Wang 0050, Tong Xiao 0001, Tongran Liu
AAAI3
2020 Learning Architectures from an Extended Search Space for Language Modeling
abstract
Neural architecture search (NAS) has advanced significantly in recent years but most NAS systems restrict search to learning architectures of a recurrent or convolutional cell.In this paper, we extend the search space of NAS.In particular, we present a general approach to learn both intra-cell and inter-cell architectures (call it ESS).For a better search result, we design a joint learning method to perform intra-cell and inter-cell NAS simultaneously.We implement our model in a differentiable architecture search system.For recurrent neural language modeling, it outperforms a strong baseline significantly on the PTB and Wiki-Text data, with a new state-of-the-art on PTB.Moreover, the learned architectures show good transferability to other systems.E.g., they improve state-of-the-art systems on the CoNLL and WNUT named entity recognition (NER) tasks and CoNLL chunking task, indicating a promising line of research on large-scale prelearned architectures.
Yinqiao Li, Chi Hu, Nuo Xu 0010, Yufan Jiang, Tong Xiao 0001, Tongran Liu, Changliang Li
ACL6
2020 Does Multi-Encoder Help? A Case Study on Context-Aware Neural Machine Translation
abstract
In encoder-decoder neural models, multiple encoders are in general used to represent the contextual information in addition to the individual sentence.In this paper, we investigate multi-encoder approaches in document-level neural machine translation (NMT).Surprisingly, we find that the context encoder does not only encode the surrounding sentences but also behaves as a noise generator.This makes us rethink the real benefits of multi-encoder in context-aware translation -some of the improvements come from robust training.We compare several methods that introduce noise and/or well-tuned dropout setup into the training of these encoders.Experimental results show that noisy training plays an important role in multi-encoder-based NMT, especially when the training data is small.Also, we establish a new state-of-the-art on IWSLT Fr-En task by careful use of noise generation and dropout methods.
Yufan Jiang, Tong Xiao 0001, Tongran Liu, Changliang Li
ACL5
2020 A Simple and Effective Approach to Robust Unsupervised Bilingual Dictionary Induction
abstract
Unsupervised Bilingual Dictionary Induction methods based on the initialization and the selflearning have achieved great success in similar language pairs, e.g., English-Spanish.But they still fail and have an accuracy of 0% in many distant language pairs, e.g., English-Japanese.In this work, we show that this failure results from the gap between the actual initialization performance and the minimum initialization performance for the self-learning to succeed.We propose Iterative Dimension Reduction to bridge this gap.Our experiments show that this simple method does not hamper the performance of similar language pairs and achieves an accuracy of 13.64∼55.53%between English and four distant languages, i.e., Chinese, Japanese, Vietnamese and Thai.
Yanyang Li, Yingfeng Luo, Quan Du, Huizhen Wang, Shujian Huang, Tong Xiao 0001
COLING7
2020 Layer-Wise Multi-View Learning for Neural Machine Translation
abstract
Traditional neural machine translation is limited to the topmost encoder layer's context representation and cannot directly perceive the lower encoder layers.Existing solutions usually rely on the adjustment of network architecture, making the calculation more complicated or introducing additional structural restrictions.In this work, we propose layer-wise multi-view learning to solve this problem, circumventing the necessity to change the model structure.We regard each encoder layer's off-the-shelf output, a by-product in layer-by-layer encoding, as the redundant view for the input sentence.In this way, in addition to the topmost encoder layer (referred to as the primary view), we also incorporate an intermediate encoder layer as the auxiliary view.We feed the two views to a partially shared decoder to maintain independent prediction.Consistency regularization based on KL divergence is used to encourage the two views to learn from each other.Extensive experimental results on five translation tasks show that our approach yields stable improvements over multiple strong baselines.As another bonus, our method is agnostic to network architectures and can maintain the same inference speed as the original model.
Qiang Wang 0050, Changliang Li, Yue Zhang 0004, Tong Xiao 0001
COLING4
2020 Dynamic Curriculum Learning for Low-Resource Neural Machine Translation
abstract
Large amounts of data has made neural machine translation (NMT) a big success in recent years.But it is still a challenge if we train these models on small-scale corpora.In this case, the way of using data appears to be more important.Here, we investigate the effective use of training data for low-resource NMT.In particular, we propose a dynamic curriculum learning (DCL) method to reorder training samples in training.Unlike previous work, we do not use a static scoring function for reordering.Instead, the order of training samples is dynamically determined in two ways -loss decline and model competence.This eases training by highlighting easy samples that the current model has enough competence to learn.We test our DCL method in a Transformerbased system.Experimental results show that DCL outperforms several strong baselines on three low-resource machine translation benchmarks and different sized data of WMT'16 En-De.
Chen Xu 0008, Bojie Hu, Yufan Jiang, Zeyang Wang, Shen Huang, Qi Ju 0002, Tong Xiao 0001
COLING8
2020 Shallow-to-Deep Training for Neural Machine Translation
abstract
Deep encoders have been proven to be effective in improving neural machine translation (NMT) systems, but training an extremely deep encoder is time consuming.Moreover, why deep models help NMT is an open question.In this paper, we investigate the behavior of a well-tuned deep Transformer system.We find that stacking layers is helpful in improving the representation ability of N-MT models and adjacent layers perform similarly.This inspires us to develop a shallowto-deep training method that learns deep models by stacking shallow models.In this way, we successfully train a Transformer system with a 54-layer encoder.Experimental results on WMT'16 English-German and WMT'14 English-French translation tasks show that it is 1.4 × faster than training from scratch, and achieves a BLEU score of 30.33 and 43.29 on two tasks.The code is publicly available at https://github.com/libeineu/ SDT-Training.
Yufan Jiang, Quan Du, Tong Xiao 0001, Huizhen Wang
EMNLP (1)6
2020 Towards Fully 8-bit Integer Inference for the Transformer Model
abstract
8-bit integer inference, as a promising direction in reducing both the latency and storage of deep neural networks, has made great progress recently. On the other hand, previous systems still rely on 32-bit floating point for certain functions in complex models (e.g., Softmax in Transformer), and make heavy use of quantization and de-quantization. In this work, we show that after a principled modification on the Transformer architecture, dubbed Integer Transformer, an (almost) fully 8-bit integer inference algorithm Scale Propagation could be derived. De-quantization is adopted when necessary, which makes the network more efficient. Our experiments on WMT16 En<->Ro, WMT14 En<->De and En->Fr translation tasks as well as the WikiText-103 language modelling task show that the fully 8-bit Transformer system achieves comparable performance with the floating point baseline but requires nearly 4x less memory footprint.
Yanyang Li, Tengbo Liu, Tong Xiao 0001, Tongran Liu
IJCAI4
2020 Towards Differentially Private Text Representations
abstract
Most deep learning frameworks require users to pool their local data or model updates to a trusted server to train or maintain a global model. The assumption of a trusted server who has access to user information is ill-suited in many applications. To tackle this problem, we develop a new deep learning framework under an untrusted server setting, which includes three modules: (1) embedding module, (2) randomization module, and (3) classifier module. For the randomization module, we propose a novel local differentially private (LDP) protocol to reduce the impact of privacy parameter ε on accuracy, and provide enhanced flexibility in choosing randomization probabilities for LDP. Analysis and experiments show that our framework delivers comparable or even better performance than the non-private framework and existing LDP protocols, demonstrating the advantages of our LDP protocol.
Lingjuan Lyu, Yitong Li 0002, Xuanli He, Tong Xiao 0001
SIGIR4
2019 Shared-Private Bilingual Word Embeddings for Neural Machine Translation
abstract
Word embedding is central to neural machine translation (NMT), which has attracted intensive research interest in recent years.In NMT, the source embedding plays the role of the entrance while the target embedding acts as the terminal.These layers occupy most of the model parameters for representation learning.Furthermore, they indirectly interface via a soft-attention mechanism, which makes them comparatively isolated.In this paper, we propose shared-private bilingual word embeddings, which give a closer relationship between the source and target embeddings, and which also reduce the number of model parameters.For similar source and target words, their embeddings tend to share a part of the features and they cooperatively learn these common representation units.Experiments on 5 language pairs belonging to 6 different language families and written in 5 different alphabets demonstrate that the proposed model provides a significant performance boost over the strong baselines with dramatically fewer model parameters.
Xuebo Liu 0002, Derek F. Wong, Yang Liu 0005, Lidia S. Chao, Tong Xiao 0001
ACL (1)5
2019 Learning Deep Transformer Models for Machine Translation
abstract
Transformer is the state-of-the-art model in recent machine translation evaluations. Two strands of research are promising to improve models of this kind: the first uses wide networks (a.k.a. Transformer-Big) and has been the de facto standard for development of the Transformer system, and the other uses deeper language representation but faces the difficulty arising from learning deep networks. Here, we continue the line of research on the latter. We claim that a truly deep Transformer model can surpass the Transformer-Big counterpart by 1) proper use of layer normalization and 2) a novel way of passing the combination of previous layers to the next. On WMT’16 English-German and NIST OpenMT’12 Chinese-English tasks, our deep system (30/25-layer encoder) outperforms the shallow Transformer-Big/Base baseline (6-layer encoder) by 0.4-2.4 BLEU points. As another bonus, the deep model is 1.6X smaller in size and 3X faster in training than Transformer-Big.
Qiang Wang 0050, Tong Xiao 0001, Changliang Li, Derek F. Wong, Lidia S. Chao
ACL (1)3
2019 Improved Differentiable Architecture Search for Language Modeling and Named Entity Recognition
abstract
Yufan Jiang, Chi Hu, Tong Xiao, Chunliang Zhang, Jingbo Zhu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yufan Jiang, Chi Hu, Tong Xiao 0001, Chunliang Zhang
EMNLP/IJCNLP (1)3
2019 Sharing Attention Weights for Fast Transformer
abstract
Recently, the Transformer machine translation system has shown strong results by stacking attention layers on both the source and target-language sides. But the inference of this model is slow due to the heavy use of dot-product attention in auto-regressive decoding. In this paper we speed up Transformer via a fast and lightweight attention model. More specifically, we share attention weights in adjacent layers and enable the efficient re-use of hidden states in a vertical manner. Moreover, the sharing policy can be jointly learned with the MT model. We test our approach on ten WMT and NIST OpenMT tasks. Experimental results show that it yields an average of 1.3X speed-up (with almost no decrease in BLEU) on top of a state-of-the-art implementation that has already adopted a cache for fast inference. Also, our approach obtains a 1.8X speed-up when it works with the AAN model. This is even 16 times faster than the baseline with no use of the attention cache.
Tong Xiao 0001, Yinqiao Li, Zhengtao Yu 0001, Tongran Liu
IJCAI1
2019 Analysis of Back-Translation Methods for Low-Resource Neural Machine Translation
Nuo Xu 0010, Yinqiao Li, Chen Xu 0008, Yanyang Li, Tong Xiao 0001
NLPCC (2)6
2018 Multi-layer Representation Fusion for Neural Machine Translation
abstract
Neural machine translation systems require a number of stacked layers for deep models. But the prediction depends on the sentence representation of the top-most layer with no access to low-level representations. This makes it more difficult to train the model and poses a risk of information loss to prediction. In this paper, we propose a multi-layer representation fusion (MLRF) approach to fusing stacked layers. In particular, we design three fusion functions to learn a better representation from the stack. Experimental results show that our approach yields improvements of 0.92 and 0.56 BLEU points over the strong Transformer baseline on IWSLT German-English and NIST Chinese-English MT tasks respectively. The result is new state-of-the-art in German-English translation.
Qiang Wang 0050, Fuxue Li, Tong Xiao 0001, Yanyang Li, Yinqiao Li
COLING3
2018 Source Segment Encoding for Neural Machine Translation
Qiang Wang 0050, Tong Xiao 0001
NLPCC (1)2
2018 Linguistic Knowledge-Aware Neural Machine Translation
abstract
Recently, researchers have shown an increasing interest in incorporating linguistic knowledge into neural machine translation (NMT). To this end, previous works choose either to alter the architecture of NMT encoder to incorporate syntactic information into the translation model, or to generalize the embedding layer of the encoder to encode additional linguistic features. The former approach mainly focuses on injecting the syntactic structure of the source sentence into the encoding process, leading to a complicated model that lacks the flexibility to incorporate other types of knowledge. The latter extends word embeddings by considering additional linguistic knowledge as features to enrich the word representation. It thus does not explicitly balance the contribution from word embeddings and the contribution from additional linguistic knowledge. To address these limitations, this paper proposes a knowledge-aware NMT approach that models additional linguistic features in parallel to the word feature. The core idea is that we propose modeling a series of linguistic features at the word level (knowledge block) using a recurrent neural network (RNN). And in sentence level, those word-corresponding feature blocks are further encoded using a RNN encoder. In decoding, we propose a knowledge gate and an attention gate to dynamically control the proportions of information contributing to the generation of target words from different sources. Extensive experiments show that our approach is capable of better accounting for importance of additional linguistic, and we observe significant improvements from 1.0 to 2.3 BLEU points on Chinese$\leftrightarrow$English and English$\rightarrow$German translation tasks.
Qiang Li 0022, Derek F. Wong, Lidia S. Chao, Muhua Zhu, Tong Xiao 0001, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.5
2017 Towards Bidirectional Hierarchical Representations for Attention-based Neural Machine Translation
abstract
This paper proposes a hierarchical attentional neural translation model which focuses on enhancing source-side hierarchical representations by covering both local and global semantic information using a bidirectional tree-based encoder.To maximize the predictive likelihood of target words, a weighted variant of an attention mechanism is used to balance the attentive information between lexical and phrase vectors.Using a tree-based rare word encoding, the proposed model is extended to sub-word level to alleviate the out-of-vocabulary (OOV) problem.Empirical results reveal that the proposed model significantly outperforms sequence-to-sequence attention-based and tree-based neural translation models in English-Chinese translation tasks.
Baosong Yang, Derek F. Wong, Tong Xiao 0001, Lidia S. Chao
EMNLP3
2017 Fast Parallel Training of Neural Language Models
abstract
Training neural language models (NLMs) is very time consuming and we need parallelization for system speedup. However, standard training methods have poor scalability across multiple devices (e.g., GPUs) due to the huge time cost required to transmit data for gradient sharing in the back-propagation process. In this paper we present a sampling-based approach to reducing data transmission for better scaling of NLMs. As a ''bonus'', the resulting model also improves the training speed on a single device. Our approach yields significant speed improvements on a recurrent neural network-based language model. On four NVIDIA GTX1080 GPUs, it achieves a speedup of 2.1+ times over the standard asynchronous stochastic gradient descent baseline, yet with no increase in perplexity. This is even 4.2 times faster than the naive single GPU counterpart.
Tong Xiao 0001, Tongran Liu, Chunliang Zhang
IJCAI1
2017 Implicit Syntactic Features for Target-dependent Sentiment Analysis
abstract
Targeted sentiment analysis investigates the sentiment polarities on given target mentions from input texts. Different from sentence level sentiment, it offers more fine-grained knowledge on each entity mention. While early work leveraged syntactic information, recent research has used neural representation learning to induce features automatically, thereby avoiding error propagation of syntactic parsers, which are particularly severe on social media texts. We study a method to leverage syntactic information without explicitly building the parser outputs, by training an encoder-decoder structure parser model on standard syntactic treebanks, and then leveraging its hidden encoder layers when analysing tweets. Such hidden vectors do not contain explicit syntactic outputs, yet encode rich syntactic features. We use them to augment the inputs to a baseline state-of-the-art targeted sentiment classifier, observing significant improvements on various benchmark datasets. We obtain the best accuracies on all test sets.
Yuze Gao, Yue Zhang 0004, Tong Xiao 0001
IJCNLP(1)3
2016 Syntactic Skeleton-Based Translation
abstract
In this paper we propose an approach to modeling syntactically-motivated skeletal structure of source sentence for machine translation. This model allows for application of high-level syntactic transfer rules and low-level non-syntactic rules. It thus involves fully syntactic, non-syntactic, and partially syntactic derivations via a single grammar and decoding paradigm. On large-scale Chinese-English and English-Chinese translation tasks, we obtain an average improvement of +0.9 BLEU across the newswire and web genres.
Tong Xiao 0001, Chunliang Zhang, Tongran Liu
AAAI1
2016 A Loss-Augmented Approach to Training Syntactic Machine Translation Systems
abstract
Current syntactic machine translation (MT) systems implicitly use beam-width unlimited search in learning model parameters (e.g., feature values for each translation rule). However, a limited beam-width has to be adopted in decoding new sentences, and the MT output is in general evaluated by various metrics, such as BLEU and TER. In this paper, we address: 1) the mismatch of adopted beam-widths between training and decoding; and 2) the mismatch of training criteria and MT evaluation metrics. Unlike previous work, we model the two problems in a single training paradigm simultaneously. We design a loss-augmented approach that explicitly considers the limited beam-width and evaluation metric in training, and present a simple but effective method to learn the model. By using beam search and BLEU-related losses, our approach improves a state-of-the-art syntactic MT system by +1.0 BLEU on Chinese-to-English and English-to-Chinese translation tasks. It even outperforms seven previous training approaches over 0.8 BLEU points. More interestingly, promising improvements are observed when our approach works with TER.
Tong Xiao 0001, Derek F. Wong
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Improving syntactic rule extraction through deleting spurious links with translation span alignment
abstract
Abstract Most statistical machine translation systems typically rely on word alignments to extract translation rules. This approach would suffer from a practical problem that even one spurious word alignment link can prevent some desirable translation rules from being extracted. To address this issue, this paper presents two approaches, referred to as sub-tree alignment and phrase-based forced decoding methods, to automatically learn translation span alignments from parallel data. Then, we improve the translation rule extraction by deleting spurious links and inserting new links based on bilingual translation span correspondences. Some comparison experiments are designed to demonstrate the effectiveness of the proposed approaches.
Qiang Li 0022, Tong Xiao 0001
Nat. Lang. Eng.3
2014 Effective Incorporation of Source Syntax into Hierarchical Phrase-based Translation
Tong Xiao 0001, Adrià de Gispert, William J. Byrne
COLING1
2013 Bagging and Boosting statistical machine translation systems
Tong Xiao 0001, Tongran Liu
Artif. Intell.1
2013 Unsupervised Sub-tree Alignment for Tree-to-Tree Translation
abstract
This article presents a probabilistic sub-tree alignment model and its application to tree-to-tree machine translation. Unlike previous work, we do not resort to surface heuristics or expensive annotated data, but instead derive an unsupervised model to infer the syntactic correspondence between two languages. More importantly, the developed model is syntactically-motivated and does not rely on word alignments. As a by-product, our model outputs a sub-tree alignment matrix encoding a large number of diverse alignments between syntactic structures, from which machine translation systems can efficiently extract translation rules that are often filtered out due to the errors in 1-best alignment. Experimental results show that the proposed approach outperforms three state-of-the-art baseline approaches in both alignment accuracy and grammar quality. When applied to machine translation, our approach yields a +1.0 BLEU improvement and a -0.9 TER reduction on the NIST machine translation evaluation corpora. With tree binarization and fuzzy decoding, it even outperforms a state-of-the-art hierarchical phrase-based system.
Tong Xiao 0001
J. Artif. Intell. Res.1
2012 Easy-First Chinese POS Tagging and Dependency Parsing
Tong Xiao 0001, Feiliang Ren
COLING2
2011 Document-level Consistency Verification in Machine Translation
Tong Xiao 0001, Shujie Yao
MTSummit1
2011 Language Modeling for Syntax-Based Machine Translation Using Tree Substitution Grammars: A Case Study on Chinese-English Translation
abstract
The poor grammatical output of Machine Translation (MT) systems appeals syntax-based approaches within language modeling. However, previous studies showed that syntax-based language modeling using (Context-Free) Treebank Grammars was not very helpful in improving BLEU scores for Chinese-English machine translation. In this article we further study this issue in the context of Chinese-English syntax-based Statistical Machine Translation (SMT) where Synchronous Tree Substitution Grammars (STSGs) are utilized to model the translation process. In particular, we develop a Tree Substitution Grammar-based language model for syntax-based MT, and present three methods to efficiently integrate the proposed language model into MT decoding. In addition, we design a simple and effective method to adapt syntax-based language models for MT tasks. We demonstrate that the proposed methods are able to benefit a state-of-the-art syntax-based MT system. On the NIST Chinese-English MT evaluation corpora, we finally achieve an improvement of 0.6 BLEU points over the baseline.
Tong Xiao 0001, Muhua Zhu
ACM Trans. Asian Lang. Inf. Process.1
2011 Automatic Treebank Conversion via Informed Decoding - A Case Study on Chinese Treebanks
abstract
Treebanks are valuable resources for syntactic parsing. For some languages such as Chinese, we can obtain multiple constituency treebanks which are developed by different organizations. However, due to discrepancies of underlying annotation standards, such treebanks in general cannot be used together through direct data combination. To enlarge training data for syntactic parsing, we focus in this article on the challenge of unifying standards of disparate treebanks by automatically converting one treebank (source treebank) to fit a different standard which is exhibited by another treebank (target treebank). We propose to convert a treebank in two sequential steps which correspond to the part-of-speech level and syntactic structure level (including tree structures and grammar labels), respectively. Approaches used in both levels can be unified as an informed decoding procedure, where information derived from original annotation in a source treebank is used to guide the conversion conducted by a POS tagger (or a parser in the syntactic structure level) trained on a target treebank. We take two Chinese treebanks as a case study, and experiments on these two treebanks show significant improvements in conversion accuracy over baseline systems, especially in situations where a target treebank is small in size.
Muhua Zhu, Tong Xiao 0001
ACM Trans. Asian Lang. Inf. Process.3
2010 Boosting-Based System Combination for Machine Translation
Tong Xiao 0001, Muhua Zhu, Huizhen Wang
ACL1
2010 Heterogeneous Parsing via Collaborative Decoding
Muhua Zhu, Tong Xiao 0001
COLING3
2009 The Feature Subspace Method for SMT System Combination
Nan Duan 0001, Mu Li 0001, Tong Xiao 0001, Ming Zhou 0001
EMNLP3
2009 Better Synchronous Binarization for Machine Translation
Tong Xiao 0001, Mu Li 0001, Dongdong Zhang 0001, Ming Zhou 0001
EMNLP1