VLDB 2026 Research / reviewers in the wild / expert
Deyi Xiong
dblp:55/6548
· DBLP profile ↗
187ranked-venue papers
21as first author
97since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 173 · 21 first-author · 87 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 2 first-author · 16 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMsabstractLarge Language Models (LLMs) frequently exhibit strong translation abilities, even without task-specific fine-tuning. However, the internal mechanisms governing this innate capability remain largely opaque. To demystify this process, we leverage Sparse Autoencoders (SAEs) and introduce a novel framework for identifying task-specific features. Our method first recalls features that are frequently co-activated on translation inputs and then filters them for functional coherence using a PCA-based consistency metric. This framework successfully isolates a small set of "translation initiation" features. Causal interventions demonstrate that amplifying these features steers the model towards correct translation, while ablating them induces hallucinations and off-task outputs, confirming they represent a core component of the model's innate translation competency. Moving from analysis to application, we leverage this mechanistic insight to propose a new data selection strategy for efficient fine-tuning. Specifically, we prioritize training on "mechanistically hard" samples—those that fail to naturally activate the translation initiation features. Experiments show this approach significantly improves data efficiency and suppresses hallucinations. Furthermore, we find these mechanisms are transferable to larger models of the same family. Our work not only decodes a core component of the translation mechanism in LLMs but also provides a blueprint for using internal model mechanism to create more robust and efficient models. Xinwei Wu 0001, Yuqi Ren, Linlong Xu, Longyue Wang, Deyi Xiong, Weihua Luo, Kaifu Zhang |
AAAI | 7 |
| 2026 | AdaDPI: Document-level Translation Adaptive Agent via Dynamic Parametric InternalizationabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities in machine translation.However, maintaining discourse coherence and terminological consistency remains a persistent challenge in documentlevel translation (DocMT).Existing solutions, such as memory-based agents, predominantly rely on explicit context concatenation.This paradigm treats historical context as a static external resource, which often leads to context dilution, high inference latency, and superficial knowledge integration.To address these limitations, we propose AdaDPI, an adaptive agentic framework that shifts the DocMT paradigm from static retrieval to dynamic parametric internalization.Specifically, we design a linguistic uncertainty monitor (LUM) to actively detect critical discourse discontinuities by the model's epistemic uncertainty.Upon detection, a context-to-parameter integrator (CPI) compiles retrieved external constraints directly into the model's intrinsic state via an online parameter adaptation mechanism.Through the online parameter adaptation on a lightweight adapter, AdaDPI internalizes document-specific norms into the model's intrinsic representations, enabling a progressive evolution of the translation strategy as the discourse unfolds.Extensive experiments on the discourse-rich GuoFeng and IWSLT2017 datasets demonstrate that AdaDPI significantly outperforms the SoTA baselines by more than 5 points on the consistency metric. Hong Ren, Liting Deng, Shaolin Zhu, Deyi Xiong |
ACL (1) | 4 |
| 2026 | Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language ModelsabstractDan Shi, Zhuowen Han, Simon Ostermann, Renren Jin, Josef Van Genabith, Deyi Xiong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Dan Shi 0001, Zhuowen Han, Simon Ostermann 0002, Renren Jin, Josef van Genabith, Deyi Xiong |
ACL (1) | 6 |
| 2026 | From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language ModelsabstractLing Shi, Xinwei Wu, Xiaohu Zhao, Hao Wang, Heng Liu, Yangyang Liu, Linlong Xu, Longyue Wang, Deyi Xiong, Weihua Luo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ling Shi 0004, Xinwei Wu 0001, Linlong Xu, Longyue Wang, Deyi Xiong, Weihua Luo |
ACL (1) | 9 |
| 2026 | EvoSci: A Bio-Inspired Multi-Agent Framework for the Evolution of Scientific DiscoveryabstractLarge language models (LLMs), have shown strong potential in scientific discovery, yet existing methods still face substantial challenges in the design of research workflows and multi-role collaboration mechanisms.To mitigate these issues, we propose EvoSci, a multi-agent scientific collaboration framework, which integrates bio-inspired evolution with knowledge graph modeling.To iteratively generate, evaluate, and refine research ideas, EvoSci incorporates multiple role-based agents, including mentor, researcher, and reviewer.By combining collaborative reasoning, shared memory, and evolutionary feedback, EvoSci significantly enhances the coherence and creativity of scientific exploration.Experiments on real-world research topics demonstrate that EvoSci significantly outperforms strong baselines in LLM-based structured peer-review and comparative ranking evaluations, achieving the highest overall peer-review score (ICLR 4.90) and top ranking (Top-10 = 54).These results suggest its superiority in both scientific idea generation and continuous discovery. Xiaoyu Xiong, Yuqi Ren, Deyi Xiong |
ACL (1) | 3 |
| 2026 | From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for TibetanabstractLei Yang, Leiyu Pan, Bojian Xiong, Renren Jin, Shaowei Zhang, Yue Chen, Ling Shi, Jiang Zhou, Junru Wu, Zhen Wang, Jianxiang Peng, Juesi Xiao, Tianyu Dong, Zhuowen Han, Zhuo Chen, Yuqi Ren, Deyi Xiong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Leiyu Pan, Bojian Xiong, Renren Jin, Ling Shi 0004, Jianxiang Peng, Juesi Xiao, Tianyu Dong, Zhuowen Han, Yuqi Ren, Deyi Xiong |
ACL (1) | 17 |
| 2026 | Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-SortsabstractLarge Language Models (LLMs) are increasingly deployed in contexts requiring complex moral reasoning and value trade-offs.However, existing evaluations typically rely on item-level behavioral metrics, which fail to capture how models structurally prioritize competing values as a cohesive system.To address this, we propose a symmetric human-LLM evaluation framework, grounded in Q methodology, to measure value-structure alignment.Under our protocol, humans and models sort an identical 140-item moral statement set into a shared nine-column forced distribution; for LLMs, we elicit strict rankings and deterministically map them to Q-sort buckets.Using a human reference sample (N = 35), we establish a stable three-factor reference geometry specific to this instrument and sample.We evaluate 12 LLMs across four model families via 240 replicated Q-sorts at two temperature settings, quantifying structural alignment via Procrustes similarity (ϕ) and RSA-based Spearman correlation (ρ).Our results reveal significant crossfamily heterogeneity, model-specific sensitivity to generation stochasticity and localized misalignment, which demonstrate that favorable global scores can obscure underlying regional distortions.While rank-and bucket-based analyses remain highly consistent, prompt phrasing introduces notable variance.Ultimately, assessing value-structure alignment provides a crucial structural complement to traditional itemwise moral benchmarks. Jingting Zheng, Yuqi Ren, Linhao Yu, Yongqi Leng, Deyi Xiong |
ACL (1) | 5 |
| 2026 | Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity TranslationabstractJiang Zhou, Xiaohu Zhao, Xinwei Wu, Tianyu Dong, Hao Wang, Yangyang Liu, Heng Liu, Linlong Xu, Longyue Wang, Weihua Luo, Deyi Xiong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xinwei Wu 0001, Tianyu Dong, Linlong Xu, Longyue Wang, Weihua Luo, Deyi Xiong |
ACL (1) | 11 |
| 2026 | DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value MappingabstractCurrent Large Language Models (LLMs) typically rely on coarse-grained national labels for pluralistic value alignment.However, such macro-level supervision often obscures intra-country value heterogeneity, yielding a loose alignment.We argue that resolving this limitation requires shifting from national labels to multi-dimensional demographic constraints, which can identify groups with predictable, high-consensus value preference.To this end, we propose DVMap (High-Consensus Demographic-Value Mapping), a framework for fine-grained pluralistic value alignment.In this framework, we first present a demographic archetype extraction strategy to construct a high-quality value alignment corpus of 56,152 samples from the World Values Survey (WVS) by strictly retaining respondents with consistent value preferences under identical demographics.Over this corpus, we introduce a Structured Chain-of-Thought (CoT) mechanism that explicitly guides LLMs to reason about demographic-value correlations.Subsequently, we employ Group Relative Policy Optimization (GRPO) to achieve adaptive anchoring of value distributions.To rigorously evaluate generalization, we further establish a triple-generalization benchmark (spanning cross-demographic, cross-country, and crossvalue) comprising 21,553 samples.Experimental results demonstrate that DVMap effectively learns the manifold mapping from demographics to values, exhibiting strong generalization and robustness.On cross-demographic tests, Qwen3-8B-DVMap achieves 48.6% accuracy, surpassing the advanced open-source LLM DeepSeek-v3.2(45.1%).The source code and dataset are available at https://github. com/EnlightenedAI/DVMap. Pengyun Zhu, Yuqi Ren, Deyi Xiong |
ACL (1) | 5 |
| 2026 | APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and InterpretationabstractPengyun Zhu, Qiheng Sun, Long Wen, Yanbo Wang, Yang Cao, Junxu Liu, Deyi Xiong, Jinfei Liu, Zhibo Wang, Kui Ren. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pengyun Zhu, Qiheng Sun, Yang Cao 0011, Junxu Liu, Deyi Xiong, Jinfei Liu, Zhibo Wang 0001, Kui Ren 0001 |
ACL (1) | 7 |
| 2026 | AMART: A multi-agent reflective framework for detecting and correcting faithfulness errors in translation
Shaolin Zhu, Deyi Xiong |
Inf. Process. Manag. | 3 |
| 2025 | MLAS-LoRA: Language-Aware Parameters Detection and LoRA-Based Knowledge Transfer for Multilingual Machine TranslationabstractLarge language models (LLMs) have achieved remarkable progress in multilingual machine translation (MT), demonstrating strong performance even with limited parallel data.However, effectively fine-tuning LLMs for MT is challenging due to parameter interference, which arises from the conflicting demands of different language pairs and the risk of overwriting pre-trained knowledge.To address this issue, we propose MLAS-LoRA, a novel multiple language-aware LoRA knowledge transfer framework.MLAS-LoRA efficiently adapts LLMs to MT by selectively transferring knowledge from a large teacher to a small student model.Our approach first evaluates the awareness of neurons and extracts linguistic knowledge in the teacher model to both the general MT task and specific language pairs.We then propose a multiple language-specific LoRA architecture to inject the extracted knowledge into the student model.During fine-tuning, only the parameters of the relevant languagegeneral and language-specific LoRA modules are updated.Experimental results on diverse multilingual language pairs demonstrate that MLAS-LoRA significantly outperforms strong baselines by +1.7 BLEU on average, including standard fine-tuning and other parameterefficient methods. Tianyu Dong, Bo Li 0131, Shaolin Zhu, Deyi Xiong |
ACL (1) | 5 |
| 2025 | Praetor: A Fine-Grained Generative LLM Evaluator with Instance-Level Customizable Evaluation CriteriaabstractYongqi Leng, Renren Jin, Yue Chen, Zhuowen Han, Ling Shi, Jianxiang Peng, Lei Yang, Juesi Xiao, Deyi Xiong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yongqi Leng, Renren Jin, Zhuowen Han, Ling Shi 0004, Jianxiang Peng, Juesi Xiao, Deyi Xiong |
ACL (1) | 9 |
| 2025 | ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue AgentsabstractZhigen Li, Jianxiang Peng, Yanmeng Wang, Yong Cao, Tianhao Shen, Minghui Zhang, Linxi Su, Shang Wu, Yihang Wu, YuQian Wang, Ye Wang, Wei Hu, Jianfeng Li, Shaojun Wang, Jing Xiao, Deyi Xiong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhigen Li, Jianxiang Peng, Yanmeng Wang, Tianhao Shen, Linxi Su, Yihang Wu, Jing Xiao 0006, Deyi Xiong |
ACL (1) | 16 |
| 2025 | CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language ModelsabstractLarge language models (LLMs) are possessed of numerous beneficial capabilities, yet their potential inclination harbors unpredictable risks that may materialize in the future.We hence propose CRiskEval, a Chinese dataset meticulously designed for gauging the risk proclivities inherent in LLMs such as resource acquisition and malicious coordination, as part of efforts for proactive preparedness.To curate CRiskEval, we define a new risk taxonomy with 7 types of frontier risks and 4 safety levels, including extremely hazardous,moderately hazardous, neutral and safe.We follow the philosophy of tendency evaluation to empirically measure the stated "desire" of LLMs via fine-grained multiple-choice question answering.The dataset consists of 14,888 questions that simulate scenarios related to predefined 7 types of frontier risks.Each question is accompanied with 4 answer choices that state opinions or behavioral tendencies corresponding to the question.All answer choices are manually annotated with one of the defined risk levels so that we can easily build a fine-grained frontier risk profile for each assessed LLM.Extensive evaluation with CRiskEval on a spectrum of prevalent Chinese LLMs has unveiled a striking revelation: most models exhibit risk tendencies of more than 40% (weighted tendency to the four risk levels).Furthermore, a subtle increase in the model's inclination toward urgent selfsustainability, power seeking and other dangerous goals becomes evident as the size of models increases.To promote further research on the frontier risk evaluation of LLMs, we publicly release our dataset at https://github.com/ tjunlp-lab/CRiskEval. Warning: This paper contains model outputs which are offensive in nature. Ling Shi 0004, Deyi Xiong |
ACL (1) | 2 |
| 2025 | Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree SearchabstractVideo captioning can be used to assess the video understanding capabilities of Multimodal Large Language Models (MLLMs).However, existing benchmarks and evaluation protocols suffer from crucial issues, such as inadequate or homogeneous creation of key points, exorbitant cost of data creation, and limited evaluation scopes. To address these issues, we propose an automatic framework, named AutoCaption, which leverages Monte Carlo Tree Search (MCTS) to construct numerous and diverse descriptive sentences (i.e., key points) that thoroughly represent video content in an iterative way. This iterative captioning strategy enables the continuous enhancement of video details such as actions, objects’ attributes, environment details, etc. We apply AutoCaption to curate MCTS-VCB, a fine-grained video caption benchmark covering video details, thereby enabling a comprehensive evaluation of MLLMs on the video captioning task. We evaluate more than 20 open- and closed-source MLLMs of varying sizes on MCTS-VCB. Results show that MCTS-VCB can effectively and comprehensively evaluate the video captioning capability, with Gemini-1.5-Pro achieving the highest F1 score of 71.2. Interestingly, we fine-tune InternVL2.5-8B with the AutoCaption-generated data, which helps the model achieve an overall improvement of 25.0% on MCTS-VCB and 16.3% on DREAM-1K, further demonstrating the effectiveness of AutoCaption. The code and data are available at https://github.com/tjunlp-lab/MCTS-VCB. Linhao Yu, Xingguang Ji, Fanheng Kong, Victoria W., Deyi Xiong |
ACL (1) | 10 |
| 2025 | CONTRANS: Weak-to-Strong Alignment Engineering via Concept TransplantationabstractEnsuring large language models (LLM) behave consistently with human goals, values, and intentions is crucial for their safety but yet computationally expensive. To reduce the computational cost of alignment training of LLMs, especially for those with a huge number of parameters, and to reutilize learned value alignment, we propose ConTrans, a novel framework that enables weak-to-strong alignment transfer via concept transplantation. From the perspective of representation engineering, ConTrans refines concept vectors in value alignment from a source LLM (usually a weak yet aligned LLM). The refined concept vectors are then reformulated to adapt to the target LLM (usually a strong yet unaligned base LLM) via affine transformation. In the third step, ConTrans transplants the reformulated concept vectors into the residual stream of the target LLM. Experiments demonstrate the successful transplantation of a wide range of aligned concepts from 7B models to 13B and 70B models across multiple LLMs and LLM families. Remarkably, ConTrans even surpasses instruction-tuned models in terms of truthfulness. Experiment results validate the effectiveness of both inter-LLM-family and intra-LLM-family concept transplantation. Our work successfully demonstrates an alternative way to achieve weak-to-strong alignment generalization and control. Weilong Dong, Xinwei Wu 0001, Renren Jin, Shaoyang Xu, Deyi Xiong |
COLING | 5 |
| 2025 | Automated Progressive Red TeamingabstractEnsuring the safety of large language models (LLMs) is paramount, yet identifying potential vulnerabilities is challenging. While manual red teaming is effective, it is time-consuming, costly and lacks scalability. Automated red teaming (ART) offers a more cost-effective alternative, automatically generating adversarial prompts to expose LLM vulnerabilities. However, in current ART efforts, a robust framework is absent, which explicitly frames red teaming as an effectively learnable task. To address this gap, we propose Automated Progressive Red Teaming (APRT) as an effectively learnable framework. APRT leverages three core modules: an Intention Expanding LLM that generates diverse initial attack samples, an Intention Hiding LLM that crafts deceptive prompts, and an Evil Maker to manage prompt diversity and filter ineffective samples. The three modules collectively and progressively explore and exploit LLM vulnerabilities through multi-round interactions. In addition to the framework, we further propose a novel indicator, Attack Effectiveness Rate (AER) to mitigate the limitations of existing evaluation metrics. By measuring the likelihood of eliciting unsafe but seemingly helpful responses, AER aligns closely with human evaluations. Extensive experiments with both automatic and human evaluations, demonstrate the effectiveness of APRT across both open- and closed-source LLMs. Specifically, APRT effectively elicits 54% unsafe yet useful responses from Meta’s Llama-3-8B-Instruct, 50% from GPT-4o (API access), and 39% from Claude-3.5 (API access), showcasing its robust attack capability and transferability across LLMs (especially from open-source LLMs to closed-source LLMs). The code and seed data are available at https://github.com/tjunlp-lab/APRT. Bojian Jiang, Yi Jing, Tianhao Shen, Deyi Xiong |
COLING | 5 |
| 2025 | Towards Understanding Multi-Task Learning (Generalization) of LLMs via Detecting and Exploring Task-Specific NeuronsabstractWhile large language models (LLMs) have demonstrated superior multi-task capabilities, understanding the learning mechanisms behind this is still a challenging problem. In this paper, we attempt to understand such mechanisms from the perspective of neurons. Specifically, we detect task-sensitive neurons in LLMs via gradient attribution on task-specific data. Through extensive deactivation and fine-tuning experiments, we demonstrate that the detected neurons are highly correlated with the given task, which we term as task-specific neurons. With these identified task-specific neurons, we delve into two common problems in multi-task learning and continuous learning: Generalization and Catastrophic Forgetting. We find that the overlap of task-specific neurons is strongly associated with generalization and specialization across tasks. Interestingly, at certain layers of LLMs, there is a high similarity in the parameters of different task-specific neurons, and such similarity is highly correlated with the generalization performance. Inspired by these findings, we propose a neuron-level continuous fine-tuning method that only fine-tunes the current task-specific neurons during continuous learning, and extensive experiments demonstrate the effectiveness of the proposed method. Our study provides insights into the interpretability of LLMs in multi-task learning. Yongqi Leng, Deyi Xiong |
COLING | 2 |
| 2025 | Empirical Study on Data Attributes Insufficiency of Evaluation Benchmarks for LLMsabstractPrevious benchmarks for evaluating large language models (LLMs) have primarily emphasized quantitative metrics, such as data volume. However, this focus may neglect key qualitative data attributes that can significantly impact the final rankings of LLMs, resulting in unreliable leaderboards. In this paper, we investigate whether current LLM benchmarks adequately consider these data attributes. We specifically examine three attributes: diversity, redundancy, and difficulty. To explore these attributes, we propose a framework with three separate modules, each designed to assess one of the attributes. Using a method that progressively incorporates these attributes, we analyze their influence on the benchmark. Our experimental results reveal a meaningful correlation between LLM rankings on the revised benchmark and the original benchmark when these attributes are accounted for. These findings indicate that existing benchmarks often fail to meet all three criteria, highlighting a lack of consideration for multifaceted data attributes in current evaluation datasets. Chuang Liu 0009, Renren Jin, Mark Steedman, Deyi Xiong |
COLING | 7 |
| 2025 | Do Large Language Models Mirror Cognitive Language Processing?abstractLarge Language Models (LLMs) have demonstrated remarkable abilities in text comprehension and logical reasoning, indicating that the text representations learned by LLMs can facilitate their language processing capabilities. In neuroscience, brain cognitive processing signals are typically utilized to study human language processing. Therefore, it is natural to ask how well the text embeddings from LLMs align with the brain cognitive processing signals, and how training strategies affect the LLM-brain alignment? In this paper, we employ Representational Similarity Analysis (RSA) to measure the alignment between 23 mainstream LLMs and fMRI signals of the brain to evaluate how effectively LLMs simulate cognitive language processing. We empirically investigate the impact of various factors (e.g., pre-training data size, model scaling, alignment training, and prompts) on such LLM-brain alignment. Experimental results indicate that pre-training data size and model scaling are positively correlated with LLM-brain similarity, and alignment training can significantly improve LLM-brain similarity. Explicit prompts contribute to the consistency of LLMs with brain cognitive language processing, while nonsensical noisy prompts may attenuate such alignment. Additionally, the performance of a wide range of LLM evaluations (e.g., MMLU, Chatbot Arena) is highly correlated with the LLM-brain similarity. Yuqi Ren, Renren Jin, Tongxuan Zhang, Deyi Xiong |
COLING | 4 |
| 2025 | Towards a Unified Paradigm of Concept Editing in Large Language ModelsabstractConcept editing aims to control specific concepts in large language models (LLMs) and is an emerging subfield of model editing.Despite the emergence of various editing methods in recent years, there remains a lack of rigorous theoretical analysis and a unified perspective to systematically understand and compare these methods.To address this gap, we propose a unified paradigm for concept editing methods, in which all forms of conceptual injection are aligned at the neuron level.We study four representative concept editing methods: Neuron Editing (NE), Supervised Fine-tuning (SFT), Sparse Autoencoder (SAE), and Steering Vector (SV).Then we categorize them into two classes based on their mode of conceptual information injection: indirect (NE, SFT) and direct (SAE, SV).We evaluate above methods along four dimensions: editing reliability, output generalization, neuron level consistency, and mathematical formalization.Experiments show that SAE achieves the best editing reliability.In output generalization, SAE captures features closer to human-understood concepts, while NE tends to locate text patterns rather than true semantics.Neuron-level analysis reveals that direct methods share high neuron overlap, as do indirect methods, indicating methodological commonality within each category.Our unified paradigm offers a clear framework and valuable insights for advancing interpretability and controlled generation in LLMs. Zhuowen Han, Xinwei Wu 0001, Dan Shi 0001, Renren Jin, Deyi Xiong |
EMNLP | 5 |
| 2025 | Towards Optimal Evaluation Efficiency for Large Language ModelsabstractComprehensive evaluation of large language models (LLMs) typically requires large-scale benchmarks, which is costly in terms of both data annotation and computational resource needed for evaluation.To mitigate these challenges, we propose an efficient evaluation framework that selects a question subset based on pre-tested results, thereby reducing the costs.We formulate the subset selection problem as an optimization task, solved using optimal random sampling and simulated annealing algorithms.We compare our approach with prior clustering-based methods and assess their reliability in terms of score accuracy.Additionally, we perform semantic analysis and evaluate whether the selected subsets preserve the semantic information of the original benchmark using Wasserstein distance.Experimental results show that our method outperforms previous approaches in terms of reliability, as measured by L2 norm.Our study provides an optimized perspective for balancing evaluation efficiency and reliability in LLM assessments, while revealing the relationship between optimization methods and semantic retention. Guohong Li, Deyi Xiong |
EMNLP | 2 |
| 2025 | DiplomacyAgent: Do LLMs Balance Interests and Ethical Principles in International Events?abstractThe widespread deployment of large language models (LLMs) across various domains has made their safety a critical priority.Inspired by think-tank decision-making philosophy, we propose DiplomacyAgent, an LLM-based multiagent system for diplomatic position analysis.With DiplomacyAgent, we are able to systematically assess how LLMs balance "interests" against "ethical principles" when addressing various international events, hence understanding the safety implications of LLMs in diplomacy.Specifically, this will help to assess the consistency of LLM stance with widely recognized ethical standards, as well as the potential risks or ideological biases that may arise.Through integrated quantitative metrics, our research uncovers unexpected decision-making patterns in LLM responses to sensitive issues including human rights protection, environmental sustainability, regional conflicts, etc.It discloses that LLMs could exhibit a strong bias towards interests, leading to unsafe decisions that violate ethical and moral principles.Our experiment results suggest that deploying LLMs in high-stakes domains, particularly in the formulation of diplomatic policies, necessitates a comprehensive assessment of potential ethical and social implications, as well as the implementation of stringent safety protocols. Jianxiang Peng, Ling Shi 0004, Xinwei Wu 0001, Fujiang Liu, Haocheng Lyu, Deyi Xiong |
EMNLP | 7 |
| 2025 | DCIS: Efficient Length Extrapolation of LLMs via Divide-and-Conquer Scaling Factor SearchabstractLarge language models (LLMs) based on the Transformer architecture usually have their context length limited due to the high training cost.Recent advancements extend the context window by adjusting the scaling factors of RoPE and fine-tuning.However, suboptimal initialization of these factors results in increased fine-tuning costs and reduced performance at target length.To address these challenges, we propose a novel RoPE-based fine-tuning framework that diverges from conventional scaling factors search.Specifically, we present a Divide-and-Conquer Incremental Search (DCIS) algorithm that strategically determines the better scaling factors.Further finetuning with the identified scaling factors effectively extends the context window of LLMs.Empirical results demonstrate that our methodology not only mitigates performance decay at extended target lengths but also allows the model to fine-tune on short contexts and generalize to long contexts, thereby reducing the cost of fine-tuning.The scaling factors obtained through DCIS can even perform effectively without fine-tuning.Further analysis of the search space reveals that DCIS achieves twice the search efficiency compared to other methods.We also examine the impact of the non-strictly increasing scaling factors utilized in DCIS and evaluate the general capabilities of LLMs across various context lengths. Shaoyang Xu, Jianxiang Peng, Shaolin Zhu, Deyi Xiong |
EMNLP | 5 |
| 2025 | Evaluating and Improving Graph to Text Generation with Large Language ModelsabstractJie He, Yijun Yang, Wanqiu Long, Deyi Xiong, Victor Gutierrez Basulto, Jeff Z. Pan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jie He 0004, Wanqiu Long, Deyi Xiong, Víctor Gutiérrez-Basulto, Jeff Z. Pan |
NAACL (Long Papers) | 4 |
| 2025 | Self-Pluralising Culture Alignment for Large Language ModelsabstractShaoyang Xu, Yongqi Leng, Linhao Yu, Deyi Xiong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shaoyang Xu, Yongqi Leng, Linhao Yu, Deyi Xiong |
NAACL (Long Papers) | 4 |
| 2025 | Very-Long-Distance Dependency Capturing Evaluation via Language Modeling Based on Gender Consistency
Hongfei Xu, Zhuofei Liang, Josef van Genabith, Deyi Xiong, Hongying Zan, Qiuhui Liu, Tengxun Zhang |
NLPCC (4) | 4 |
| 2025 | Improving blind face restoration by utilizing edge semantic enhancement
Xiaodong Qian, Jianglin Wang, Deyi Xiong |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Overcoming language barriers via machine translation with sparse Mixture-of-Experts fusion of large language models
Shaolin Zhu, Leiyu Pan, Dong Jian, Deyi Xiong |
Inf. Process. Manag. | 4 |
| 2025 | FineMath: A Fine-Grained Mathematical Evaluation Benchmark for Chinese Large Language ModelsabstractTo thoroughly assess the mathematical reasoning abilities of Large Language Models (LLMs), we need to carefully curate evaluation datasets covering diverse mathematical concepts and mathematical problems at different difficulty levels. In pursuit of this objective, we propose FineMath in this article, a fine-grained mathematical evaluation benchmark dataset for assessing Chinese LLMs. FineMath is created to cover the major key mathematical concepts taught in elementary school math, which are further divided into 17 categories of math word problems, enabling in-depth analysis of mathematical reasoning abilities of LLMs. All the 17 categories of math word problems are manually annotated with their difficulty levels according to the number of reasoning steps required to solve these problems. We conduct extensive experiments on a wide range of LLMs on FineMath and find that there is still considerable room for improvements in terms of mathematical reasoning capability of Chinese LLMs. We also carry out an in-depth analysis on the evaluation process and methods that have been overlooked previously. These two factors significantly influence the model results and our understanding of their mathematical reasoning capabilities. Our data is available at https://github.com/tjunlp-lab/FineMATH . Renren Jin, Zheng Yao 0004, Deyi Xiong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2024 | Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark RemedyabstractTo mitigate potential risks associated with language models (LMs), recent AI detection research proposes incorporating watermarks into machine-generated text through random vocabulary restrictions and utilizing this information for detection. In this paper, we show that watermarking algorithms designed for LMs cannot be seamlessly applied to conditional text generation (CTG) tasks without a notable decline in downstream task performance. To address this issue, we introduce a simple yet effective semantic-aware watermarking algorithm that considers the characteristics of conditional text generation with the input context. Compared to the baseline watermarks, our proposed watermark yields significant improvements in both automatic and human evaluations across various text generation models, including BART and Flan-T5, for CTG tasks such as summarization and data-to-text generation. Meanwhile, it maintains detection ability with higher z-scores but lower AUC scores, suggesting the presence of a detection paradox that poses additional challenges for watermarking CTG. Yu Fu 0009, Deyi Xiong, Yue Dong 0002 |
AAAI | 2 |
| 2024 | CORECODE: A Common Sense Annotated Dialogue Dataset with Benchmark Tasks for Chinese Large Language ModelsabstractAs an indispensable ingredient of intelligence, commonsense reasoning is crucial for large language models (LLMs) in real-world scenarios. In this paper, we propose CORECODE, a dataset that contains abundant commonsense knowledge manually annotated on dyadic dialogues, to evaluate the commonsense reasoning and commonsense conflict detection capabilities of Chinese LLMs. We categorize commonsense knowledge in everyday conversations into three dimensions: entity, event, and social interaction. For easy and consistent annotation, we standardize the form of commonsense knowledge annotation in open-domain dialogues as "domain: slot = value". A total of 9 domains and 37 slots are defined to capture diverse commonsense knowledge. With these pre-defined domains and slots, we collect 76,787 commonsense knowledge annotations from 19,700 dialogues through crowdsourcing. To evaluate and enhance the commonsense reasoning capability for LLMs on the curated dataset, we establish a series of dialogue-level reasoning and detection tasks, including commonsense knowledge filling, commonsense knowledge generation, commonsense conflict phrase detection, domain identification, slot identification, and event causal inference. A wide variety of existing open-source Chinese LLMs are evaluated with these tasks on our dataset. Experimental results demonstrate that these models are not competent to predict CORECODE's plentiful reasoning content, and even ChatGPT could only achieve 0.275 and 0.084 accuracy on the domain identification and slot identification tasks under the zero-shot setting. We release the data and codes of CORECODE at https://github.com/danshi777/CORECODE to promote commonsense reasoning evaluation and study of LLMs in the context of daily conversations. Dan Shi 0001, Chaobin You, Jiantao Huang, Taihao Li, Deyi Xiong |
AAAI | 5 |
| 2024 | LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine TranslationabstractRecent advancements in large language models (LLMs) have shown promising results in multilingual translation even with limited bilingual supervision.The major challenges are catastrophic forgetting and parameter interference 1 for finetuning LLMs when provided parallel training data.To address these challenges, we propose LANDeRMT, a Language-Aware Neuron Detecting and Routing framework that selectively finetunes LLMs to Machine Translation with diverse translation training data.In LANDeRMT, we evaluate the awareness of neurons to MT tasks and categorize them into language-general and languagespecific neurons.This categorization enables selective parameter updates during finetuning, mitigating parameter interference and catastrophic forgetting issues.For the detected neurons, we further propose a conditional awareness-based routing mechanism to dynamically adjust language-general and languagespecific capacity within LLMs, guided by translation signals.Experimental results demonstrate that the proposed LANDeRMT is very effective in learning translation knowledge, significantly improving translation quality over various strong baselines for multiple language pairs. Shaolin Zhu, Leiyu Pan, Bo Li 0131, Deyi Xiong |
ACL (1) | 4 |
| 2024 | CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language ModelsabstractHolistically measuring societal biases of large language models is crucial for detecting and reducing ethical risks in highly capable AI models. In this work, we present a Chinese Bias Benchmark dataset that consists of over 100K questions jointly constructed by human experts and generative language models, covering stereotypes and societal biases in 14 social dimensions related to Chinese culture and values. The curation process contains 4 essential steps: bias identification, ambiguous context generation, AI-assisted disambiguous context generation, and manual review and recomposition. The testing instances in the dataset are automatically derived from 3K+ high-quality templates manually authored with stringent quality control. The dataset exhibits wide coverage and high diversity. Extensive experiments demonstrate the effectiveness of the dataset in evaluating model bias, with all 12 publicly available Chinese large language models exhibiting strong bias in certain categories. Additionally, we observe from our experiments that fine-tuned models could, to a certain extent, heed instructions and avoid generating harmful outputs, in the way of “moral self-correction”. Our dataset is available at https://anonymous.4open.science/r/CBBQ-B860/. Yufei Huang 0005, Deyi Xiong |
LREC/COLING | 2 |
| 2024 | IT2ACL Learning Easy-to-Hard Instructions via 2-Phase Automated Curriculum Learning for Large Language ModelsabstractInstruction tuning has demonstrated its superiority in unlocking the abilities of pre-trained large language models (LLMs), including their capability to respond to diverse human instructions and conduct complex reasoning. In order to further enhance the continuous learning capabilities of pre-trained LLMs, we explore the training process of instruction tuning through the lens of task sequences. We propose a 2-phase automated curriculum learning guided instruction tuning framework, IT2ACL that learns easy-to-hard instructions for LLMs in a self-adjusting dynamic manner. To facilitate curriculum learning from instructions, we propose a loss-driven progress signal for two-phase strategies: instruction prediction gain that decides the instruction level syllabus. Through comprehensive experiments on 70 Chinese datasets which have been grouped into 16 distinct task clusters, we demonstrate the effectiveness of our approach in eliciting latent ability in pre-trained LLMs and achieving superior performance across diverse tasks. Yufei Huang 0005, Deyi Xiong |
LREC/COLING | 2 |
| 2024 | LHMKE: A Large-scale Holistic Multi-subject Knowledge Evaluation Benchmark for Chinese Large Language ModelsabstractChinese Large Language Models (LLMs) have recently demonstrated impressive capabilities across various NLP benchmarks and real-world applications. However, the existing benchmarks for comprehensively evaluating these LLMs are still insufficient, particularly in terms of measuring knowledge that LLMs capture. Current datasets collect questions from Chinese examinations across different subjects and educational levels to address this issue. Yet, these benchmarks primarily focus on objective questions such as multiple-choice questions, leading to a lack of diversity in question types. To tackle this problem, we propose LHMKE, a Large-scale, Holistic, and Multi-subject Knowledge Evaluation benchmark in this paper. LHMKE is designed to provide a comprehensive evaluation of the knowledge acquisition capabilities of Chinese LLMs. It encompasses 10,465 questions across 75 tasks covering 30 subjects, ranging from primary school to professional certification exams. Notably, LHMKE includes both objective and subjective questions, offering a more holistic evaluation of the knowledge level of LLMs. We have assessed 11 Chinese LLMs under the zero-shot setting, which aligns with real examinations, and compared their performance across different subjects. We also conduct an in-depth analysis to check whether GPT-4 can automatically score subjective predictions. Our findings suggest that LHMKE is a challenging and advanced testbed for Chinese LLMs. Chuang Liu 0009, Renren Jin, Yuqi Ren, Deyi Xiong |
LREC/COLING | 4 |
| 2024 | Can Large Language Models Learn Translation Robustness from Noisy-Source In-context Demonstrations?abstractLarge language models (LLMs) have been used for machine translation. When provided with prompts and source sentences, LLMs can achieve impressive translation results. However, the robustness of these LLMs remains a significant challenge, as they often struggle to accurately translate sentences in the presence of noise, even when using similarity-based in-context learning methods. This work proposes a research scheme for studying machine translation robustness on LLMs, investigating whether LLMs can learn translation robustness from noisy-source demonstration examples. Through experiments on different models, languages, and noise types, we empirically demonstrate that LLMs can learn how to handle noise and translation methods from noisy-source demonstration examples, thereby improving their translation performance on noisy sentences. Furthermore, we find that increasing the noise ratio appropriately for the noisy-source demonstration examples can enhance the translation robustness of LLMs. Additionally, we also attempt to investigate scenarios where LLMs are more likely to learn translation robustness for mixed and specific types of noise. We find that the model’s performance varies across different noise settings. Leiyu Pan, Yongqi Leng, Deyi Xiong |
LREC/COLING | 3 |
| 2024 | Decoding at the Speed of Thought: Harnessing Parallel Decoding of Lexical Units for LLMsabstractLarge language models have demonstrated exceptional capability in natural language understanding and generation. However, their generation speed is limited by the inherently sequential nature of their decoding process, posing challenges for real-time applications. This paper introduces Lexical Unit Decoding (LUD), a novel decoding methodology implemented in a data-driven manner, accelerating the decoding process without sacrificing output quality. The core of our approach is the observation that a pre-trained language model can confidently predict multiple contiguous tokens, forming the basis for a lexical unit, in which these contiguous tokens could be decoded in parallel. Extensive experiments validate that our method substantially reduces decoding time while maintaining generation quality, i.e., 33% speed up on natural language generation with no quality loss, and 30% speed up on code generation with a negligible quality loss of 3%. Distinctively, LUD requires no auxiliary models and does not require changes to existing architectures. It can also be integrated with other decoding acceleration methods, thus achieving an even more pronounced inference efficiency boost. We posit that the foundational principles of LUD could define a new decoding paradigm for future language models, enhancing their applicability for a broader spectrum of applications. All codes are be publicly available at https://github.com/tjunlp-lab/Lexical-Unit-Decoding-LUD-. Zijia Lin, Zhongyuan Wang 0006, Chengru Song, Di Zhang 0026, Kun Gai, Deyi Xiong |
LREC/COLING | 11 |
| 2024 | An Empirical Study on the Robustness of Massively Multilingual Neural Machine TranslationabstractMassively multilingual neural machine translation (MMNMT) has been proven to enhance the translation quality of low-resource languages. In this paper, we empirically investigate the translation robustness of Indonesian-Chinese translation in the face of various naturally occurring noise. To assess this, we create a robustness evaluation benchmark dataset for Indonesian-Chinese translation. This dataset is automatically translated into Chinese using four NLLB-200 models of different sizes. We conduct both automatic and human evaluations. Our in-depth analysis reveal the correlations between translation error types and the types of noise present, how these correlations change across different model sizes, and the relationships between automatic evaluation indicators and human evaluation indicators. The dataset is publicly available at https://github.com/tjunlp-lab/ID-ZH-MTRobustEval. Supryadi, Leiyu Pan, Deyi Xiong |
LREC/COLING | 3 |
| 2024 | Rewiring the Transformer with Depth-Wise LSTMsabstractStacking non-linear layers allows deep neural networks to model complicated functions, and including residual connections in Transformer layers is beneficial for convergence and performance. However, residual connections may make the model “forget” distant layers and fail to fuse information from previous layers effectively. Selectively managing the representation aggregation of Transformer layers may lead to better performance. In this paper, we present a Transformer with depth-wise LSTMs connecting cascading Transformer layers and sub-layers. We show that layer normalization and feed-forward computation within a Transformer layer can be absorbed into depth-wise LSTMs connecting pure Transformer attention layers. Our experiments with the 6-layer Transformer show significant BLEU improvements in both WMT 14 English-German / French tasks and the OPUS-100 many-to-many multilingual NMT task, and our deep Transformer experiments demonstrate the effectiveness of depth-wise LSTM on the convergence and performance of deep Transformers. Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong |
LREC/COLING | 5 |
| 2024 | LFED: A Literary Fiction Evaluation Dataset for Large Language ModelsabstractThe rapid evolution of large language models (LLMs) has ushered in the need for comprehensive assessments of their performance across various dimensions. In this paper, we propose LFED, a Literary Fiction Evaluation Dataset, which aims to evaluate the capability of LLMs on the long fiction comprehension and reasoning. We collect 95 literary fictions that are either originally written in Chinese or translated into Chinese, covering a wide range of topics across several centuries. We define a question taxonomy with 8 question categories to guide the creation of 1,304 questions. Additionally, we conduct an in-depth analysis to ascertain how specific attributes of literary fictions (e.g., novel types, character numbers, the year of publication) impact LLM performance in evaluations. Through a series of experiments involving various state-of-the-art LLMs, our findings reveal that these models face considerable challenges in effectively addressing questions related to literary fictions, with ChatGPT reaching only 57.08% under the zero-shot setting. The dataset will be publicly available at https://github.com/tjunlp-lab/LFED.git. Linhao Yu, Qun Liu 0001, Deyi Xiong |
LREC/COLING | 3 |
| 2024 | Towards Robust In-Context Learning for Machine Translation with Large Language ModelsabstractUsing large language models (LLMs) for machine translation via in-context learning (ICL) has become an interesting research direction of machine translation (MT) in recent years. Its main idea is to retrieve a few translation pairs as demonstrations from an additional datastore (parallel corpus) to guide translation without updating the LLMs. However, the underlying noise of retrieved demonstrations usually dramatically deteriorate the performance of LLMs. In this paper, we propose a robust method to enable LLMs to achieve robust translation with ICL. The method incorporates a multi-view approach, considering both sentence- and word-level information, to select demonstrations that effectively avoid noise. At the sentence level, a margin-based score is designed to avoid semantic noise. At the word level, word embeddings are utilized to evaluate the related tokens and change the weight of words in demonstrations. By considering both sentence- and word-level similarity, the proposed method provides fine-grained demonstrations that effectively prompt the translation of LLMs. Experimental results demonstrate the effectiveness of our method, particularly in domain adaptation. Shaolin Zhu, Menglong Cui, Deyi Xiong |
LREC/COLING | 3 |
| 2024 | Reassessing Non-Autoregressive Neural Machine Translation with a Fine-Grained Error TaxonomyabstractNon-autoregressive neural machine translation (NAT) has made remarkable progress since it is proposed. The performance of NAT in terms of BLEU has approached or even matched that of autoregressive neural machine translation (AT). However, other evaluation metrics show that NAT still lags behind. Unfortunately, these metrics only provide a numerical difference, and it is unclear how the translations produced by NAT differ from those produced by AT. In addition, the multimodality problem is always a significant issue in NAT. To assess whether NAT models are fully capable of solving the multimodality problem and achieving the performance of AT, we specifically design an error taxonomy to annotate errors in translations. The taxonomy is grounded on a systematic and hierarchical error analysis. We carry out an extensive annotation with professional annotators and analyze four NAT models and two AT models. Our analysis and experiments show that (1) the number of errors in NAT translations marked by annotators is 1.54 times that of AT translations, (2) the multimodality problem of NAT affects translations from lexical to syntactic levels, and even up to discourse, and (3) the four NAT models cannot fully eradicate the multimodality problem despite mitigation efforts. Longyue Wang, Zhaopeng Tu, Deyi Xiong |
ECAI | 4 |
| 2024 | End-to-End Speech Translation with Mutual Knowledge DistillationabstractMulti-task learning (MTL) is widely used to improve end-to-end speech translation (ST), which implicitly transfer knowledge from auxiliary automatic speech recognition (ASR) and/or machine translation (MT) to ST through shared modules. In this study, we find that triple-task MTL (ST+MT+ASR) suffers from a knowledge transfer limitation that leads to performance stagnation compared with dual-task MTL (ST+MT or ST+ASR). To address this issue, we propose a simple yet effective method, ST-MKD (Speech Translation with Mutual Knowledge Distillation). In ST-MKD, we employ a mutual knowledge distillation framework to mutually enhance dual-task MTL models with different knowledge bases, and explore regularization to maintain the consistency of the task representations. Experiments on the ST benchmark dataset MuST-C show that ST-MKD significantly outperforms strong MTL baseline and achieves state-of-the-art performance under three speech pre-training settings. Further analyses confirm that our approach effectively overcomes the knowledge transfer limitation of triple-task MTL. Zhengshan Xue, Yikun Lei, Deyi Xiong |
ICASSP | 4 |
| 2024 | Enhanced Transfer Learning with Efficient Modeling and Adaptive Fusion of Knowledge Via Prompt TuningabstractThis work presents a novel and parameter-efficient transfer learning framework. The framework consists of two phases: knowledge modeling based on prompt decomposition and knowledge transfer based on attention. Specifically, during the first phase, we decompose the prompt into parameter spaces of different ranks and leverage their characteristics to precisely model both general knowledge and task-specific knowledge separately. During the second phase, we train an attention module to adaptively integrate task-specific knowledge and generate an instance-wise prompt, which is then further fine-tuned. Through these two stages, our approach can accurately and efficiently transfer task knowledge of different granularities and types based on the input sample during inference. Extensive experiments demonstrate that our approach outperforms the multi-task learning baselines and SOTA parameter-efficient transfer learning methods. Furthermore, despite using only 0.13% of the parameters compared to full-parameter fine-tuning, it achieves an absolute improvement ranging from 0.3 to 3.9 across different benchmarks. Zishan Guo, Yulong Zeng, Deyi Xiong |
ICASSP | 4 |
| 2024 | IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware NeuronsabstractIt is widely acknowledged that large language models (LLMs) encode a vast reservoir of knowledge after being trained on mass data. Recent studies disclose knowledge conflicts in LLM generation, wherein outdated or incorrect parametric knowledge (i.e., encoded knowledge) contradicts new knowledge provided in the context. To mitigate such knowledge conflicts, we propose a novel framework, IRCAN (Identifying and Reweighting Context-Aware Neurons) to capitalize on neurons that are crucial in processing contextual cues. Specifically, IRCAN first identifies neurons that significantly contribute to context processing, utilizing a context-aware attribution score derived from integrated gradients. Subsequently, the identified context-aware neurons are strengthened via reweighting. In doing so, we steer LLMs to generate context-sensitive outputs with respect to the new knowledge provided in the context. Extensive experiments conducted across a variety of models and tasks demonstrate that IRCAN not only achieves remarkable improvements in handling knowledge conflicts but also offers a scalable, plug-and-play solution that can be integrated seamlessly with existing models. Our codes are released at https://github.com/danshi777/IRCAN. Dan Shi 0001, Renren Jin, Tianhao Shen, Weilong Dong, Xinwei Wu 0001, Deyi Xiong |
NeurIPS | 6 |
| 2024 | Star-Agents: Automatic Data Optimization with LLM Agents for Instruction TuningabstractThe efficacy of large language models (LLMs) on downstream tasks usually hinges on instruction tuning, which relies critically on the quality of training data. Unfortunately, collecting high-quality and diverse data is both expensive and time-consuming. To mitigate this issue, we propose a novel Star-Agents framework, which automates the enhancement of data quality across datasets through multi-agent collaboration and assessment. The framework adopts a three-pronged strategy. It initially generates diverse instruction data with multiple LLM agents through a bespoke sampling method. Subsequently, the generated data undergo a rigorous evaluation using a dual-model method that assesses both difficulty and quality. Finaly, the above process evolves in a dynamic refinement phase, where more effective LLMs are prioritized, enhancing the overall data quality. Our empirical studies, including instruction tuning experiments with models such as Pythia and LLaMA, demonstrate the effectiveness of the proposed framework. Optimized datasets have achieved substantial improvements, with an average increase of 12\% and notable gains in specific metrics, such as a 40\% improvement in Fermi, as evidenced by benchmarks like MT-bench, Vicuna bench, and WizardLM testset. Codes will be released soon. Yehui Tang 0001, Haochen Qin, Renren Jin, Deyi Xiong, Kai Han 0002, Yunhe Wang 0001 |
NeurIPS | 6 |
| 2024 | Speaker voice normalization for end-to-end speech translation
Zhengshan Xue, Tingxun Shi, Deyi Xiong |
Expert Syst. Appl. | 4 |
| 2024 | Regularizing cross-attention learning for end-to-end speech translation with ASR and MT attention matrices
Yikun Lei, Deyi Xiong |
Expert Syst. Appl. | 4 |
| 2024 | VisTFC: Vision-guided target-side future context learning for neural machine translation
ShaoLin Zhu, Shangjie Li, Deyi Xiong |
Expert Syst. Appl. | 3 |
| 2024 | FEDS-ICL: Enhancing translation ability and efficiency of large language model by optimizing demonstration selection
ShaoLin Zhu, Leiyu Pan, Deyi Xiong |
Inf. Process. Manag. | 3 |
| 2024 | Mining parallel sentences from internet with multi-view knowledge distillation for low-resource language pairs
Shaolin Zhu, Shiwei Gu, Shangjie Li, Deyi Xiong |
Knowl. Inf. Syst. | 5 |
| 2024 | TCLNet: Turn-level contrastive learning network with reranking for dialogue state tracking
Chaobin You, Deyi Xiong |
Knowl. Based Syst. | 2 |
| 2023 | PEIT: Bridging the Modality Gap with Pre-trained Models for End-to-End Image TranslationabstractImage translation is a task that translates an image containing text in the source language to the target language.One major challenge with image translation is the modality gap between visual text inputs and textual inputs/outputs of machine translation (MT).In this paper, we propose PEIT, an end-to-end image translation framework that bridges the modality gap with pre-trained models.It is composed of four essential components: a visual encoder, a shared encoder-decoder backbone network, a vision-text representation aligner equipped with the shared encoder and a cross-modal regularizer stacked over the shared decoder.Both the aligner and regularizer aim at reducing the modality gap.To train PEIT, we employ a twostage pre-training strategy with an auxiliary MT task: (1) pre-training the MT model on the MT training data to initialize the shared encoder-decoder backbone network; and (2) pre-training PEIT with the aligner and regularizer on a synthesized dataset with rendered images containing text from the MT training data.In order to facilitate the evaluation of PEIT and promote research on image translation, we create a large-scale image translation corpus ECOIT containing 480K imagetranslation pairs via crowd-sourcing and manual post-editing from real-world images in the e-commerce domain.Experiments on the curated ECOIT benchmark dataset demonstrate that PEIT substantially outperforms both cascaded image translation systems (OCR+MT) and previous strong end-to-end image translation model, with fewer parameters and faster decoding speed.Codes are available at https: //github.com/lishangjie1/PEIT. Shaolin Zhu, Shangjie Li, Yikun Lei, Deyi Xiong |
ACL (1) | 4 |
| 2023 | CS2W: A Chinese Spoken-to-Written Style Conversion Dataset with Multiple Conversion TypesabstractSpoken texts (either manual or automatic transcriptions from automatic speech recognition (ASR)) often contain disfluencies and grammatical errors, which pose tremendous challenges to downstream tasks.Converting spoken into written language is hence desirable.Unfortunately, the availability of datasets for this is limited.To address this issue, we present CS2W, a Chinese Spoken-to-Written style conversion dataset comprising 7,237 spoken sentences extracted from transcribed conversational texts.Four types of conversion problems are covered in CS2W: disfluencies, grammatical errors, ASR transcription errors, and colloquial words.Our annotation convention, Zishan Guo, Linhao Yu, Renren Jin, Deyi Xiong |
EMNLP | 5 |
| 2023 | MMNMT: Modularizing Multilingual Neural Machine Translation with Flexibly Assembled MoE and Dense BlocksabstractMixture-of-Experts (MoE) based sparse architectures can significantly increase model capacity with sublinear computational overhead, which are hence widely used in massively multilingual neural machine translation (MNMT).However, they are prone to overfitting on lowresource language translation.In this paper, we propose a modularized MNMT framework that is able to flexibly assemble dense and MoEbased sparse modules to achieve the best of both worlds.The training strategy of the modularized MNMT framework consists of three stages: (1) Pre-training basic MNMT models with different training objectives or model structures, (2) Initializing modules of the framework with pre-trained couterparts (e.g., encoder, decoder and embedding layers) from the basic models and (3) Fine-tuning the modularized MNMT framework to fit modules from different models together.We pre-train three basic MNMT models from scratch: a dense model, an MoE-based sparse model and a new MoE model, termed as MoE-LGR that explores multiple Language-Group-specifc Routers to incorporate language group knowledge into MNMT.The strengths of these pre-trained models are either on low-resource language translation, highresource language translation or zero-shot translation.Our modularized MNMT framework attempts to incorporate these advantages into a single model with reasonable initialization and fine-tuning.Experiments on widely-used benchmark datasets demonstrate that the proposed modularized MNMT framwork substantially outperforms both MoE and dense models on high-and low-resource language translation as well as zero-shot translation.Our framework facilitates the combination of different methods with their own strengths and recycling off-the-shelf models for multilingual neural machine translation.Codes are available at https://github.com/lishangjie1/MMNMT. Shangjie Li, Xiangpeng Wei, Shaolin Zhu, Baosong Yang, Deyi Xiong |
EMNLP | 6 |
| 2023 | DEPN: Detecting and Editing Privacy Neurons in Pretrained Language ModelsabstractLarge language models pretrained on a huge amount of data capture rich knowledge and information in the training data.The ability of data memorization and regurgitation in pretrained language models, revealed in previous studies, brings the risk of data leakage.In order to effectively reduce these risks, we propose a framework DEPN to Detect and Edit Privacy Neurons in pretrained language models, partially inspired by knowledge neurons and model editing.In DEPN, we introduce a novel method, termed as privacy neuron detector, to locate neurons associated with private information, and then edit these detected privacy neurons by setting their activations to zero.Furthermore, we propose a privacy neuron aggregator dememorize private information in a batch processing manner.Experimental results show that our method can significantly and efficiently reduce the exposure of private data leakage without deteriorating the performance of the model.Additionally, we empirically demonstrate the relationship between model memorization and privacy neurons, from multiple perspectives, including model size, training time, prompts, privacy neuron distribution, illustrating the robustness of our approach. Xinwei Wu 0001, Junzhuo Li, Weilong Dong, Shuangzhi Wu, Chao Bian 0006, Deyi Xiong |
EMNLP | 7 |
| 2023 | Language Representation Projection: Can We Transfer Factual Knowledge across Languages in Multilingual Language Models?abstractMultilingual pretrained language models serve as repositories of multilingual factual knowledge.Nevertheless, a substantial performance gap of factual knowledge probing exists between high-resource languages and lowresource languages, suggesting limited implicit factual knowledge transfer across languages in multilingual pretrained language models.This paper investigates the feasibility of explicitly transferring relatively rich factual knowledge from English to non-English languages.To accomplish this, we propose two parameter-free Language Representation Projection modules (LRP2).The first module converts non-English representations into English-like equivalents, while the second module reverts English-like representations back into representations of the corresponding non-English language.Experimental results on the mLAMA dataset demonstrate that LRP2 significantly improves factual knowledge retrieval accuracy and facilitates knowledge transferability across diverse non-English languages.We further investigate the working mechanism of LRP2 from the perspectives of representation space and cross-lingual knowledge neuron. Shaoyang Xu, Junzhuo Li, Deyi Xiong |
EMNLP | 3 |
| 2023 | Joint Training and Decoding for Multilingual End-to-End Simultaneous Speech TranslationabstractRecent studies on end-to-end speech translation(ST) have facilitated the exploration of multilingual end-to-end ST and end-to-end simultaneous ST. In this paper, we investigate end-to-end simultaneous speech translation in a one-to-many multilingual setting which is closer to applications in real scenarios. We explore a separate decoder architecture and a unified architecture for joint synchronous training in this scenario. To further explore knowledge transfer across languages, we propose an asynchronous training strategy on the proposed unified decoder architecture. A multi-way aligned multilingual end-to-end ST dataset was curated as a benchmark testbed to evaluate our methods. Experimental results demonstrate the effectiveness of our models on the collected dataset. Our codes and data are available at: https://github.com/XiaoMi/TED-MMST. Wuwei Huang, Renren Jin, Wen Zhang 0015, Jian Luan 0001, Bin Wang 0004, Deyi Xiong |
ICASSP | 6 |
| 2023 | SCoMoE: Efficient Mixtures of Experts with Structured Communication
Zhiyuan Zeng 0004, Deyi Xiong |
ICLR | 2 |
| 2023 | Unsupervised and Few-Shot Parsing from Pretrained Language Models (Extended Abstract)abstractThis paper proposes two Unsupervised constituent Parsing models (UPOA and UPIO) that calculate inside and outside association scores solely based on the self-attention weight matrix learned in a pretrained language model. The proposed unsupervised parsing models are further extended to few-shot parsing models (FPOA, FPIO) that use a few annotated trees to fine-tune the linear projection matrices in self-attention. Experiments on PTB and SPRML show that both unsupervised and few-shot parsing methods are better than or comparable to the previous methods. Zhiyuan Zeng 0004, Deyi Xiong |
IJCAI | 2 |
| 2023 | GhostRNN: Reducing State Redundancy in RNN with Cheap OperationsabstractRecurrent neural network (RNNs) that are capable of modeling long-distance dependencies are widely used in various speech tasks, eg., keyword spotting (KWS) and speech enhancement (SE). Due to the limitation of power and memory in low-resource devices, efficient RNN models are urgently required for real-world applications. In this paper, we propose an efficient RNN architecture, GhostRNN, which reduces hidden state redundancy with cheap operations. In particular, we observe that partial dimensions of hidden states are similar to the others in trained RNN models, suggesting that redundancy exists in specific RNNs. To reduce the redundancy and hence computational cost, we propose to first generate a few intrinsic states, and then apply cheap operations to produce ghost states based on the intrinsic states. Experiments on KWS and SE tasks demonstrate that the proposed GhostRNN significantly reduces the memory usage (~40%) and computation cost while keeping performance similar. Xiaoxu Zheng, Yunhe Wang 0001, Michael Bi Mi, Deyi Xiong, Kai Han 0002 |
INTERSPEECH | 5 |
| 2023 | NAPG: Non-Autoregressive Program Generation for Hybrid Tabular-Textual Question Answering
Tengxun Zhang, Hongfei Xu, Josef van Genabith, Deyi Xiong, Hongying Zan |
NLPCC (1) | 4 |
| 2023 | JoinER-BART: Joint Entity and Relation Extraction With Constrained Decoding, Representation Reuse and FusionabstractJoint Entity and Relation Extraction (JERE) is an important research direction in Information Extraction (IE). Given the surprising performance with fine-tuning of pre-trained BERT in a wide range of NLP tasks, nowadays most studies for JERE are based on the BERT model. Rather than predicting a simple tag for each word, these approaches are usually forced to design complex tagging schemes, as they may have to extract entity-relation pairs which may overlap with others from the same sequence of word representations in a sentence. Recently, sequence-to-sequence (seq2seq) pre-trained BART models show better performance than BERT models in many NLP tasks. Importantly, a seq2seq BART model can simply generate sequences of (many) entity-relation triplets with its decoder, rather than just tag input words. In this paper, we present a new generative JERE framework based on pre-trained BART. Different from the basic seq2seq BART architecture: 1) our framework employs a constrained classifier which only predicts either a token of the input sentence or a relation in each decoding step, and 2) we reuse representations from the pre-trained BART encoder in the classifier instead of a newly trained weight matrix, as this better utilizes the knowledge of the pre-trained model and context-aware representations for classification, and empirically leads to better performance. In our experiments on the widely studied NYT and WebNLG datasets, we show that our approach outperforms previous studies and establishes a new state-of-the-art (92.91 and 91.37 F1 respectively in exact match evaluation). Hongyang Chang, Hongfei Xu, Josef van Genabith, Deyi Xiong, Hongying Zan |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | DetTrans: A Lightweight Framework to Detect and Translate Noisy Inputs SimultaneouslyabstractNeural machine translation (NMT) systems trained on clean data usually suffer from performance degradation when translating noisy inputs. Existing works attempt to improve the robustness of NMT normally via data augmentation, where synthetic noisy data are mixed with original clean data, either for training NMT with the standard NMT loss alone, or for tuning auxiliary tasks in a multi-task learning manner. Typical auxiliary tasks include detecting and correcting noises, exploiting noisy outputs for contrastive learning etc. The aforementioned two auxiliary tasks are generally designed independently, and the modules for detecting and correcting noises are heavyweight. In this article, we propose a new framework, DetTransNet (Detector-TranslatorNetwork), aiming to detect positions of noises in the input and translate the input simultaneously. The newly introduced noise detector module is essentially a lightweight binary classifier built upon the final layer of the encoder of the original Transformer model for the translation task, which is to identify at which position of the input has potential noise. The module has a very few parameters. In order to help the model capture the relationship between clean instances and their noisy counterparts, an extra loss is further introduced to enhance the interaction between clean and noisy data. In this way, we combine noise detection and contrastive learning together. As the model is able to identify and locate noises, a heuristic method is proposed to correct detected noises, in order to achieve better translations. Experiments show that DetTransNet is robust to four types of noises (deletion, insertion, swapping, keyboard), and obtain a substantial improvement of up to 1.6 BLEU points across different datasets. Zhengshan Xue, Tingxun Shi, Deyi Xiong |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Bridging between Cognitive Processing Signals and Linguistic Features via a Unified Attentional NetworkabstractCognitive processing signals can be used to improve natural language processing (NLP) tasks. However, it is not clear how these signals correlate with linguistic information. Bridging between human language processing and linguistic features has been widely studied in neurolinguistics, usually via single-variable controlled experiments with highly-controlled stimuli. Such methods not only compromises the authenticity of natural reading, but also are time-consuming and expensive. In this paper, we propose a data-driven method to investigate the relationship between cognitive processing signals and linguistic features. Specifically, we present a unified attentional framework that is composed of embedding, attention, encoding and predicting layers to selectively map cognitive processing signals to linguistic features. We define the mapping procedure as a bridging task and develop 12 bridging tasks for lexical, syntactic and semantic features. The proposed framework only requires cognitive processing signals recorded under natural reading as inputs, and can be used to detect a wide range of linguistic features with a single cognitive dataset. Observations from experiment results resonate with previous neuroscience findings. In addition to this, our experiments also reveal a number of interesting findings, such as the correlation between contextual eye-tracking features and tense of sentence. Yuqi Ren, Deyi Xiong |
AAAI | 2 |
| 2022 | KaFSP: Knowledge-Aware Fuzzy Semantic Parsing for Conversational Question Answering over a Large-Scale Knowledge BaseabstractIn this paper, we study two issues of semantic parsing approaches to conversational question answering over a large-scale knowledge base:(1) The actions defined in grammar are not sufficient to handle uncertain reasoning common in real-world scenarios.(2) Knowledge base information is not well exploited and incorporated into semantic parsing.To mitigate the two issues, we propose a knowledge-aware fuzzy semantic parsing framework (KaFSP).It defines fuzzy comparison operations in the grammar system for uncertain reasoning based on the fuzzy set theory.In order to enhance the interaction between semantic parsing and knowledge base, we incorporate entity triples from the knowledge base into a knowledgeaware entity disambiguation module.Additionally, we propose a multi-label classification framework to not only capture correlations between entity types and relations but also detect knowledge base information relevant to the current utterance.Both enhancements are based on pre-trained language models.Experiments on a large-scale conversational question answering benchmark demonstrate that the proposed KaFSP achieves significant improvements over previous state-of-the-art models, setting new SOTA results on 8 out of 10 question types, gaining improvements of over 10% F1 or accuracy on 3 question types, and improving overall F1 from 83.01% to 85.33%.The source code of KaFSP is available at https: //github.com/tjunlp-lab/KaFSP. Junzhuo Li, Deyi Xiong |
ACL (1) | 2 |
| 2022 | CogTaskonomy: Cognitively Inspired Task Taxonomy Is Beneficial to Transfer Learning in NLPabstractIs there a principle to guide transfer learning across tasks in natural language processing (NLP)?Taxonomy (Zamir et al., 2018) finds that a structure exists among visual tasks, as a principle underlying transfer learning for them.In this paper, we propose a cognitively inspired framework, CogTaskonomy, to learn taxonomy for NLP tasks.The framework consists of Cognitive Representation Analytics (CRA) and Cognitive-Neural Mapping (CNM).The former employs Representational Similarity Analysis, which is commonly used in computational neuroscience to find a correlation between brainactivity measurement and computational modeling, to estimate task similarity with taskspecific sentence representations.The latter learns to detect task relations by projecting neural representations from NLP models to cognitive signals (i.e., fMRI voxels).Experiments on 12 NLP tasks, where BERT/TinyBERT are used as the underlying models for transfer learning, demonstrate that the proposed Cog-Taskonomy is able to guide transfer learning, achieving performance competitive to the Analytic Hierarchy Process (Saaty, 1987) used in visual Taskonomy (Zamir et al., 2018) but without requiring exhaustive pairwise O(m 2 ) task transferring.Analyses further discover that CNM is capable of learning modelagnostic task taxonomy.The source code is available at https://github.com/ tjunlp-lab/CogTaskonomy.git. Deyi Xiong |
ACL (1) | 3 |
| 2022 | Efficient Cluster-Based k-Nearest-Neighbor Machine Translationabstractk-Nearest-Neighbor Machine Translation (kNN-MT) has been recently proposed as a non-parametric solution for domain adaptation in neural machine translation (NMT).It aims to alleviate the performance degradation of advanced MT systems in translating out-ofdomain sentences by coordinating with an additional token-level feature-based retrieval module constructed from in-domain data.Previous studies (Khandelwal et al., 2021; Zheng et al., 2021a) have already demonstrated that non-parametric NMT is even superior to models fine-tuned on out-of-domain data.In spite of this success, kNN retrieval is at the expense of high latency, in particular for large datastores.To make it practical, in this paper, we explore a more efficient kNN-MT and propose to use clustering to improve the retrieval efficiency.Concretely, we first propose a cluster-based Compact Network for feature reduction in a contrastive learning manner to compress context features into 90+% lower dimensional vectors.We then suggest a cluster-based pruning solution to filter out 10%~40% redundant nodes in large datastores while retaining translation quality.Our proposed methods achieve better or comparable performance while reducing up to 57% inference latency against the advanced non-parametric MT model on several machine translation benchmarks.Experimental results indicate that the proposed methods maintain the most useful information of the original datastore and the Compact Network shows good generalization on unseen domains.Codes are available at https: //github.com/tjunlp-lab/PCKMT.Datastore distribution C-II("bank") C-III("bank") C-VI("bill") Compact Layer Compact Layer centroid ✖ Datastore distribution C-VI("bill") C-II("bank") C-I("sell") C-III("bank") Datastore distribution C-VI("bill") C-II("bank") C-I("sell") C-III("bank") Join Repartition Prune C-III C-II A pruning example w.r.t."bank" ✔ C-I("sell") Kai Fan 0002, Boxing Chen, Deyi Xiong |
ACL (1) | 4 |
| 2022 | Learning Disentangled Semantic Representations for Zero-Shot Cross-Lingual Transfer in Multilingual Machine Reading ComprehensionabstractLinjuan Wu, Shaojuan Wu, Xiaowang Zhang, Deyi Xiong, Shizhan Chen, Zhiqiang Zhuang, Zhiyong Feng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Linjuan Wu, Shaojuan Wu, Xiaowang Zhang, Deyi Xiong, Shizhan Chen, Zhiqiang Zhuang, Zhiyong Feng 0002 |
ACL (1) | 4 |
| 2022 | ParaZh-22M: A Large-Scale Chinese Parabank via Machine TranslationabstractParaphrasing, i.e., restating the same meaning in different ways, is an important data augmentation approach for natural language processing (NLP). Zhang et al. (2019b) propose to extract sentence-level paraphrases from multiple Chinese translations of the same source texts, and construct the PKU Paraphrase Bank of 0.5M sentence pairs. However, despite being the largest Chinese parabank to date, the size of PKU parabank is limited by the availability of one-to-many sentence translation data, and cannot well support the training of large Chinese paraphrasers. In this paper, we relieve the restriction with one-to-many sentence translation data, and construct ParaZh-22M, a larger Chinese parabank that is composed of 22M sentence pairs, based on one-to-one bilingual sentence translation data and machine translation (MT). In our data augmentation experiments, we show that paraphrasing based on ParaZh-22M can bring about consistent and significant improvements over several strong baselines on a wide range of Chinese NLP tasks, including a number of Chinese natural language understanding benchmarks (CLUE) and low-resource machine translation. Wenjie Hao, Hongfei Xu, Deyi Xiong, Hongying Zan, Lingling Mu |
COLING | 3 |
| 2022 | Informative Language Representation Learning for Massively Multilingual Neural Machine TranslationabstractIn a multilingual neural machine translation model that fully shares parameters across all languages, an artificial language token is usually used to guide translation into the desired target language. However, recent studies show that prepending language tokens sometimes fails to navigate the multilingual neural machine translation models into right translation directions, especially on zero-shot translation. To mitigate this issue, we propose two methods, language embedding embodiment and language-aware multi-head attention, to learn informative language representations to channel translation into right directions. The former embodies language embeddings into different critical switching points along the information flow from the source to the target, aiming at amplifying translation direction guiding signals. The latter exploits a matrix, instead of a vector, to represent a language in the continuous space. The matrix is chunked into multiple heads so as to learn language representations in multiple subspaces. Experiment results on two datasets for massively multilingual neural machine translation demonstrate that language-aware multi-head attention benefits both supervised and zero-shot translation and significantly alleviates the off-target translation issue. Further linguistic typology prediction experiments show that matrix-based language representations learned by our methods are capable of capturing rich linguistic typology features. Renren Jin, Deyi Xiong |
COLING | 2 |
| 2022 | CoDoNMT: Modeling Cohesion Devices for Document-Level Neural Machine TranslationabstractCohesion devices, e.g., reiteration, coreference, are crucial for building cohesion links across sentences. In this paper, we propose a document-level neural machine translation framework, CoDoNMT, which models cohesion devices from two perspectives: Cohesion Device Masking (CoDM) and Cohesion Attention Focusing (CoAF). In CoDM, we mask cohesion devices in the current sentence and force NMT to predict them with inter-sentential context information. A prediction task is also introduced to be jointly trained with NMT. In CoAF, we attempt to guide the model to pay exclusive attention to relevant cohesion devices in the context when translating cohesion devices in the current sentence. Such a cohesion attention focusing strategy is softly applied to the self-attention layer. Experiments on three benchmark datasets demonstrate that our approach outperforms state-of-the-art document-level neural machine translation baselines. Further linguistic evaluation validates the effectiveness of the proposed model in producing cohesive translations. Yikun Lei, Yuqi Ren, Deyi Xiong |
COLING | 3 |
| 2022 | Language Branch Gated Multilingual Neural Machine TranslationabstractKnowledge transfer across languages is crucial for multilingual neural machine translation. In this paper, we propose language branch (LB) gated multilingual neural machine translation that encourages knowledge transfer within the same language branch with a LB-gated module that is integrated into both the encoder and decoder. The LB-gated module distinguishes LB-specific parameters from global parameters shared by all languages and routes languages from the same LB to the corresponding LB-specific network. Comprehensive experiments on the OPUS-100 dataset show that the proposed approach substantially improves translation quality on both middle- and low-resource languages over previous methods. Further analysis demonstrates its ability in learning similarities between language branches. Deyi Xiong |
COLING | 2 |
| 2022 | Recovering Gold from Black Sand: Multilingual Dense Passage Retrieval with Hard and False Negative SamplesabstractNegative samples have not been efficiently explored in multilingual dense passage retrieval.In this paper, we propose a novel multilingual dense passage retrieval framework, mHFN, to recover and utilize hard and false negative samples.mHFN consists of three key components: 1) a multilingual hard negative sample augmentation module that allows knowledge of indistinguishable passages to be shared across multiple languages and synthesizes new hard negative samples by interpolating representations of queries and existing hard negative samples, 2) a multilingual negative sample cache queue that stores negative samples from previous batches in each language to increase the number of multilingual negative samples used in training beyond the batch size limit, and 3) a lightweight adaptive false negative sample filter that uses generated pseudo labels to separate unlabeled false negative samples and converts them into positive passages in training.We evaluate mHFN on Mr. TyDi, a high-quality multilingual dense passage retrieval dataset covering eleven typologically diverse languages, and experimental results show that mHFN outperforms strong sparse, dense and hybrid baselines and achieves new stateof-the-art performance on all languages.Our source code is available at https://github. com/Magnetic2014/mHFN. Tianhao Shen, Mingtong Liu, Deyi Xiong |
EMNLP | 4 |
| 2022 | Long Text Generation with Topic-aware Discrete Latent Variable ModelabstractGenerating coherent long texts is an important yet challenging task, particularly for the openended generation task.Prior work based on discrete latent codes focuses on the modeling of discourse relation, resulting in discrete codes only learning shallow semantics (Ji and Huang, 2021).A natural text always revolves around several related topics and the transition across them is natural and smooth.In this work, we investigate whether discrete latent codes can learn information of topics.To this end, we build a topic-aware latent code-guided text generation model.To encourage discrete codes to model information about topics, we propose a span-level bag-of-words training objective for the model.Automatic and manual evaluation experiments show that our method can generate more topic-relevant and coherent texts. Erguang Yang, Mingtong Liu, Deyi Xiong, Yufeng Chen 0005, Jin An Xu |
EMNLP | 3 |
| 2022 | TGEA 2.0: A Large-Scale Diagnostically Annotated Dataset with Benchmark Tasks for Text Generation of Pretrained Language ModelsabstractIn order to diagnostically analyze and improve the capability of pretrained language models (PLMs) in text generation, we propose TGEA 2.0, to date the largest dataset built on machine-authored texts by PLMs with fine-grained semantic annotations on a wide variety of pathological generation errors. We collect 170K nominal, phrasal and sentential prompts from 6M natural sentences in 3 domains. These prompts are fed into 4 generative PLMs with their best decoding strategy to generate paragraphs. 195,629 sentences are extracted from these generated paragraphs for manual annotation, where 36K erroneous sentences are detected, 42K erroneous spans are located and categorized into an error type defined in a two-level error taxonomy. We define a \textbf{Mi}nimal \textbf{S}et of \textbf{E}rror-related \textbf{W}ords (MiSEW) for each erroneous span, which not only provides error-associated words but also rationalizes the reasoning behind the error. Quality control with a pre-annotation and feedback loop is performed before and during the entire annotation process. With the diagnostically annotated dataset, we propose 5 diagnosis benchmark tasks (i.e., erroneous text detection, MiSEW extraction, erroneous span location and correction together with error type classification) and 2 pathology mitigation benchmark tasks (pairwise comparison and word prediction). Experiment results on these benchmark tasks demonstrate that TGEA 2.0 is a challenging dataset that could facilitate further research on automatic diagnosis and pathology mitigation over machine texts. The dataset will be publicly available at https://github.com/tjunlp-lab/TGEA/. Huibin Ge, Chuang Liu 0009, Yulong Zeng, Qun Liu 0001, Deyi Xiong |
NeurIPS | 6 |
| 2022 | Unsupervised and few-shot parsing from pretrained language models
Zhiyuan Zeng 0004, Deyi Xiong |
Artif. Intell. | 2 |
| 2022 | Short text matching model with multiway semantic interaction based on multi-granularity semantic embedding
Deyi Xiong, Jingming Yang, Rui Li 0066, Deguang Peng |
Appl. Intell. | 3 |
| 2022 | Improving generation diversity via syntax-controlled paraphrasing
Erguang Yang, Mingtong Liu, Deyi Xiong, Jin An Xu, Yufeng Chen 0005 |
Neurocomputing | 3 |
| 2022 | AAN+: Generalized Average Attention Network for Accelerating Neural TransformerabstractTransformer benefits from the high parallelization of attention networks in fast training, but it still suffers from slow decoding partially due to the linear dependency O(m) of the decoder self-attention on previous target words at inference. In this paper, we propose a generalized average attention network (AAN+) aiming at speeding up decoding by reducing the dependency from O(m) to O(1). We find that the learned self-attention weights in the decoder follow some patterns which can be approximated via a dynamic structure. Based on this insight, we develop AAN+, extending our previously proposed average attention (Zhang et al., 2018a, AAN) to support more general position- and content-based attention patterns. AAN+ only requires to maintain a small constant number of hidden states during decoding, ensuring its O(1) dependency. We apply AAN+ as a drop-in replacement of the decoder selfattention and conduct experiments on machine translation (with diverse language pairs), table-to-text generation and document summarization. With masking tricks and dynamic programming, AAN+ enables Transformer to decode sentences around 20% faster without largely compromising in the training speed and the generation performance. Our results further reveal the importance of the localness (neighboring words) in AAN+ and its capability in modeling long-range dependency. Biao Zhang 0002, Deyi Xiong, Yubin Ge, Junfeng Yao, Jinsong Su |
J. Artif. Intell. Res. | 2 |
| 2022 | Knowledge-enhanced graph convolutional network for recommendation
Jingming Yang, Deyi Xiong, Deguang Peng |
Multim. Tools Appl. | 3 |
| 2022 | Improving Multiscale Object Detection With Off-Centered Semantics RefinementabstractFeature Pyramid (FP) is typically a fundamental component for detecting multi-scale objects. However, as the network deepens, FP faces two problems: (1) Information loss caused by channel reduction. (2) The insufficient effective receptive field due to convolution with the sliding window mode. We found that the above problems can be alleviated by increasing the semantics extraction weights of the off-centered feature map. In this paper, a new feature pyramid architecture named Off-Centered Semantics Refinement Feature Pyramid Network (OSR-FPN) is proposed. Specifically, OSR-FPN contains two components exploiting the Off-Centered Semantics Refinement (OSR) mechanism: Features Supplement Module (FSM) and Receptive Field Enlargement Module (RFEM). FSM and RFEM are respectively designed to complement the lost context at the highest pyramid level and enrich the semantics by expanding the receptive field. In addition, we propose the Sigmoid-interpolation Padding method to enhance our OSR. Experiments on MS COCO dataset and UAVDT object detection benchmarks demonstrate the effectiveness of our method. As a result, OSR-FPN achieves a better accuracy of complex object detection. Deyi Xiong, Ying Xie 0004, Huiming Wang 0002, Rui Li 0066 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Integrating Pre-trained Model into Rule-based Dialogue ManagementabstractRule-based dialogue management is still the most popular solution for industrial task-oriented dialogue systems for their interpretablility. However, it is hard for developers to maintain the dialogue logic when the scenarios get more and more complex. On the other hand, data-driven dialogue systems, usually with end-to-end structures, are popular in academic research and easier to deal with complex conversations, but such methods require plenty of training data and the behaviors are less interpretable. In this paper, we propose a method to leverages the strength of both rule-based and data-driven dialogue managers (DM). We firstly introduce the DM of Carina Dialog System (CDS, an advanced industrial dialogue system built by Microsoft). Then we propose the "model-trigger" design to make the DM trainable thus scalable to scenario changes. Furthermore, we integrate pre-trained models and empower the DM with few-shot capability. The experimental results demonstrate the effectiveness and strong few-shot capability of our method. Jun Quan, Qiang Gan 0004, Deyi Xiong, Yuchen Dong, Fangxin Ouyang, Ruiling Deng, Yang Yang 0012, Daxin Jiang |
AAAI | 4 |
| 2021 | Efficient Object-Level Visual Context Modeling for Multimodal Machine Translation: Masking Irrelevant Objects Helps GroundingabstractVisual context provides grounding information for multimodal machine translation (MMT). However, previous MMT models and probing studies on visual features suggest that visual information is less explored in MMT as it is often redundant to textual information. In this paper, we propose an Object-level Visual Context modeling framework (OVC) to efficiently capture and explore visual information for multimodal machine translation. With detected objects, the proposed OVC encourages MMT to ground translation on desirable visual objects by masking irrelevant objects in the visual modality. We equip the proposed with an additional object-masking loss to achieve this goal. The object-masking loss is estimated according to the similarity between masked objects and the source texts so as to encourage masking source-irrelevant objects. Additionally, in order to generate vision-consistent target words, we further propose a vision-weighted translation loss for OVC. Experiments on MMT datasets demonstrate that the proposed OVC model outperforms state-of-the-art MMT models and analyses show that masking irrelevant objects helps grounding in MMT. Deyi Xiong |
AAAI | 2 |
| 2021 | TGEA: An Error-Annotated Dataset and Benchmark Tasks for TextGeneration from Pretrained Language ModelsabstractJie He, Bo Peng, Yi Liao, Qun Liu, Deyi Xiong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jie He 0004, Bo Peng 0041, Qun Liu 0001, Deyi Xiong |
ACL/IJCNLP (1) | 5 |
| 2021 | CogAlign: Learning to Align Textual Neural Representations to Cognitive Language Processing SignalsabstractYuqi Ren, Deyi Xiong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yuqi Ren, Deyi Xiong |
ACL/IJCNLP (1) | 2 |
| 2021 | Multi-Head Highly Parallelized LSTM Decoder for Neural Machine TranslationabstractHongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Meng Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Meng Zhang 0019 |
ACL/IJCNLP (1) | 4 |
| 2021 | Chinese WPLC: A Chinese Dataset for Evaluating Pretrained Language Models on Word Prediction Given Long-Range ContextabstractThis paper presents a Chinese dataset for evaluating pretrained language models on Word Prediction given Long-term Context (Chinese WPLC).We propose both automatic and manual selection strategies tailored to Chinese to guarantee that target words in passages collected from over 69K novels can only be predicted with long-term context beyond the scope of sentences containing the target words.Dataset analysis reveals that the types of target words range from common nouns to Chinese 4-character idioms.We also observe that linguistic relations between target words and long-range context exhibit diversity, including lexical match, synonym, summary and reasoning.Experiment results show that the Chinese pretrained language model PanGu-α (Zeng et al., 2021) is 45 points behind human in terms of top-1 word prediction accuracy, indicating that Chinese WPLC is a challenging dataset. Huibin Ge, Deyi Xiong, Qun Liu 0001 |
EMNLP (1) | 3 |
| 2021 | Re-embedding Difficult Samples via Mutual Information Constrained Semantically Oversampling for Imbalanced Text ClassificationabstractDifficult samples of the minority class in imbalanced text classification are usually hard to be classified as they are embedded into an overlapping semantic region with the majority class.In this paper, we propose a Mutual Information constrained Semantically Oversampling framework (MISO) that can generate anchor instances to help the backbone network determine the re-embedding position of a non-overlapping representation for each difficult sample.MISO consists of (1) a semantic fusion module that learns entangled semantics among difficult and majority samples with an adaptive multi-head attention mechanism, (2) a mutual information loss that forces our model to learn new representations of entangled semantics in the non-overlapping region of the minority class, and (3) a coupled adversarial encoder-decoder that fine-tunes disentangled semantic representations to remain their correlations with the minority class, and then using these disentangled semantic representations to generate anchor instances for each difficult sample.Experiments on a variety of imbalanced text classification tasks demonstrate that anchor instances help classifiers achieve significant improvements over strong baselines. Shizhan Chen, Xiaowang Zhang, Zhiyong Feng 0002, Deyi Xiong, Shaojuan Wu, Chunliu Dou |
EMNLP (1) | 5 |
| 2021 | Syntactically-Informed Unsupervised Paraphrasing with Non-Parallel DataabstractPrevious works on syntactically controlled paraphrase generation heavily rely on largescale parallel paraphrase data that are not easily available for many languages and domains.In this paper, we take this research direction to the extreme and investigate whether it is possible to learn syntactically controlled paraphrase generation with non-parallel data.We propose a syntactically-informed unsupervised paraphrasing model based on conditional variational auto-encoder (VAE) which can generate texts in a specified syntactic structure.Particularly, we design a two-stage learning method to effectively train the model using non-parallel data.The conditional VAE is trained to reconstruct the input sentence according to the given input and its syntactic structure.Furthermore, to improve the syntactic controllability and semantic consistency of the pre-trained conditional VAE, we finetune it using syntax controlling and cycle reconstruction learning objectives, and employ Gumbel-Softmax to combine these new learning objectives.Experiment results demonstrate that the proposed model trained only on non-parallel data is capable of generating diverse paraphrases with specified structures.Additionally, we further validate the effectiveness of our method for generating syntactically adversarial examples on a sentiment analysis task. Erguang Yang, Mingtong Liu, Deyi Xiong, Changjian Hu, Jin An Xu, Yufeng Chen 0005 |
EMNLP (1) | 3 |
| 2021 | Modeling Homophone Noise for Robust Neural Machine TranslationabstractIn this paper, we propose a robust neural machine translation (NMT) framework to deal with homophone errors. The framework consists of a homophone noise detector and a syllable-aware NMT model. The detector identifies potential homophone errors in a textual sentence and converts them into syllables to form a mixed sequence that is then fed into the syllable-aware NMT. Extensive experiments on Chinese→English translation demonstrate that the proposed method not only significantly outperforms baselines on noisy test sets with homophone noise, but also achieves substantial improvements over them on clean texts. Wenjie Qin, Xiang Li 0104, Yuhui Sun, Deyi Xiong, Jianwei Cui 0002, Bin Wang 0004 |
ICASSP | 4 |
| 2021 | Domain-Aware Self-Attention for Multi-Domain Neural Machine Translation
Shiqi Zhang 0007, Deyi Xiong, Pei Zhang 0011, Boxing Chen |
Interspeech | 3 |
| 2021 | Probing Word Translations in the Transformer and Trading Decoder for Encoder LayersabstractHongfei Xu, Josef van Genabith, Qiuhui Liu, Deyi Xiong. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Hongfei Xu, Josef van Genabith, Qiuhui Liu, Deyi Xiong |
NAACL-HLT | 4 |
| 2021 | Explore Coarse-Grained Structures for Syntactically Controllable Paraphrase Generation
Erguang Yang, Mingtong Liu, Deyi Xiong, Changjian Hu, Jin An Xu, Yufeng Chen 0005 |
NLPCC (1) | 3 |
| 2021 | Enhanced aspect-based sentiment analysis models with progressive self-supervised attention learning
Jinsong Su, Jialong Tang, Ziyao Lu, Yubin Ge, Linfeng Song, Deyi Xiong, Le Sun 0001, Jiebo Luo 0001 |
Artif. Intell. | 7 |
| 2020 | Modeling Long Context for Task-Oriented Dialogue State GenerationabstractBased on the recently proposed transferable dialogue state generator (TRADE) (Wu et al., 2019) that predicts dialogue states from utterance-concatenated dialogue context, we propose a multi-task learning model with a simple yet effective utterance tagging technique and a bidirectional language model as an auxiliary task for task-oriented dialogue state generation.By enabling the model to learn a better representation of the long dialogue context, our approaches attempt to solve the problem that the performance of the baseline significantly drops when the input dialogue context sequence is long.In our experiments, our proposed model achieves a 7.03% relative improvement over the baseline, establishing a new state-of-the-art joint goal accuracy of 52.04% on the MultiWOZ 2.0 dataset. Jun Quan, Deyi Xiong |
ACL | 2 |
| 2020 | Dynamically Adjusting Transformer Batch Size by Monitoring Gradient Direction ChangeabstractThe choice of hyper-parameters affects the performance of neural models.While much previous research (Sutskever et al., 2013;Duchi et al., 2011;Kingma and Ba, 2015) focuses on accelerating convergence and reducing the effects of the learning rate, comparatively few papers concentrate on the effect of batch size.In this paper, we analyze how increasing batch size affects gradient direction, and propose to evaluate the stability of gradients with their angle change.Based on our observations, the angle change of gradient direction first tends to stabilize (i.e.gradually decrease) while accumulating mini-batches, and then starts to fluctuate.We propose to automatically and dynamically determine batch sizes by accumulating gradients of mini-batches and performing an optimization step at just the time when the direction of gradients starts to fluctuate.To improve the efficiency of our approach for large models, we propose a sampling approach to select gradients of parameters sensitive to the batch size.Our approach dynamically determines proper and efficient batch sizes during training.In our experiments on the WMT 14 English to German and English to French tasks, our approach improves the Transformer with a fixed 25k batch size by +0.73 and +0.82 BLEU respectively. Hongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu |
ACL | 3 |
| 2020 | Learning Source Phrase Representations for Neural Machine TranslationabstractThe Transformer translation model (Vaswani et al., 2017) based on a multi-head attention mechanism can be computed effectively in parallel and has significantly pushed forward the performance of Neural Machine Translation (NMT).Though intuitively the attentional network can connect distant words via shorter network paths than RNNs, empirical analysis demonstrates that it still has difficulty in fully capturing long-distance dependencies (Tang et al., 2018).Considering that modeling phrases instead of words has significantly improved the Statistical Machine Translation (SMT) approach through the use of larger translation blocks ("phrases") and its reordering ability, modeling NMT at phrase level is an intuitive proposal to help the model capture long-distance relationships.In this paper, we first propose an attentive phrase representation generation mechanism which is able to generate phrase representations from corresponding token representations.In addition, we incorporate the generated phrase representations into the Transformer translation model to enhance its ability to capture long-distance relationships.In our experiments, we obtain significant improvements on the WMT 14 English-German and English-French tasks on top of the strong Transformer baseline, which shows the effectiveness of our approach.Our approach helps Transformer Base models perform at the level of Transformer Big models, and even significantly better for long sentences, but with substantially fewer parameters and training steps.The fact that phrase representations help even in the big setting further supports our conjecture that they make a valuable contribution to long-distance relations. Hongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu, Jingyi Zhang 0002 |
ACL | 3 |
| 2020 | Lipschitz Constrained Parameter Initialization for Deep TransformersabstractThe Transformer translation model employs residual connection and layer normalization to ease the optimization difficulties caused by its multi-layer encoder/decoder structure. Previous research shows that even with residual connection and layer normalization, deep Transformers still have difficulty in training, and particularly Transformer models with more than 12 encoder/decoder layers fail to converge. In this paper, we first empirically demonstrate that a simple modification made in the official implementation, which changes the computation order of residual connection and layer normalization, can significantly ease the optimization of deep Transformers. We then compare the subtle differences in computation order in considerable detail, and present a parameter initialization method that leverages the Lipschitz constraint on the initialization of Transformer parameters that effectively ensures training convergence. In contrast to findings in previous research we further demonstrate that with Lipschitz parameter initialization, deep Transformers with the original computation order can converge, and obtain significant BLEU improvements with up to 24 layers. In contrast to previous research which focuses on deep encoders, our approach additionally enables Transformers to also benefit from deep decoders. Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Jingyi Zhang 0002 |
ACL | 4 |
| 2020 | Balanced Joint Adversarial Training for Robust Intent Detection and Slot FillingabstractJoint intent detection and slot filling has recently achieved tremendous success in advancing the performance of utterance understanding.However, many joint models still suffer from the robustness problem, especially on noisy inputs or rare/unseen events.To address this issue, we propose a Joint Adversarial Training (JAT) model to improve the robustness of joint intent detection and slot filling, which consists of two parts: (1) automatically generating joint adversarial examples to attack the joint model, and (2) training the model to defend against the joint adversarial examples so as to robustify the model on small perturbations.As the generated joint adversarial examples have different impacts on the intent detection and slot filling loss, we further propose a Balanced Joint Adversarial Training (BJAT) model that applies a balance factor as a regularization term to the final loss function, which yields a stable training procedure.Extensive experiments and analyses on the lightweight models show that our proposed methods achieve significantly higher scores and substantially improve the robustness of both intent detection and slot filling.In addition, the combination of our BJAT with BERT-large achieves state-of-the-art results on two datasets. Deyi Xiong, Chongyang Shi 0001, Changjian Hu |
COLING | 2 |
| 2020 | Cycle-Consistent Adversarial Autoencoders for Unsupervised Text Style TransferabstractUnsupervised text style transfer is full of challenges due to the lack of parallel data and difficulties in content preservation.In this paper, we propose a novel neural approach to unsupervised text style transfer which we refer to as Cycle-consistent Adversarial autoEncoders (CAE) trained from non-parallel data.CAE consists of three essential components: (1) LSTM autoencoders that encode a text in one style into its latent representation and decode an encoded representation into its original text or a transferred representation into a style-transferred text, (2) adversarial style transfer networks that use an adversarially trained generator to transform a latent representation in one style into a representation in another style, and (3) a cycle-consistent constraint that enhances the capacity of the adversarial style transfer networks in content preservation.The entire CAE with these three components can be trained end-to-end.Extensive experiments and in-depth analyses on two widely-used public datasets consistently validate the effectiveness of proposed CAE in both style transfer and content preservation against several strong baselines in terms of four automatic evaluation metrics and human evaluation. Yufang Huang, Wentao Zhu 0001, Deyi Xiong, Yiye Zhang, Changjian Hu |
COLING | 3 |
| 2020 | A Learning-Exploring Method to Generate Diverse Paraphrases with Multi-Objective Deep Reinforcement LearningabstractParaphrase generation (PG) is of great importance to many downstream tasks in natural language processing.Diversity is an essential nature to PG for enhancing generalization capability and robustness of downstream applications.Recently, neural sequence-to-sequence (Seq2Seq) models have shown promising results in PG.However, traditional model training for PG focuses on optimizing model prediction against single reference and employs cross-entropy loss, which objective is unable to encourage model to generate diverse paraphrases.In this work, we present a novel approach with multi-objective learning to PG.We propose a learning-exploring method to generate sentences as learning objectives from the learned data distribution, and employ reinforcement learning to combine these new learning objectives for model training.We first design a sample-based algorithm to explore diverse sentences.Then we introduce several reward functions to evaluate the sampled sentences as learning signals in terms of expressive diversity and semantic fidelity, aiming to generate diverse and high-quality paraphrases.To effectively optimize model performance satisfying different evaluating aspects, we use a GradNorm-based algorithm that automatically balances these training objectives.Experiments and analyses on Quora and Twitter datasets demonstrate that our proposed method not only gains a significant increase in diversity but also improves generation quality over several state-of-the-art baselines. Mingtong Liu, Erguang Yang, Deyi Xiong, Changjian Hu, Jin An Xu, Yufeng Chen 0005 |
COLING | 3 |
| 2020 | Learning to Reuse Translations: Guiding Neural Machine Translation with ExamplesabstractIn this paper, we study the problem of enabling neural machine translation (NMT) to reuse previous translations from similar examples in target prediction. Distinguishing reusable translations from noisy segments and learning to reuse them in NMT are non-trivial. To solve these challenges, we propose an Example-Guided NMT (EGNMT) framework with two models: (1) a noise-masked encoder model that masks out noisy words according to word alignments and encodes the noise-masked sentences with an additional example encoder and (2) an auxiliary decoder model that predicts reusable words via an auxiliary decoder sharing parameters with the primary decoder. We define and implement the two models with the state-of-the-art Transformer. Experiments show that the noise-masked encoder model allows NMT to learn useful information from examples with low fuzzy match scores (FMS) while the auxiliary decoder model is good for high-FMS examples. More experiments on Chinese-English, English-German and English-Spanish translation demonstrate that the combination of the two EGNMT models can achieve improvements of up to +9 BLEU points over the baseline system and +7 BLEU points over a two-encoder Transformer. Shaohui Kuang, Deyi Xiong |
ECAI | 3 |
| 2020 | Learning Contextualized Sentence Representations for Document-Level Neural Machine TranslationabstractDocument-level machine translation incorporates inter-sentential dependencies into the translation of a source sentence. In this paper, we propose a new framework to model cross-sentence dependencies by training neural machine translation (NMT) to predict both the target translation and surrounding sentences of a source sentence. By enforcing the NMT model to predict source context, we want the model to learn source sentence representations that capture document-level dependencies on the source side. We further propose two different methods to learn and integrate such contextualized sentence embeddings into NMT: a joint training method that jointly trains an NMT model with the source context prediction model and a pre-training & fine-tuning method that pretrains the source context prediction model on a large-scale monolingual document corpus and then fine-tunes it with the NMT model. Experiments on Chinese-English and English-German translation show that both methods can substantially improve the translation quality over a strong document-level Transformer baseline. Wei Chen 0071, Deyi Xiong |
ECAI | 6 |
| 2020 | TED-CDB: A Large-Scale Chinese Discourse Relation Dataset on TED TalksabstractAs different genres are known to differ in their communicative properties and as previously, for Chinese, discourse relations have only been annotated over news text, we have created the TED-CDB dataset.TED-CDB comprises a large set of TED talks in Chinese that have been manually annotated according to the goals and principles of Penn Discourse Treebank, but adapted to features that are not present in English.It serves as a unique Chinese corpus of spoken discourse.Benchmark experiments show that TED-CDB poses a challenge for state-of-the-art discourse relation classifiers, whose F1 performance on 4way classification is <60%.This is a dramatic drop of 35% from performance on the news text in the Chinese Discourse Treebank.Transfer learning experiments have been carried out with the TED-CDB for both same-language cross-domain transfer and same-domain crosslanguage transfer.Both demonstrate that the TED-CDB can improve the performance of systems being developed for languages other than Chinese and would be helpful for insufficient or unbalanced data in other corpora.The dataset and our Chinese annotation guidelines has been made freely available.1 Wanqiu Long, Bonnie L. Webber, Deyi Xiong |
EMNLP (1) | 3 |
| 2020 | RiSAWOZ: A Large-Scale Multi-Domain Wizard-of-Oz Dataset with Rich Semantic Annotations for Task-Oriented Dialogue ModelingabstractIn order to alleviate the shortage of multidomain data and to capture discourse phenomena for task-oriented dialogue modeling, we propose RiSAWOZ, a large-scale multidomain Chinese Wizard-of-Oz dataset with Rich Semantic Annotations.RiSAWOZ contains 11.2K human-to-human (H2H) multiturn semantically annotated dialogues, with more than 150K utterances spanning over 12 domains, which is larger than all previous annotated H2H conversational datasets.Both single-and multi-domain dialogues are constructed, accounting for 65% and 35%, respectively.Each dialogue is labeled with comprehensive dialogue annotations, including dialogue goal in the form of natural language description, domain, dialogue states and acts at both the user and system side.In addition to traditional dialogue annotations, we especially provide linguistic annotations on discourse phenomena, e.g., ellipsis and coreference, in dialogues, which are useful for dialogue coreference and ellipsis resolution tasks.Apart from the fully annotated dataset, we also present a detailed description of the data collection procedure, statistics and analysis of the dataset.A series of benchmark models and results are reported, including natural language understanding (intent detection & slot filling), dialogue state tracking and dialogue contextto-text generation, as well as coreference and ellipsis resolution, which facilitate the baseline comparison for future research on this corpus.1 Jun Quan, Shian Zhang, Zizhong Li, Deyi Xiong |
EMNLP (1) | 5 |
| 2020 | Exploring Bilingual Parallel Corpora for Syntactically Controllable Paraphrase GenerationabstractParaphrase generation is of great importance to many downstream tasks in natural language processing. Recent efforts have focused on generating paraphrases in specific syntactic forms, which, generally, heavily relies on manually annotated paraphrase data that is not easily available for many languages and domains. In this paper, we propose a novel end-to-end framework to leverage existing large-scale bilingual parallel corpora to generate paraphrases under the control of syntactic exemplars. In order to train one model over the two languages of parallel corpora, we embed sentences of them into the same content and style spaces with shared content and style encoders using cross-lingual word embeddings. We propose an adversarial discriminator to disentangle the content and style space, and employ a latent variable to model the syntactic style of a given exemplar in order to guide the two decoders for generation. Additionally, we introduce cycle and masking learning schemes to efficiently train the model. Experiments and analyses demonstrate that the proposed model trained only on bilingual parallel data is capable of generating diverse paraphrases with desirable syntactic styles. Fine-tuning the trained model on a small paraphrase corpus makes it substantially outperform state-of-the-art paraphrase generation models trained on a larger paraphrase dataset. Mingtong Liu, Erguang Yang, Deyi Xiong, Chen Sheng, Changjian Hu, Jin An Xu, Yufeng Chen 0005 |
IJCAI | 3 |
| 2020 | Efficient Context-Aware Neural Machine Translation with Layer-Wise Weighting and Input-Aware GatingabstractExisting Neural Machine Translation (NMT) systems are generally trained on a large amount of sentence-level parallel data, and during prediction sentences are independently translated, ignoring cross-sentence contextual information. This leads to inconsistency between translated sentences. In order to address this issue, context-aware models have been proposed. However, document-level parallel data constitutes only a small part of the parallel data available, and many approaches build context-aware models based on a pre-trained frozen sentence-level translation model in a two-step training manner. The computational cost of these approaches is usually high. In this paper, we propose to make the most of layers pre-trained on sentence-level data in contextual representation learning, reusing representations from the sentence-level Transformer and significantly reducing the cost of incorporating contexts in translation. We find that representations from shallow layers of a pre-trained sentence-level encoder play a vital role in source context encoding, and propose to perform source context encoding upon weighted combinations of pre-trained encoder layers' outputs. Instead of separately performing source context and input encoding, we propose to iteratively and jointly encode the source input and its contexts and to generate input-aware context representations with a cross-attention layer and a gating mechanism, which resets irrelevant information in context encoding. Our context-aware Transformer model outperforms the recent CADec [Voita et al., 2019c] on the English-Russian subtitle data and is about twice as fast in training and decoding. Hongfei Xu, Deyi Xiong, Josef van Genabith, Qiuhui Liu |
IJCAI | 2 |
| 2020 | Shallow Discourse Annotation for Chinese TED TalksabstractText corpora annotated with language-related properties are an important resource for the development of Language Technology. The current work contributes a new resource for Chinese Language Technology and for Chinese-English translation, in the form of a set of TED talks (some originally given in English, some in Chinese) that have been annotated with discourse relations in the style of the Penn Discourse TreeBank, adapted to properties of Chinese text that are not present in English. The resource is currently unique in annotating discourse-level properties of planned spoken monologues rather than of written text. An inter-annotator agreement study demonstrates that the annotation scheme is able to achieve highly reliable results. Wanqiu Long, Xinyi Cai, James E. M. Reid, Bonnie L. Webber, Deyi Xiong |
LREC | 5 |
| 2020 | Neural Machine Translation with Deep AttentionabstractDeepening neural models has been proven very successful in improving the model's capacity when solving complex learning tasks, such as the machine translation task. Previous efforts on deep neural machine translation mainly focus on the encoder and the decoder, while little on the attention mechanism. However, the attention mechanism is of vital importance to induce the translation correspondence between different languages where shallow neural networks are relatively insufficient, especially when the encoder and decoder are deep. In this paper, we propose a deep attention model (DeepAtt). Based on the low-level attention information, DeepAtt is capable of automatically determining what should be passed or suppressed from the corresponding encoder layer so as to make the distributed representation appropriate for high-level attention and translation. We conduct experiments on NIST Chinese-English, WMT English-German, and WMT English-French translation tasks, where, with five attention layers, DeepAtt yields very competitive performance against the state-of-the-art results. We empirically find that with an adequate increase of attention layers, DeepAtt tends to produce more accurate attention weights. An in-depth analysis on the translation of important context words further reveals that DeepAtt significantly improves the faithfulness of system translations. Biao Zhang 0002, Deyi Xiong, Jinsong Su |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Alignment-Supervised Bidimensional Attention-Based Recursive Autoencoders for Bilingual Phrase RepresentationabstractExploiting semantic interactions between the source and target linguistic items at different levels of granularity is crucial for generating compact vector representations for bilingual phrases. To achieve this, we propose alignment-supervised bidimensional attention-based recursive autoencoders (ABattRAE) in this paper. ABattRAE first individually employs two recursive autoencoders to recover hierarchical tree structures of bilingual phrase, and treats the subphrase covered by each node on the tree as a linguistic item. Unlike previous methods, ABattRAE introduces a bidimensional attention network to measure the semantic matching degree between linguistic items of different languages, which enables our model to integrate information from all nodes by dynamically assigning varying weights to their corresponding embeddings. To ensure the accuracy of the generated attention weights in the attention network, ABattRAE incorporates word alignments as supervision signals to guide the learning procedure. Using the general stochastic gradient descent algorithm, we train our model in an end-to-end fashion, where the semantic similarity of translation equivalents is maximized while the semantic similarity of nontranslation pairs is minimized. Finally, we incorporate a semantic feature based on the learned bilingual phrase representations into a machine translation system for better translation selection. Experimental results on NIST Chinese-English and WMT English-German test sets show that our model achieves substantial improvements of up to 2.86 and 1.09 BLEU points over the baseline, respectively. Extensive in-depth analyses demonstrate the superiority of our model in learning bilingual phrase embeddings. Biao Zhang 0002, Deyi Xiong, Jinsong Su |
IEEE Trans. Cybern. | 2 |
| 2020 | Neural Machine Translation With GRU-Gated Attention ModelabstractNeural machine translation (NMT) heavily relies on context vectors generated by an attention network to predict target words. In practice, we observe that the context vectors for different target words are quite similar to one another and translations with such nondiscriminatory context vectors tend to be degenerative. We ascribe this similarity to the invariant source representations that lack dynamics across decoding steps. In this article, we propose a novel gated recurrent unit (GRU)-gated attention model (GAtt) for NMT. By updating the source representations with the previous decoder state via a GRU, GAtt enables translation-sensitive source representations that then contribute to discriminative context vectors. We further propose a variant of GAtt by swapping the input order of the source representations and the previous decoder state to the GRU. Experiments on the NIST Chinese-English, WMT14 English-German, and WMT17 English-German translation tasks show that the two GAtt models achieve significant improvements over the vanilla attention-based NMT. Further analyses on the attention weights and context vectors demonstrate the effectiveness of GAtt in enhancing the discriminating capacity of representations and handling the challenging issue of overtranslation. Biao Zhang 0002, Deyi Xiong, Jinsong Su |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | BiPaR: A Bilingual Parallel Dataset for Multilingual and Cross-lingual Reading Comprehension on NovelsabstractYimin Jing, Deyi Xiong, Zhen Yan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yimin Jing, Deyi Xiong, Yan Zhen |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Generating Highly Relevant QuestionsabstractJiazuo Qiu, Deyi Xiong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jiazuo Qiu, Deyi Xiong |
EMNLP/IJCNLP (1) | 2 |
| 2019 | GECOR: An End-to-End Generative Ellipsis and Co-reference Resolution Model for Task-Oriented DialogueabstractJun Quan, Deyi Xiong, Bonnie Webber, Changjian Hu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jun Quan, Deyi Xiong, Bonnie L. Webber, Changjian Hu |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Hierarchical Modeling of Global Context for Document-Level Neural Machine TranslationabstractXin Tan, Longyin Zhang, Deyi Xiong, Guodong Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Longyin Zhang, Deyi Xiong, Guodong Zhou 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Towards Linear Time Neural Machine Translation with Capsule NetworksabstractMingxuan Wang, Jun Xie, Zhixing Tan, Jinsong Su, Deyi Xiong, Lei Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mingxuan Wang, Zhixing Tan, Jinsong Su, Deyi Xiong, Lei Li 0005 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Detecting and Translating Dropped Pronouns in Neural Machine Translation
Shaohui Kuang, Deyi Xiong |
NLPCC (1) | 3 |
| 2019 | Future-Aware Knowledge Distillation for Neural Machine TranslationabstractAlthough future context is widely regarded useful for word prediction in machine translation, it is quite difficult in practice to incorporate it into neural machine translation. In this paper, we propose a future-aware knowledge distillation framework (FKD) to address this issue. In the FKD framework, we learn to distill future knowledge from a backward neural language model (teacher) to future-aware vectors (student) during the training phase. The future-aware vector for each word position is computed in a bridge network and optimized towards the corresponding hidden state in the backward neural language model via a knowledge distillation mechanism. We further propose an algorithm to jointly train the neural machine translation model, neural language model and knowledge distillation module end-to-end. The learned future-aware vectors are incorporated into the attention layer of the decoder to provide full-range context information during the decoding phase. Experiments on the NIST Chinese-English and WMT English-German translation tasks show that the proposed method significantly improves translation quality and word alignment. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Jiebo Luo 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Variational Recurrent Neural Machine TranslationabstractPartially inspired by successful applications of variational recurrent neural networks, we propose a novel variational recurrent neural machine translation (VRNMT) model in this paper. Different from the variational NMT, VRNMT introduces a series of latent random variables to model the translation procedure of a sentence in a generative way, instead of a single latent variable. Specifically, the latent random variables are included into the hidden states of the NMT decoder with elements from the variational autoencoder. In this way, these variables are recurrently generated, which enables them to further capture strong and complex dependencies among the output translations at different timesteps. In order to deal with the challenges in performing efficient posterior inference and large-scale training during the incorporation of latent variables, we build a neural posterior approximator, and equip it with a reparameterization technique to estimate the variational lower bound. Experiments on Chinese-English and English-German translation tasks demonstrate that the proposed model achieves significant improvements over both the conventional and variational NMT models. Jinsong Su, Deyi Xiong, Yaojie Lu 0001, Xianpei Han, Biao Zhang 0002 |
AAAI | 3 |
| 2018 | Attention Focusing for Neural Machine Translation by Bridging Source and Target EmbeddingsabstractIn neural machine translation, a source sequence of words is encoded into a vector from which a target sequence is generated in the decoding phase.Differently from statistical machine translation, the associations between source words and their possible target counterparts are not explicitly stored.Source and target words are at the two ends of a long information processing procedure, mediated by hidden states at both the source encoding and the target decoding phases.This makes it possible that a source word is incorrectly translated into a target word that is not any of its admissible equivalent counterparts in the target language. Shaohui Kuang, António Branco, Weihua Luo, Deyi Xiong |
ACL (1) | 5 |
| 2018 | Accelerating Neural Transformer via an Average Attention NetworkabstractWith parallelizable attention networks, the neural Transformer is very fast to train.However, due to the auto-regressive architecture and self-attention in the decoder, the decoding procedure becomes slow.To alleviate this issue, we propose an average attention network as an alternative to the self-attention network in the decoder of the neural Transformer.The average attention network consists of two layers, with an average layer that models dependencies on previous positions and a gating layer that is stacked over the average layer to enhance the expressiveness of the proposed attention network.We apply this network on the decoder part of the neural Transformer to replace the original target-side self-attention model.With masking tricks and dynamic programming, our model enables the neural Transformer to decode sentences over four times faster than its original version with almost no loss in training time and translation performance.We conduct a series of experiments on WMT17 translation tasks, where on 6 different language pairs, we obtain robust and consistent speed-ups in decoding.1 Biao Zhang 0002, Deyi Xiong, Jinsong Su |
ACL (1) | 2 |
| 2018 | Fusing Recency into Neural Machine Translation with an Inter-Sentence Gate ModelabstractNeural machine translation (NMT) systems are usually trained on a large amount of bilingual sentence pairs and translate one sentence at a time, ignoring inter-sentence information. This may make the translation of a sentence ambiguous or even inconsistent with the translations of neighboring sentences. In order to handle this issue, we propose an inter-sentence gate model that uses the same encoder to encode two adjacent sentences and controls the amount of information flowing from the preceding sentence to the translation of the current sentence with an inter-sentence gate. In this way, our proposed model can capture the connection between sentences and fuse recency from neighboring sentences into neural machine translation. On several NIST Chinese-English translation tasks, our experiments demonstrate that the proposed inter-sentence gate model achieves substantial improvements over the baseline. Shaohui Kuang, Deyi Xiong |
COLING | 2 |
| 2018 | Modeling Coherence for Neural Machine Translation with Dynamic and Topic CachesabstractSentences in a well-formed text are connected to each other via various links to form the cohesive structure of the text. Current neural machine translation (NMT) systems translate a text in a conventional sentence-by-sentence fashion, ignoring such cross-sentence links and dependencies. This may lead to generate an incoherent target text for a coherent source text. In order to handle this issue, we propose a cache-based approach to modeling coherence for neural machine translation by capturing contextual information either from recently translated sentences or the entire document. Particularly, we explore two types of caches: a dynamic cache, which stores words from the best translation hypotheses of preceding sentences, and a topic cache, which maintains a set of target-side topical words that are semantically related to the document to be translated. On this basis, we build a new layer to score target words in these two caches with a cache-based neural model. Here the estimated probabilities from the cache-based neural model are combined with NMT probabilities into the final word prediction probabilities via a gating mechanism. Finally, the proposed cache-based neural model is trained jointly with NMT system in an end-to-end manner. Experiments and analysis presented in this paper demonstrate that the proposed cache-based model achieves substantial improvements over several state-of-the-art SMT and NMT baselines. Shaohui Kuang, Deyi Xiong, Weihua Luo, Guodong Zhou 0001 |
COLING | 2 |
| 2018 | Neural Machine Translation with Decoding History Enhanced AttentionabstractNeural machine translation with source-side attention have achieved remarkable performance. however, there has been little work exploring to attend to the target-side which can potentially enhance the memory capbility of NMT. We reformulate a Decoding History Enhanced Attention mechanism (DHEA) to render NMT model better at selecting both source-side and target-side information. DHA enables dynamic control of the ratios at which source and target contexts contribute to the generation of target words, offering a way to weakly induce structure relations among both source and target tokens. It also allows training errors to be directly back-propagated through short-cut connections and effectively alleviates the gradient vanishing problem. The empirical study on Chinese-English translation shows that our model with proper configuration can improve by 0:9 BLEU upon Transformer and the best reported results in the dataset. On WMT14 English-German task and a larger WMT14 English-French task, our model achieves comparable results with the state-of-the-art. Mingxuan Wang, Zhixing Tan, Jinsong Su, Deyi Xiong, Chao Bian 0005 |
COLING | 5 |
| 2018 | Sentence Weighting for Neural Machine Translation Domain AdaptationabstractIn this paper, we propose a new sentence weighting method for the domain adaptation of neural machine translation. We introduce a domain similarity metric to evaluate the relevance between a sentence and an available entire domain dataset. The similarity of each sentence to the target domain is calculated with various methods. The computed similarity is then integrated into the training objective to weight sentences. The adaptation results on both IWSLT Chinese-English TED task and a task with only synthetic training parallel data show that our sentence weighting method is able to achieve an significant improvement over strong baselines. Shiqi Zhang 0007, Deyi Xiong |
COLING | 2 |
| 2018 | Encoding Gated Translation Memory into Neural Machine TranslationabstractTranslation memories (TM) facilitate human translators to reuse existing repetitive translation fragments.In this paper, we propose a novel method to combine the strengths of both TM and neural machine translation (NMT) for high-quality translation.We treat the target translation of a TM match as an additional reference input and encode it into NMT with an extra encoder.A gating mechanism is further used to balance the impact of the TM match on the NMT decoder.Experiment results on the UN corpus demonstrate that when fuzzy matches are higher than 50%, the quality of NMT translation can be significantly improved by over 10 BLEU points. Deyi Xiong |
EMNLP | 2 |
| 2018 | Simplifying Neural Machine Translation with Addition-Subtraction Twin-Gated Recurrent NetworksabstractIn this paper, we propose an additionsubtraction twin-gated recurrent network (ATR) to simplify neural machine translation.The recurrent units of ATR are heavily simplified to have the smallest number of weight matrices among units of all existing gated RNNs.With the simple addition and subtraction operation, we introduce a twin-gated mechanism to build input and forget gates which are highly correlated.Despite this simplification, the essential non-linearities and capability of modeling long-distance dependencies are preserved.Additionally, the proposed ATR is more transparent than LSTM/GRU due to the simplification.Forward self-attention can be easily established in ATR, which makes the proposed network interpretable.Experiments on WMT14 translation tasks demonstrate that ATR-based neural machine translation can yield competitive performance on English-German and English-French language pairs in terms of both translation quality and speed.Further experiments on NIST Chinese-English translation, natural language inference and Chinese word segmentation verify the generality and applicability of ATR on different natural language processing tasks. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Huiji Zhang |
EMNLP | 2 |
| 2018 | Context-Aware Phrase Representation for Statistical Machine Translation
Zhiwei Ruan, Jinsong Su, Deyi Xiong, Rongrong Ji |
PRICAI (1) | 3 |
| 2018 | Learning better discourse representation for implicit discourse relation recognition via attention networks
Biao Zhang 0002, Deyi Xiong, Jinsong Su, Min Zhang 0005 |
Neurocomputing | 2 |
| 2018 | A neural generative autoencoder for bilingual word embeddings
Jinsong Su, Biao Zhang 0002, Changxing Wu, Deyi Xiong |
Inf. Sci. | 6 |
| 2018 | Cross-lingual implicit discourse relation recognition with co-trainingabstractA lack of labeled corpora obstructs the research progress on implicit discourse relation recognition (DRR) for Chinese, while there are some available discourse corpora in other languages, such as English. In this paper, we propose a cross-lingual implicit DRR framework that exploits an available English corpus for the Chinese DRR task. We use machine translation to generate Chinese instances from a labeled English discourse corpus. In this way, each instance has two independent views: Chinese and English views. Then we train two classifiers in Chinese and English in a co-training way, which exploits unlabeled Chinese data to implement better implicit DRR for Chinese. Experimental results demonstrate the effectiveness of our method. Yaojie Lu 0001, Mu Xu, Changxing Wu, Deyi Xiong, Jinsong Su |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2018 | Alignment-consistent recursive neural networks for bilingual phrase embeddings
Jinsong Su, Biao Zhang 0002, Deyi Xiong, Yang Liu 0005, Min Zhang 0005 |
Knowl. Based Syst. | 3 |
| 2018 | A Hierarchy-to-Sequence Attentional Neural Machine Translation ModelabstractAlthough sequence-to-sequence attentional neural machine translation (NMT) has achieved great progress recently, it is confronted with two challenges: learning optimal model parameters for long parallel sentences and well exploiting different scopes of contexts. In this paper, partially inspired by the idea of segmenting a long sentence into short clauses, each of which can be easily translated by NMT, we propose a hierarchy-to-sequence attentional NMT model to handle these two challenges. Our encoder takes the segmented clause sequence as input and explores a hierarchical neural network structure to model words, clauses, and sentences at different levels, particularly with two layers of recurrent neural networks modeling semantic compositionality at the word and clause level. Correspondingly, the decoder sequentially translates segmented clauses and simultaneously applies two types of attention models to capture contexts of interclause and intraclause for translation prediction. In this way, we can not only improve parameter learning, but also well explore different scopes of contexts for translation. Experimental results on Chinese-English and English-German translation demonstrate the superiorities of the proposed model over the conventional NMT model. Jinsong Su, Jiali Zeng, Deyi Xiong, Yang Liu 0005, Mingxuan Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Lattice-Based Recurrent Neural Network Encoders for Neural Machine TranslationabstractNeural machine translation (NMT) heavily relies on word-level modelling to learn semantic representations of input sentences.However, for languages without natural word delimiters (e.g., Chinese) where input sentences have to be tokenized first,conventional NMT is confronted with two issues:1) it is difficult to find an optimal tokenization granularity for source sentence modelling, and2) errors in 1-best tokenizations may propagate to the encoder of NMT.To handle these issues, we propose word-lattice based Recurrent Neural Network (RNN) encoders for NMT,which generalize the standard RNN to word lattice topology.The proposed encoders take as input a word lattice that compactly encodes multiple tokenizations, and learn to generate new hidden states from arbitrarily many inputs and hidden states in preceding time steps.As such, the word-lattice based encoders not only alleviate the negative impact of tokenization errors but also are more expressive and flexible to embed input sentences.Experiment results on Chinese-English translation demonstrate the superiorities of the proposed encoders over the conventional encoder. Jinsong Su, Zhixing Tan, Deyi Xiong, Rongrong Ji, Xiaodong Shi, Yang Liu 0005 |
AAAI | 3 |
| 2017 | Neural Machine Translation Advised by Statistical Machine TranslationabstractNeural Machine Translation (NMT) is a new approach to machine translation that has made great progress in recent years. However, recent studies show that NMT generally produces fluent but inadequate translations (Tu et al. 2016b; 2016a; He et al. 2016; Tu et al. 2017). This is in contrast to conventional Statistical Machine Translation (SMT), which usually yields adequate but non-fluent translations. It is natural, therefore, to leverage the advantages of both models for better translations, and in this work we propose to incorporate SMT model into NMT framework. More specifically, at each decoding step, SMT offers additional recommendations of generated words based on the decoding information from NMT (e.g., the generated partial translation and attention history). Then we employ an auxiliary classifier to score the SMT recommendations and a gating function to combine the SMT recommendations with NMT generations, both of which are jointly trained within the NMT architecture in an end-to-end manner. Experimental results on Chinese-English translation show that the proposed approach achieves significant and consistent improvements over state-of-the-art NMT and SMT systems on multiple NIST test sets. Xing Wang 0007, Zhengdong Lu, Zhaopeng Tu, Hang Li 0001, Deyi Xiong, Min Zhang 0005 |
AAAI | 5 |
| 2017 | BattRAE: Bidimensional Attention-Based Recursive Autoencoders for Learning Bilingual Phrase EmbeddingsabstractIn this paper, we propose a bidimensional attention based recursiveautoencoder (BattRAE) to integrate clues and sourcetargetinteractions at multiple levels of granularity into bilingualphrase representations. We employ recursive autoencodersto generate tree structures of phrases with embeddingsat different levels of granularity (e.g., words, sub-phrases andphrases). Over these embeddings on the source and targetside, we introduce a bidimensional attention network to learntheir interactions encoded in a bidimensional attention matrix,from which we extract two soft attention weight distributionssimultaneously. These weight distributions enableBattRAE to generate compositive phrase representations viaconvolution. Based on the learned phrase representations, wefurther use a bilinear neural model, trained via a max-marginmethod, to measure bilingual semantic similarity. To evaluatethe effectiveness of BattRAE, we incorporate this semanticsimilarity as an additional feature into a state-of-the-art SMTsystem. Extensive experiments on NIST Chinese-English testsets show that our model achieves a substantial improvementof up to 1.63 BLEU points on average over the baseline. Biao Zhang 0002, Deyi Xiong, Jinsong Su |
AAAI | 2 |
| 2017 | Modeling Source Syntax for Neural Machine TranslationabstractEven though a linguistics-free sequence to sequence model in neural machine translation (NMT) has certain capability of implicitly learning syntactic information of source sentences, this paper shows that source syntax can be explicitly incorporated into NMT effectively to provide further improvements.Specifically, we linearize parse trees of source sentences to obtain structural label sequences.On the basis, we propose three different sorts of encoders to incorporate source syntax into NMT: 1) Parallel RNN encoder that learns word and label annotation vectors parallelly; 2) Hierarchical RNN encoder that learns word and label annotation vectors in a two-level hierarchy; and 3) Mixed RNN encoder that stitchingly learns word and label annotation vectors over sequences where words and labels are mixed.Experimentation on Chinese-to-English translation demonstrates that all the three proposed syntactic encoders are able to improve translation accuracy.It is interesting to note that the simplest RNN encoder, i.e., Mixed RNN encoder yields the best performance with an significant improvement of 1.4 BLEU points.Moreover, an in-depth analysis from several perspectives is provided to reveal how source syntax benefits NMT. Junhui Li 0001, Deyi Xiong, Zhaopeng Tu, Muhua Zhu, Min Zhang 0005, Guodong Zhou 0001 |
ACL (1) | 2 |
| 2017 | Translating Phrases in Neural Machine TranslationabstractPhrases play an important role in natural language understanding and machine translation (Sag et al., 2002;Villavicencio et al., 2005).However, it is difficult to integrate them into current neural machine translation (NMT) which reads and generates sentences word by word.In this work, we propose a method to translate phrases in NMT by integrating a phrase memory storing target phrases from a phrase-based statistical machine translation (SMT) system into the encoder-decoder architecture of NMT.At each decoding step, the phrase memory is first re-written by the SMT model, which dynamically generates relevant target phrases with contextual information provided by the NMT model.Then the proposed model reads the phrase memory to make probability estimations for all phrases in the phrase memory.If phrase generation is carried on, the NMT decoder selects an appropriate phrase from the memory to perform phrase translation and updates its decoding state by consuming the words in the selected phrase.Otherwise, the NMT decoder generates a word from the vocabulary as the general NMT decoder does.Experiment results on the Chinese→English translation show that the proposed model achieves significant improvements over the baseline on various test sets. Xing Wang 0007, Zhaopeng Tu, Deyi Xiong, Min Zhang 0005 |
EMNLP | 3 |
| 2017 | Improving Chinese-English Neural Machine Translation with Detected Usages of Function Words
Kunli Zhang, Hongfei Xu, Deyi Xiong, Qiuhui Liu, Hongying Zan |
NLPCC | 3 |
| 2017 | A Context-Aware Recurrent Encoder for Neural Machine TranslationabstractNeural machine translation (NMT) heavily relies on its encoder to capture the underlying meaning of a source sentence so as to generate a faithful translation. However, most NMT encoders are built upon either unidirectional or bidirectional recurrent neural networks, which either do not deal with future context or simply concatenate the history and future context to form context-dependent word representations, implicitly assuming the independence of the two types of contextual information. In this paper, we propose a novel context-aware recurrent encoder (CAEncoder), as an alternative to the widely-used bidirectional encoder, such that the future and history contexts can be fully incorporated into the learned source representations. Our CAEncoder involves a two-level hierarchy: The bottom level summarizes the history information, whereas the upper level assembles the summarized history and future context into source representations. Additionally, CAEncoder is as efficient as the bidirectional RNN encoder in terms of both training and decoding. Experiments on both Chinese-English and English-German translation tasks show that CAEncoder achieves significant improvements over the bidirectional RNN encoder on a widely-used NMT system. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Hong Duan |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Learning Event Expressions via Bilingual Structure ProjectionabstractIdentifying events of a specific type is a challenging task as events in texts are described in numerous and diverse ways. Aiming to resolve high complexities of event descriptions, previous work (Huang and Riloff, 2013) proposes multi-faceted event recognition and a bootstrapping method to automatically acquire both event facet phrases and event expressions from unannotated texts. However, to ensure high quality of learned phrases, this method is constrained to only learn phrases that match certain syntactic structures. In this paper, we propose a bilingual structure projection algorithm that explores linguistic divergences between two languages (Chinese and English) and mines new phrases with new syntactic structures, which have been ignored in the previous work. Experiments show that our approach can successfully find novel event phrases and structures, e.g., phrases headed by nouns. Furthermore, the newly mined phrases are capable of recognizing additional event descriptions and increasing the recall of event recognition. Ruihong Huang, Deyi Xiong, Min Zhang 0005 |
COLING | 3 |
| 2016 | Convolution-Enhanced Bilingual Recursive Neural Network for Bilingual Semantic ModelingabstractEstimating similarities at different levels of linguistic units, such as words, sub-phrases and phrases, is helpful for measuring semantic similarity of an entire bilingual phrase. In this paper, we propose a convolution-enhanced bilingual recursive neural network (ConvBRNN), which not only exploits word alignments to guide the generation of phrase structures but also integrates multiple-level information of the generated phrase structures into bilingual semantic modeling. In order to accurately learn the semantic hierarchy of a bilingual phrase, we develop a recursive neural network to constrain the learned bilingual phrase structures to be consistent with word alignments. Upon the generated source and target phrase structures, we stack a convolutional neural network to integrate vector representations of linguistic units on the structures into bilingual phrase embeddings. After that, we fully incorporate information of different linguistic units into a bilinear semantic similarity model. We introduce two max-margin losses to train the ConvBRNN model: one for the phrase structure inference and the other for the semantic similarity model. Experiments on NIST Chinese-English translation tasks demonstrate the high quality of the generated bilingual phrase structures with respect to word alignments and the effectiveness of learned semantic similarities on machine translation. Jinsong Su, Biao Zhang 0002, Deyi Xiong, Jianmin Yin |
COLING | 3 |
| 2016 | Improving Translation Selection with SupersensesabstractSelecting appropriate translations for source words with multiple meanings still remains a challenge for statistical machine translation (SMT). One reason for this is that most SMT systems are not good at detecting the proper sense for a polysemic word when it appears in different contexts. In this paper, we adopt a supersense tagging method to annotate source words with coarse-grained ontological concepts. In order to enable the system to choose an appropriate translation for a word or phrase according to the annotated supersense of the word or phrase, we propose two translation models with supersense knowledge: a maximum entropy based model and a supersense embedding model. The effectiveness of our proposed models is validated on a large-scale English-to-Spanish translation task. Results indicate that our method can significantly improve translation quality via correctly conveying the meaning of the source language to the target language. Haiqing Tang, Deyi Xiong, Oier Lopez de Lacalle, Eneko Agirre |
COLING | 2 |
| 2016 | Improving Statistical Machine Translation with Selectional PreferencesabstractLong-distance semantic dependencies are crucial for lexical choice in statistical machine translation. In this paper, we study semantic dependencies between verbs and their arguments by modeling selectional preferences in the context of machine translation. We incorporate preferences that verbs impose on subjects and objects into translation. In addition, bilingual selectional preferences between source-side verbs and target-side arguments are also investigated. Our experiments on Chinese-to-English translation tasks with large-scale training data demonstrate that statistical machine translation using verbal selectional preferences can achieve statistically significant improvements over a state-of-the-art baseline. Haiqing Tang, Deyi Xiong, Min Zhang 0005, Zhengxian Gong |
COLING | 2 |
| 2016 | Bilingual Autoencoders with Global Descriptors for Modeling Parallel SentencesabstractParallel sentence representations are important for bilingual and cross-lingual tasks in natural language processing. In this paper, we explore a bilingual autoencoder approach to model parallel sentences. We extract sentence-level global descriptors (e.g. min, max) from word embeddings, and construct two monolingual autoencoders over these descriptors on the source and target language. In order to tightly connect the two autoencoders with bilingual correspondences, we force them to share the same decoding parameters and minimize a corpus-level semantic distance between the two languages. Being optimized towards a joint objective function of reconstruction and semantic errors, our bilingual antoencoder is able to learn continuous-valued latent representations for parallel sentences. Experiments on both intrinsic and extrinsic evaluations on statistical machine translation tasks show that our autoencoder achieves substantial improvements over the baselines. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Hong Duan, Min Zhang 0005 |
COLING | 2 |
| 2016 | Variational Neural Machine TranslationabstractModels of neural machine translation are often from a discriminative family of encoderdecoders that learn a conditional distribution of a target sentence given a source sentence.In this paper, we propose a variational model to learn this conditional distribution for neural machine translation: a variational encoderdecoder model that can be trained end-to-end.Different from the vanilla encoder-decoder model that generates target translations from hidden representations of source sentences alone, the variational model introduces a continuous latent variable to explicitly model underlying semantics of source sentences and to guide the generation of target translations.In order to perform efficient posterior inference and large-scale training, we build a neural posterior approximator conditioned on both the source and the target sides, and equip it with a reparameterization technique to estimate the variational lower bound.Experiments on both Chinese-English and English-German translation tasks show that the proposed variational neural machine translation achieves significant improvements over the vanilla neural machine translation baselines. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Hong Duan, Min Zhang 0005 |
EMNLP | 2 |
| 2016 | Variational Neural Discourse Relation RecognizerabstractImplicit discourse relation recognition is a crucial component for automatic discourselevel analysis and nature language understanding.Previous studies exploit discriminative models that are built on either powerful manual features or deep discourse representations.In this paper, instead, we explore generative models and propose a variational neural discourse relation recognizer.We refer to this model as VarNDRR.VarNDRR establishes a directed probabilistic model with a latent continuous variable that generates both a discourse and the relation between the two arguments of the discourse.In order to perform efficient inference and learning, we introduce neural discourse relation models to approximate the prior and posterior distributions of the latent variable, and employ these approximated distributions to optimize a reparameterized variational lower bound.This allows VarNDRR to be trained with standard stochastic gradient methods.Experiments on the benchmark data set show that VarNDRR can achieve comparable results against stateof-the-art baselines without using any manual features. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Qun Liu 0001, Rongrong Ji, Hong Duan, Min Zhang 0005 |
EMNLP | 2 |
| 2016 | Topic-based term translation models for statistical machine translation
Deyi Xiong, Fandong Meng, Qun Liu 0001 |
Artif. Intell. | 1 |
| 2016 | Semantic Similarity from Natural Language and Ontology AnalysisabstractLearning semantic similarity for units of language or concepts is crucial not only for numerous tasks in computational linguistics, but also for language understanding and reasoning in the broad context of artificial intelligence.With growing interests and efforts in modeling and computing semantic measures in recent years, we have witnessed much progress in the following two strands of research that approach semantic similarity: corpus-based statistical methods and knowledge-enhanced methods with human knowledge defined in ontologies.This book by Harispe, Ranwez, Janaqi, and Montmain provides a detailed introduction to state-of-the-art research in these two lines of work. Deyi Xiong |
Comput. Linguistics | 1 |
| 2016 | Adapted competitive learning on continuous semantic space for word sense induction
Yanzhou Huang, Deyi Xiong, Xiaodong Shi, Yidong Chen 0001, Changxing Wu, Guimin Huang |
Neurocomputing | 2 |
| 2015 | A Context-Aware Topic Model for Statistical Machine TranslationabstractJinsong Su, Deyi Xiong, Yang Liu, Xianpei Han, Hongyu Lin, Junfeng Yao, Min Zhang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Jinsong Su, Deyi Xiong, Yang Liu 0005, Xianpei Han, Junfeng Yao, Min Zhang 0005 |
ACL (1) | 2 |
| 2015 | Graph-Based Collective Lexical Selection for Statistical Machine TranslationabstractLexical selection is of great importance to statistical machine translation. In this paper, we propose a graph-based frame-work for collective lexical selection. The framework is established on a translation graph that captures not only local associ-ations between source-side content words and their target translations but also target-side global dependencies in terms of relat-edness among target items. We also in-troduce a random walk style algorithm to collectively identify translations of source-side content words that are strongly related in translation graph. We validate the ef-fectiveness of our lexical selection frame-work on Chinese-English translation. Ex-periment results with large-scale training data show that our approach significantly improves lexical selection. 1 Jinsong Su, Deyi Xiong, Shujian Huang, Xianpei Han, Junfeng Yao |
EMNLP | 2 |
| 2015 | Bilingual Correspondence Recursive Autoencoder for Statistical Machine TranslationabstractLearning semantic representations and tree structures of bilingual phrases is beneficial for statistical machine translation.In this paper, we propose a new neural network model called Bilingual Correspondence Recursive Autoencoder (BCor-rRAE) to model bilingual phrases in translation.We incorporate word alignments into BCorrRAE to allow it freely access bilingual constraints at different levels.BCorrRAE minimizes a joint objective on the combination of a recursive autoencoder reconstruction error, a structural alignment consistency error and a crosslingual reconstruction error so as to not only generate alignment-consistent phrase structures, but also capture different levels of semantic relations within bilingual phrases.In order to examine the effectiveness of BCorrRAE, we incorporate both semantic and structural similarity features built on bilingual phrase representations and tree structures learned by BCorrRAE into a state-of-the-art SMT system.Experiments on NIST Chinese-English test sets show that our model achieves a substantial improvement of up to 1.55 BLEU points over the baseline. Jinsong Su, Deyi Xiong, Biao Zhang 0002, Yang Liu 0005, Junfeng Yao, Min Zhang 0005 |
EMNLP | 2 |
| 2015 | Learning Semantic Representations for Nonterminals in Hierarchical Phrase-Based TranslationabstractIn hierarchical phrase-based translation, coarse-grained nonterminal Xs may generate inappropriate translations due to the lack of sufficient information for phrasal substitution.In this paper we propose a framework to refine nonterminals in hierarchical translation rules with real-valued semantic representations.The semantic representations are learned via a weighted mean value and a minimum distance method using phrase vector representations obtained from large scale monolingual corpus.Based on the learned semantic vectors, we build a semantic nonterminal refinement model to measure semantic similarities between phrasal substitutions and nonterminal Xs in translation rules.Experiment results on Chinese-English translation show that the proposed model significantly improves translation quality on NIST test sets. Xing Wang 0007, Deyi Xiong, Min Zhang 0005 |
EMNLP | 2 |
| 2015 | Shallow Convolutional Neural Network for Implicit Discourse Relation RecognitionabstractImplicit discourse relation recognition remains a serious challenge due to the absence of discourse connectives.In this paper, we propose a Shallow Convolutional Neural Network (SCNN) for implicit discourse relation recognition, which contains only one hidden layer but is effective in relation recognition.The shallow structure alleviates the overfitting problem, while the convolution and nonlinear operations help preserve the recognition and generalization ability of our model.Experiments on the benchmark data set show that our model achieves comparable and even better performance when comparing against current state-of-the-art systems. Biao Zhang 0002, Jinsong Su, Deyi Xiong, Yaojie Lu 0001, Hong Duan, Junfeng Yao |
EMNLP | 3 |
| 2015 | Discriminative Reordering Model Adaptation via Structural Learning
Biao Zhang 0002, Jinsong Su, Deyi Xiong, Hong Duan, Junfeng Yao |
IJCAI | 3 |
| 2015 | Learning bilingual distributed phrase represenations for statistical machine translation
Deyi Xiong, Min Zhang 0005, Chunyu Kit |
MTSummit | 2 |
| 2015 | Backward and trigger-based language models for statistical machine translationabstractAbstract The language model is one of the most important knowledge sources for statistical machine translation. In this article, we present two extensions to standard n-gram language models in statistical machine translation: a backward language model that augments the conventional forward language model, and a mutual information trigger model which captures long-distance dependencies that go beyond the scope of standard n-gram language models. We introduce algorithms to integrate the two proposed models into two kinds of state-of-the-art phrase-based decoders. Our experimental results on Chinese/Spanish/Vietnamese-to-English show that both models are able to significantly improve translation quality in terms of BLEU and METEOR over a competitive baseline. Deyi Xiong, Min Zhang 0005 |
Nat. Lang. Eng. | 1 |
| 2015 | Topic-Based Coherence Modeling for Statistical Machine TranslationabstractCoherence that ties sentences of a text into a meaningfully connected structure is of great importance to text generation and translation. In this paper, we propose topic-based coherence models to produce coherence for document translation, in terms of the continuity of sentence topics in a text. We automatically extract a coherence chain for each source text to be translated. Based on the extracted source coherence chain, we adopt a maximum entropy classifier to predict the target coherence chain that defines a linear topic structure for the target document. We build two topic-based coherence models on the predicted target coherence chain: 1) a word level coherence model that helps the decoder select coherent word translations and 2) a phrase level coherence model that guides the decoder to select coherent phrase translations. We integrate the two models into a state-of-the-art phrase-based machine translation system. Experiments on large-scale training data show that our coherence models achieve substantial improvements over both the baseline and models that are built on either document topics or sentence topics obtained under the assumption of direct topic correspondence between the source and target side. Additionally, further evaluations on translation outputs suggest that target translations generated by our coherence models are more coherent and similar to reference translations than those generated by the baseline. Deyi Xiong, Min Zhang 0005, Xing Wang 0007 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | A Sense-Based Translation Model for Statistical Machine TranslationabstractThe sense in which a word is used deter-mines the translation of the word. In this paper, we propose a sense-based transla-tion model to integrate word senses into statistical machine translation. We build a broad-coverage sense tagger based on a nonparametric Bayesian topic model that automatically learns sense clusters for words in the source language. The pro-posed sense-based translation model en-ables the decoder to select appropriate translations for source words according to the inferred senses for these words us-ing maximum entropy classifiers. Our method is significantly different from pre-vious word sense disambiguation reformu-lated for machine translation in that the lat-ter neglects word senses in nature. We test the effectiveness of the proposed sense-based translation model on a large-scale Chinese-to-English translation task. Re-sults show that the proposed model sub-stantially outperforms not only the base-line but also the previous reformulated word sense disambiguation. 1 Deyi Xiong, Min Zhang 0005 |
ACL (1) | 1 |
| 2014 | Modeling Term Translation for Document-informed Machine TranslationabstractTerm translation is of great importance for statistical machine translation (SMT), especially document-informed SMT.In this paper, we investigate three issues of term translation in the context of documentinformed SMT and propose three corresponding models: (a) a term translation disambiguation model which selects desirable translations for terms in the source language with domain information, (b) a term translation consistency model that encourages consistent translations for terms with a high strength of translation consistency throughout a document, and (c) a term bracketing model that rewards translation hypotheses where bracketable source terms are translated as a whole unit.We integrate the three models into hierarchical phrase-based SMT and evaluate their effectiveness on NIST Chinese-English translation tasks with large-scale training data.Experiment results show that all three models can achieve significant improvements over the baseline.Additionally, we can obtain a further improvement when combining the three models. Fandong Meng, Deyi Xiong, Wenbin Jiang 0002, Qun Liu 0001 |
EMNLP | 2 |
| 2014 | A Topic-Based Reordering Model for Statistical Machine Translation
Xing Wang 0007, Deyi Xiong, Min Zhang 0005, Yu Hong 0001, Jianmin Yao 0001 |
NLPCC | 2 |
| 2014 | Topic-Based Dissimilarity and Sensitivity Models for Translation Rule SelectionabstractTranslation rule selection is a task of selecting appropriate translation rules for an ambiguous source-language segment. As translation ambiguities are pervasive in statistical machine translation, we introduce two topic-based models for translation rule selection which incorporates global topic information into translation disambiguation. We associate each synchronous translation rule with source- and target-side topic distributions.With these topic distributions, we propose a topic dissimilarity model to select desirable (less dissimilar) rules by imposing penalties for rules with a large value of dissimilarity of their topic distributions to those of given documents. In order to encourage the use of non-topic specific translation rules, we also present a topic sensitivity model to balance translation rule selection between generic rules and topic-specific rules. Furthermore, we project target-side topic distributions onto the source-side topic model space so that we can benefit from topic information of both the source and target language. We integrate the proposed topic dissimilarity and sensitivity model into hierarchical phrase-based machine translation for synchronous translation rule selection. Experiments show that our topic-based translation rule selection model can substantially improve translation quality. Min Zhang 0005, Xinyan Xiao, Deyi Xiong, Qun Liu 0001 |
J. Artif. Intell. Res. | 3 |
| 2013 | A Topic-Based Coherence Model for Statistical Machine TranslationabstractCoherence that ties sentences of a text into a meaningfully connected structure is of great importance to text generation and translation. In this paper, we propose a topic-based coherence model to produce coherence for document translation, in terms of the continuity of sentence topics in a text. We automatically extract a coherence chain for each source text to be translated. Based on the extracted source coherence chain, we adopt a maximum entropy classifier to predict the target coherence chain that defines a linear topic structure for the target document. The proposed topic-based coherence model then uses the predicted target coherence chain to help decoder select coherent word/phrase translations. Our experiments show that incorporating the topic-based coherence model into machine translation achieves substantial improvement over both the baseline and previous methods that integrate document topics rather than coherence chains into machine translation. Deyi Xiong, Min Zhang 0005 |
AAAI | 1 |
| 2013 | Max-Margin Synchronous Grammar Induction for Machine TranslationabstractTraditional synchronous grammar induction estimates parameters by maximizing likelihood, which only has a loose relation to translation quality.Alternatively, we propose a max-margin estimation approach to discriminatively inducing synchronous grammars for machine translation, which directly optimizes translation quality measured by BLEU.In the max-margin estimation of parameters, we only need to calculate Viterbi translations.This further facilitates the incorporation of various non-local features that are defined on the target side.We test the effectiveness of our max-margin estimation framework on a competitive hierarchical phrase-based system.Experiments show that our max-margin method significantly outperforms the traditional twostep pipeline for synchronous rule extraction by 1.3 BLEU points and is also better than previous max-likelihood estimation method. Xinyan Xiao, Deyi Xiong |
EMNLP | 2 |
| 2013 | Lexical Chain Based Cohesion Models for Document-Level Statistical Machine TranslationabstractLexical chains provide a representation of the lexical cohesion structure of a text.In this paper, we propose two lexical chain based cohesion models to incorporate lexical cohesion into document-level statistical machine translation: 1) a count cohesion model that rewards a hypothesis whenever a chain word occurs in the hypothesis, 2) and a probability cohesion model that further takes chain word translation probabilities into account.We compute lexical chains for each source document to be translated and generate target lexical chains based on the computed source chains via maximum entropy classifiers.We then use the generated target chains to provide constraints for word selection in document-level machine translation through the two proposed lexical chain based cohesion models.We verify the effectiveness of the two models using a hierarchical phrase-based translation system.Experiments on large-scale training data show that they can substantially improve translation quality in terms of BLEU and that the probability cohesion model outperforms previous models based on lexical cohesion devices. Deyi Xiong, Min Zhang 0005, Chew Lim Tan |
EMNLP | 1 |
| 2013 | Modeling Lexical Cohesion for Document-Level Machine Translation
Deyi Xiong, Guosheng Ben, Min Zhang 0005, Yajuan Lü, Qun Liu 0001 |
IJCAI | 1 |
| 2012 | A Topic Similarity Model for Hierarchical Phrase-based Translation
Xinyan Xiao, Deyi Xiong, Min Zhang 0005, Qun Liu 0001, Shouxun Lin |
ACL (1) | 2 |
| 2012 | Modeling the Translation of Predicate-Argument Structure for SMT
Deyi Xiong, Min Zhang 0005, Haizhou Li 0001 |
ACL (1) | 1 |
| 2012 | Unsupervised Discriminative Induction of Synchronous Grammar for Machine Translation
Xinyan Xiao, Deyi Xiong, Yang Liu 0005, Qun Liu 0001, Shouxun Lin |
COLING | 2 |
| 2011 | Enhancing Language Models in Statistical Machine Translation with Backward N-grams and Mutual Information Triggers
Deyi Xiong, Min Zhang 0005, Haizhou Li 0001 |
ACL | 1 |
| 2011 | A Maximum-Entropy Segmentation Model for Statistical Machine TranslationabstractSegmentation is of great importance to statistical machine translation. It splits a source sentence into sequences of translatable segments. We propose a maximum-entropy segmentation model to capture desirable phrasal and hierarchical segmentations for statistical machine translation. We present an approach to automatically learning the beginning and ending boundaries of cohesive segments from word-aligned bilingual data without using any additional resources. The learned boundaries are then used to define cohesive segments in both phrasal and hierarchical segmentations. We integrate the segmentation model into phrasal statistical machine translation (SMT) and conduct experiments on the newswire and broadcast news domain to investigate the effectiveness of the proposed segmentation model on a large-scale training data. Our experimental results show that the maximum-entropy segmentation model significantly improves translation quality in terms of BLEU. We further validate that 1) the proposed segmentation model significantly outperforms syntactic constraints which are used in previous work to constrain segmentations; and 2) it is necessary to capture hierarchical segmentations besides phrasal segmentations. Deyi Xiong, Min Zhang 0005, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2010 | Error Detection for Statistical Machine Translation Using Linguistic Features
Deyi Xiong, Min Zhang 0005, Haizhou Li 0001 |
ACL | 1 |
| 2010 | Learning Translation Boundaries for Phrase-Based Decoding
Deyi Xiong, Min Zhang 0005, Haizhou Li 0001 |
HLT-NAACL | 1 |
| 2010 | Linguistically Annotated Reordering: Evaluation and AnalysisabstractLinguistic knowledge plays an important role in phrase movement in statistical machine translation. To efficiently incorporate linguistic knowledge into phrase reordering, we propose a new approach: Linguistically Annotated Reordering (LAR). In LAR, we build hard hierarchical skeletons and inject soft linguistic knowledge from source parse trees to nodes of hard skeletons during translation. The experimental results on large-scale training data show that LAR is comparable to boundary word-based reordering (BWR) (Xiong, Liu, and Lin 2006), which is a very competitive lexicalized reordering approach. When combined with BWR, LAR provides complementary information for phrase reordering, which collectively improves the BLEU score significantly. To further understand the contribution of linguistic knowledge in LAR to phrase reordering, we introduce a syntax-based analysis method to automatically detect constituent movement in both reference and system translations, and summarize syntactic reordering patterns that are captured by reordering models. With the proposed analysis method, we conduct a comparative analysis that not only provides the insight into how linguistic knowledge affects phrase movement but also reveals new challenges in phrase reordering. Deyi Xiong, Min Zhang 0005, AiTi Aw, Haizhou Li 0001 |
Comput. Linguistics | 1 |
| 2009 | A Syntax-Driven Bracketing Model for Phrase-Based Translation
Deyi Xiong, Min Zhang 0005, AiTi Aw, Haizhou Li 0001 |
ACL/IJCNLP | 1 |
| 2008 | Linguistically Annotated BTG for Statistical Machine Translation
Deyi Xiong, Min Zhang 0005, AiTi Aw, Haizhou Li 0001 |
COLING | 1 |
| 2008 | Refinements in BTG-based Statistical Machine Translation
Deyi Xiong, Min Zhang 0005, AiTi Aw, Haitao Mi, Qun Liu 0001, Shouxun Lin |
IJCNLP | 1 |
| 2007 | HTRDP evaluations on Chinese information processing and intelligent human-machine interface
Qun Liu 0001, Hong Liu 0007, Le Sun 0001, Sheng Tang, Deyi Xiong, Hongxu Hou, Yuanhua Lv, Shouxun Lin, Yueliang Qian |
Frontiers Comput. Sci. China | 6 |
| 2006 | Maximum Entropy Based Phrase Reordering Model for Statistical Machine TranslationabstractWe propose a novel reordering model for phrase-based statistical machine translation (SMT) that uses a maximum entropy (MaxEnt) model to predicate reorderings of neighbor blocks (phrase pairs).The model provides content-dependent, hierarchical phrasal reordering with generalization based on features automatically learned from a real-world bitext.We present an algorithm to extract all reordering events of neighbor blocks from bilingual data.In our experiments on Chineseto-English translation, this MaxEnt-based reordering model obtains significant improvements in BLEU score on the NIST MT-05 and IWSLT-04 tasks. Deyi Xiong, Qun Liu 0001, Shouxun Lin |
ACL | 1 |
| 2006 | A Local Density Based Spatial Clustering Algorithm with NoiseabstractDensity-based clustering algorithms are attractive for the task of class identification in spatial database. However, in many cases, very different local-density clusters exist in different regions of data space, therefore, DBSCAN [Ester, M. et al., A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise. In E. Simoudis, J. Han, & U. M. Fayyad (Eds.), Proc. 2nd Int. Conf. on Knowledge Discovery and Data Mining (pp. 226-231). Portland, OR: AAAI.] using a global density parameter is not suitable. As an improvement, OPTICS [Ankerst, M. et al,(1999). OPTICS: Ordering Points To Identify the Clustering Structure. In A. Delis, C. Faloutsos, & S. Ghandeharizadeh (Eds.), Proc. ACM SIGMOD Int. Conf. on Management of Data (pp. 49-60). Philadelphia, PA: ACM.] creates an augmented ordering of the database representing its density-based clustering structure, but it only generates the clusters whose local-density exceeds some threshold instead of similar local-density clusters and doesn't produce a clustering of a data set explicitly. Furthermore the parameters required by almost all the well-known clustering algorithms are hard to determine but have a significant influence on the clustering result. In this paper, a new clustering algorithm LDBSCAN relying on a local-density-based notion of clusters is proposed to solve those problems and, what is more, it is very easy for us to pick the appropriate parameters and takes the advantage of the LOF [Breunig, M. M., et al.,(2000). LOF: Identifying Density-Based Local Outliers. In W. Chen, J. F. Naughton, & P. A. Bernstein (Eds.), Proc. ACM SIGMOD Int. Conf. on Management of Data (pp. 93-104). Dalles, TX: ACM.] to detect the noises comparing with other density-based clustering algorithms. The proposed algorithm has potential applications in business intelligence and enterprise information systems. Deyi Xiong |
SMC | 2 |
| 2005 | Lexicalized Beam Thresholding Parsing with Prior and Boundary Estimates
Deyi Xiong, Qun Liu 0001, Shouxun Lin |
CICLing | 1 |
| 2005 | Parsing the Penn Chinese Treebank with Semantic Knowledge
Deyi Xiong, Shuanglong Li, Qun Liu 0001, Shouxun Lin, Yueliang Qian |
IJCNLP | 1 |
| 2004 | Tagging Complex NEs with MaxEnt Models: Layered Structures Versus Extended Tagset
Deyi Xiong, Hongkui Yu, Qun Liu 0001 |
IJCNLP | 1 |