VLDB 2026 Research / reviewers in the wild / expert
Jiali Zeng
dblp:213/5861
· DBLP profile ↗
29ranked-venue papers
9as first author
22since 2021 · last 2026
0000-0003-0808-9890ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 9 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GRAM-R²: Self-Training Generative Foundation Reward Models for Reward ReasoningabstractMajor progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs to generalist reward models. Despite this trend, developing effective reward models remains a fundamental challenge: the heavy reliance on large-scale labeled preference data. Pre-training on abundant unlabeled data offers a promising direction, but existing approaches fall short in instilling explicit reasoning capabilities into reward models. To bridge this gap, we propose a self-training approach that can leverage unlabeled data to scale up reward reasoning in reward models. Based on this approach, we develop GRAM-R² a generative reward model trained to produce not only preference labels but also accompanying reward rationales. GRAM-R² can serve as a foundation model for reward reasoning and can be applied to a wide range of tasks with minimal or no additional fine-tuning. It can support downstream applications such as policy optimization and task-specific reward tuning. Experiments on response ranking, task adaptation, and reinforcement learning from human feedback demonstrate that GRAM-R² consistently delivers strong performance, outperforming several strong discriminative and generative baselines. Chenglong Wang 0002, Yongyu Mu, Yifu Huo, Jiali Zeng, Murun Yang, Xiaoyang Hao, Chunliang Zhang, Fandong Meng, Tong Xiao 0001 |
AAAI | 6 |
| 2026 | Cross-layer Attention Sharing for Pre-trained Large Language ModelsabstractAbstract To enhance the efficiency of the attention mechanism within large language models (LLMs), previous works primarily compress the Key-Value cache or group attention heads, while largely overlooking redundancy between layers. Our comprehensive analyses across various LLMs show that highly similar attention patterns persist within most layers. It’s intuitive to reduce the redundancy by sharing attention weights across layers. However, further analysis reveals two challenges: (1) Directly sharing the weight matrix without carefully rearranging the attention heads proves to be ineffective; (2) Shallow layers are vulnerable to small deviations in attention weights. Driven by these insights, we introduce LiSA, a lightweight substitute for self-attention in well-trained LLMs. LiSA employs tiny feed-forward networks to align attention heads between adjacent layers and low-rank matrices to approximate differences in layer-wise attention weights. Evaluations encompassing 13 typical benchmarks demonstrate that LiSA maintains high response quality in terms of accuracy and perplexity while reducing redundant attention calculations within 53% −84% of the total layers. Our implementations of LiSA achieve a 6 × compression of Q and K matrices within the attention mechanism, with maximum throughput improvements 19.5%, 32.3%, and 40.1% for LLaMA3-8B, LLaMA2-7B, and LLaMA2-13B, respectively. Our code is available at https://github.com/takagi97/lisa. Yongyu Mu, Yuzhang Wu, Yuchun Fan, Chenglong Wang 0002, Jiali Zeng, Qiaozhi He, Murun Yang, Fandong Meng, Jie Zhou 0016, Tong Xiao 0001 |
Trans. Assoc. Comput. Linguistics | 6 |
| 2025 | ConCISE: Confidence-guided Compression in Step-by-step Efficient ReasoningabstractZiqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang, Lai Wei, Guanbo Wang, Fandong Meng, Jie Zhou, Ju Ren, Yaoxue Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ziqing Qiao, Yongheng Deng, Jiali Zeng, Guanbo Wang, Fandong Meng, Jie Zhou 0016, Ju Ren 0001, Yaoxue Zhang |
EMNLP | 3 |
| 2025 | DelTA: An Online Document-Level Translation Agent Based on Multi-Level MemoryabstractLarge language models (LLMs) have achieved reasonable quality improvements in machine translation (MT).
However, most current research on MT-LLMs still faces significant challenges in maintaining translation consistency and accuracy when processing entire documents.
In this paper, we introduce DelTA, a Document-levEL Translation Agent designed to overcome these limitations.
DelTA features a multi-level memory structure that stores information across various granularities and spans, including Proper Noun Records, Bilingual Summary, Long-Term Memory, and Short-Term Memory, which are continuously retrieved and updated by auxiliary LLM-based components.
Experimental results indicate that DelTA significantly outperforms strong baselines in terms of translation consistency and quality across four open/closed-source LLMs and two representative document translation datasets, achieving an increase in consistency scores by up to 4.58 percentage points and in COMET scores by up to 3.16 points on average.
DelTA employs a sentence-by-sentence translation strategy, ensuring no sentence omissions and offering a memory-efficient solution compared to the mainstream method.
Furthermore, DelTA improves pronoun and context-dependent translation accuracy, and the summary component of the agent also shows promise as a tool for query-based summarization tasks.
The code and data of our approach are released at https://github.com/YutongWang1216/DocMTAgent. Jiali Zeng, Xuebo Liu 0002, Derek F. Wong, Fandong Meng, Jie Zhou 0016, Min Zhang 0005 |
ICLR | 2 |
| 2024 | Generative Multi-Modal Knowledge Retrieval with Large Language ModelsabstractKnowledge retrieval with multi-modal queries plays a crucial role in supporting knowledge-intensive multi-modal applications. However, existing methods face challenges in terms of their effectiveness and training efficiency, especially when it comes to training and integrating multiple retrievers to handle multi-modal queries. In this paper, we propose an innovative end-to-end generative framework for multi-modal knowledge retrieval. Our framework takes advantage of the fact that large language models (LLMs) can effectively serve as virtual knowledge bases, even when trained with limited data. We retrieve knowledge via a two-step process: 1) generating knowledge clues related to the queries, and 2) obtaining the relevant document by searching databases using the knowledge clue. In particular, we first introduce an object-aware prefix-tuning technique to guide multi-grained visual learning. Then, we align multi-grained visual features into the textual feature space of the LLM, employing the LLM to capture cross-modal interactions. Subsequently, we construct instruction data with a unified format for model training. Finally, we propose the knowledge-guided generation strategy to impose prior constraints in the decoding steps, thereby promoting the generation of distinctive knowledge clues. Through experiments conducted on three benchmarks, we demonstrate significant improvements ranging from 3.0% to 14.6% across all evaluation metrics when compared to strong baselines. Xinwei Long, Jiali Zeng, Fandong Meng, Zhiyuan Ma 0005, Bowen Zhou 0002, Jie Zhou 0016 |
AAAI | 2 |
| 2024 | Teaching Large Language Models to Translate with ComparisonabstractOpen-sourced large language models (LLMs) have demonstrated remarkable efficacy in various tasks with instruction tuning. However, these models can sometimes struggle with tasks that require more specialized knowledge such as translation. One possible reason for such deficiency is that instruction tuning aims to generate fluent and coherent text that continues from a given instruction without being constrained by any task-specific requirements. Moreover, it can be more challenging to tune smaller LLMs with lower-quality training data. To address this issue, we propose a novel framework using examples in comparison to teach LLMs to learn translation. Our approach involves output comparison and preference comparison, presenting the model with carefully designed examples of correct and incorrect translations and an additional preference loss for better regularization. Empirical evaluation on four language directions of WMT2022 and FLORES-200 benchmarks shows the superiority of our proposed method over existing methods. Our findings offer a new perspective on fine-tuning LLMs for translation tasks and provide a promising solution for generating high-quality translations. Please refer to Github for more details: https://github.com/lemon0830/TIM. Jiali Zeng, Fandong Meng, Yongjing Yin, Jie Zhou 0016 |
AAAI | 1 |
| 2024 | Tree-of-Reasoning Question Decomposition for Complex Question Answering with Large Language ModelsabstractLarge language models (LLMs) have recently demonstrated remarkable performance across various Natual Language Processing tasks. In the field of multi-hop reasoning, the Chain-of-thought (CoT) prompt method has emerged as a paradigm, using curated stepwise reasoning demonstrations to enhance LLM's ability to reason and produce coherent rational pathways. To ensure the accuracy, reliability, and traceability of the generated answers, many studies have incorporated information retrieval (IR) to provide LLMs with external knowledge. However, existing CoT with IR methods decomposes questions into sub-questions based on a single compositionality type, which limits their effectiveness for questions involving multiple compositionality types. Additionally, these methods suffer from inefficient retrieval, as complex questions often contain abundant information, leading to the retrieval of irrelevant information inconsistent with the query's intent. In this work, we propose a novel question decomposition framework called TRQA for multi-hop question answering, which addresses these limitations. Our framework introduces a reasoning tree (RT) to represent the structure of complex questions. It consists of four components: the Reasoning Tree Constructor (RTC), the Question Generator (QG), the Retrieval and LLM Interaction Module (RAIL), and the Answer Aggregation Module (AAM). Specifically, the RTC predicts diverse sub-question structures to construct the reasoning tree, allowing a more comprehensive representation of complex questions. The QG generates sub-questions for leaf-node in the reasoning tree, and we explore two methods for QG: prompt-based and T5-based approaches. The IR module retrieves documents aligned with sub-questions, while the LLM formulates answers based on the retrieved information. Finally, the AAM aggregates answers along the reason tree, producing a definitive response from bottom to top. Kun Zhang 0041, Jiali Zeng, Fandong Meng, Yuanzhuo Wang, Shiqi Sun 0003, Long Bai 0002, Huawei Shen, Jie Zhou 0016 |
AAAI | 2 |
| 2024 | Understanding and Addressing the Under-Translation Problem from the Perspective of Decoding ObjectiveabstractNeural Machine Translation (NMT) has made remarkable progress over the past years.However, under-translation and over-translation remain two challenging problems in state-of-theart NMT systems.In this work, we conduct an in-depth analysis on the underlying cause of under-translation in NMT, providing an explanation from the perspective of decoding objective.To optimize the beam search objective, the model tends to overlook words it is less confident about, leading to the under-translation phenomenon.Correspondingly, the model's confidence in predicting the End Of Sentence (EOS) diminishes when under-translation occurs, serving as a mild penalty for under-translated candidates.Building upon this analysis, we propose employing the confidence of predicting EOS as a detector for under-translation, and strengthening the confidence-based penalty to penalize candidates with a high risk of under-translation.Experiments on both synthetic and real-world data show that our method can accurately detect and rectify under-translated outputs, with minor impact on other correct translations. Chenze Shao, Fandong Meng, Jiali Zeng, Jie Zhou 0016 |
ACL (1) | 3 |
| 2024 | TasTe: Teaching Large Language Models to Translate through Self-ReflectionabstractLarge language models (LLMs) have exhibited remarkable performance in various natural language processing tasks.Techniques like instruction tuning have effectively enhanced the proficiency of LLMs in the downstream task of machine translation.However, the existing approaches fail to yield satisfactory translation outputs that match the quality of supervised neural machine translation (NMT) systems.One plausible explanation for this discrepancy is that the straightforward prompts employed in these methodologies are unable to fully exploit the acquired instruction-following capabilities.To this end, we propose the TASTE framework, which stands for translating through selfreflection.The self-reflection process includes two stages of inference.In the first stage, LLMs are instructed to generate preliminary translations and conduct self-assessments on these translations simultaneously.In the second stage, LLMs are tasked to refine these preliminary translations according to the evaluation results.The evaluation results in four language directions on the WMT22 benchmark reveal the effectiveness of our approach compared to existing methods.Our work presents a promising approach to unleash the potential of LLMs and enhance their capabilities in MT.The codes and datasets are open-sourced at https://github. com/YutongWang1216/ReflectionLLMMT. Jiali Zeng, Xuebo Liu 0002, Fandong Meng, Jie Zhou 0016, Min Zhang 0005 |
ACL (1) | 2 |
| 2023 | Consistency Regularization Training for Compositional GeneralizationabstractExisting neural models have difficulty generalizing to unseen combinations of seen components.To achieve compositional generalization, models are required to consistently interpret (sub)expressions across contexts.Without modifying model architectures, we improve the capability of Transformer on compositional generalization through consistency regularization training, which promotes representation consistency across samples and prediction consistency for a single sample.Experimental results on semantic parsing and machine translation benchmarks empirically demonstrate the effectiveness and generality of our method.In addition, we find that the prediction consistency scores on in-distribution validation sets can be an alternative for evaluating models during training, when commonly-used metrics are not informative. Yongjing Yin, Jiali Zeng, Yafu Li, Fandong Meng, Jie Zhou 0016, Yue Zhang 0004 |
ACL (1) | 2 |
| 2023 | Soft Language Clustering for Multilingual Model Pre-trainingabstractJiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou 0016 |
ACL (1) | 1 |
| 2023 | Multi-modal graph contrastive encoding for neural machine translation
Yongjing Yin, Jiali Zeng, Jinsong Su, Chulun Zhou, Fandong Meng, Jie Zhou 0016, Degen Huang, Jiebo Luo 0001 |
Artif. Intell. | 2 |
| 2023 | Meta-learning based instance manipulation for implicit discourse relation recognition
Jiali Zeng, Binbin Xie, Changxing Wu, Yongjing Yin, Hualin Zeng, Jinsong Su |
Knowl. Based Syst. | 1 |
| 2022 | Learning Confidence for Transformer-based Neural Machine TranslationabstractConfidence estimation aims to quantify the confidence of the model prediction, providing an expectation of success.A well-calibrated confidence estimate enables accurate failure prediction and proper risk measurement when given noisy samples and out-of-distribution data in real-world settings.However, this task remains a severe challenge for neural machine translation (NMT), where probabilities from softmax distribution fail to describe when the model is probably mistaken.To address this problem, we propose an unsupervised confidence estimate learning jointly with the training of the NMT model.We explain confidence as how many hints the NMT model needs to make a correct prediction, and more hints indicate low confidence.Specifically, the NMT model is given the option to ask for hints to improve translation accuracy at the cost of some slight penalty.Then, we approximate their level of confidence by counting the number of hints the model uses.We demonstrate that our learned confidence estimate achieves high accuracy on extensive sentence/word-level quality estimation tasks.Analytical results verify that our confidence estimate can correctly assess underlying risk in two real-world scenarios: (1) discovering noisy samples and (2) detecting out-of-domain data.We further propose a novel confidence-based instance-specific label smoothing approach based on our learned confidence estimate, which outperforms standard label smoothing 1 . Jiali Zeng, Jiajun Zhang 0001, Shuangzhi Wu, Mu Li 0001 |
ACL (1) | 2 |
| 2022 | An Efficient Coarse-to-Fine Facet-Aware Unsupervised Summarization Framework Based on Semantic BlocksabstractUnsupervised summarization methods have achieved remarkable results by incorporating representations from pre-trained language models. However, existing methods fail to consider efficiency and effectiveness at the same time when the input document is extremely long. To tackle this problem, in this paper, we proposed an efficient Coarse-to-Fine Facet-Aware Ranking (C2F-FAR) framework for unsupervised long document summarization, which is based on the semantic block. The semantic block refers to continuous sentences in the document that describe the same facet. Specifically, we address this problem by converting the one-step ranking method into the hierarchical multi-granularity two-stage ranking. In the coarse-level stage, we proposed a new segment algorithm to split the document into facet-aware semantic blocks and then filter insignificant blocks. In the fine-level stage, we select salient sentences in each block and then extract the final summary from selected sentences. We evaluate our framework on four long document summarization datasets: Gov-Report, BillSum, arXiv, and PubMed. Our C2F-FAR can achieve new state-of-the-art unsupervised summarization results on Gov-Report and BillSum. In addition, our method speeds up 4-28 times more than previous methods. Xinnian Liang, Shuangzhi Wu, Jiali Zeng, Yufan Jiang, Mu Li 0001, Zhoujun Li 0001 |
COLING | 4 |
| 2022 | Attention Analysis and Calibration for Transformer in Natural Language GenerationabstractAttention mechanism has been ubiquitous in neural machine translation by dynamically selecting relevant contexts for different translations. Apart from performance gains, attention weights assigned to input tokens are often utilized to explain that high-attention tokens contribute more to the prediction. However, many works question whether this assumption holds in text classification by manually manipulating attention weights and observing decision flips. This article extends this question to Transformer-based neural machine translation, which heavily relies on cross-lingual attention to produce accurate translations but is relatively understudied in this context. We first design a mask perturbation model which automatically assesses each input’s contribution to model outputs. We then test whether the token contributing most to the current translation receives the highest attention weight. We find that it sometimes does not, which closely depends on the entropy of attention weights, the syntactic role of the current generation, and language pairs. We also rethink the discrepancy between attention weights and word alignments from the view of unreliable attention weights. Our observations further motivate us to calibrate the cross-lingual multi-head attention by attaching more attention to indispensable tokens, whose removal leads to a dramatic performance drop. Empirical experiments on different-scale translation tasks and text summarization tasks demonstrate that our calibration methods significantly outperform strong baselines. Jiajun Zhang 0001, Jiali Zeng, Shuangzhi Wu, Chengqing Zong |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Attention Calibration for Transformer in Neural Machine TranslationabstractYu Lu, Jiali Zeng, Jiajun Zhang, Shuangzhi Wu, Mu Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jiali Zeng, Jiajun Zhang 0001, Shuangzhi Wu, Mu Li 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | Improving Graph-based Sentence Ordering with Iteratively Predicted Pairwise OrderingsabstractShaopeng Lai, Ante Wang, Fandong Meng, Jie Zhou, Yubin Ge, Jiali Zeng, Junfeng Yao, Degen Huang, Jinsong Su. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Shaopeng Lai, Ante Wang, Fandong Meng, Jie Zhou 0016, Yubin Ge, Jiali Zeng, Junfeng Yao, Degen Huang, Jinsong Su |
EMNLP (1) | 6 |
| 2021 | Recurrent Attention for Neural Machine TranslationabstractRecent research questions the importance of the dot-product self-attention in Transformer models and shows that most attention heads learn simple positional patterns.In this paper, we push further in this research line and propose a novel substitute mechanism for self-attention: Recurrent AtteNtion (RAN).RAN directly learns attention weights without any token-to-token interaction and further improves their capacity by layer-to-layer interaction.Across an extensive set of experiments on 10 machine translation tasks, we find that RAN models are competitive and outperform their Transformer counterpart in certain scenarios, with fewer parameters and inference time.Particularly, when apply RAN to the decoder of Transformer, there brings consistent improvements by about +0.5 BLEU on 6 translation tasks and +1.0 BLEU on Turkish-English translation task.In addition, we conduct extensive analysis on the attention weights of RAN to confirm their reasonableness.Our RAN is a promising alternative to build more effective and efficient NMT models. Jiali Zeng, Shuangzhi Wu, Yongjing Yin, Yufan Jiang, Mu Li 0001 |
EMNLP (1) | 1 |
| 2021 | Enhanced Few-Shot Learning with Multiple-Pattern-Exploiting Training
Jiali Zeng, Yufan Jiang, Shuangzhi Wu, Mu Li 0001 |
NLPCC (2) | 1 |
| 2021 | Exploring Discriminative Word-Level Domain Contexts for Multi-Domain Neural Machine TranslationabstractOwing to its practical significance, multi-domain Neural Machine Translation (NMT) has attracted much attention recently. Recent studies mainly focus on constructing a unified NMT model with mixed-domain training corpora to switch translation between different domains. In these models, the words in the same sentence are not well distinguished, while intuitively, they are related to the sentence domain to varying degrees and thus should exert different effects on the multi-domain NMT model. In this article, we are committed to distinguishing and exploiting different word-level domain contexts for multi-domain NMT. For this purpose, we adopt multi-task learning to jointly model NMT and monolingual attention-based domain classification tasks, improving the NMT model in two ways: 1) One domain classifier and one adversarial domain classifier are introduced to conduct domain classifications of input sentences. During this process, two generated gating vectors are used to produce domain-specific and domain-shared annotations for decoder; 2) We equip decoder with an attentional domain classifier. Then, the derived attentional weights are utilized to refine the model training via word-level cost weighting, so that the impacts of target words can be discriminated by their relevance to sentence domain. Experimental results on several multi-domain translations demonstrate the effectiveness of our model. Jinsong Su, Jiali Zeng, Huating Wen, Yongjing Yin, Yang Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Domain Adaptive Meta-Learning for Dialogue State TrackingabstractDomain adaptation for low-resource dialogue state tracking (DST) is of significance due to the growing diversity of conversation scenarios. In this paper, we propose a novel domain adaptive model-agnostic meta-learning (DAMAML) framework. Under this framework, we equip the DST model with two domain adaptors and a unified parameter generator. The parameter generator takes a domain embedding as input to produce parameters of domain adaptors, which modulate domain-shared initial parameters to the subspace of each domain. In this way, we simultaneously model multiple individual meta-learners with each covering the distribution of one domain, allowing more efficient adaptation. Compared with the conventional MAML, this framework not only is able to seek domain-shared initial parameters that facilitate fast adaptation, but also has better capability to fit a diversified domain distribution. Experimental results and in-depth analysis demonstrate the effectiveness of the proposed framework. Jiali Zeng, Yongjing Yin, Yang Liu 0005, Yubin Ge, Jinsong Su |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Neural Simile Recognition with Cyclic Multitask Learning and Local AttentionabstractSimile recognition is to detect simile sentences and to extract simile components, i.e., tenors and vehicles. It involves two subtasks: simile sentence classification and simile component extraction. Recent work has shown that standard multitask learning is effective for Chinese simile recognition, but it is still uncertain whether the mutual effects between the subtasks have been well captured by simple parameter sharing. We propose a novel cyclic multitask learning framework for neural simile recognition, which stacks the subtasks and makes them into a loop by connecting the last to the first. It iteratively performs each subtask, taking the outputs of the previous subtask as additional inputs to the current one, so that the interdependence between the subtasks can be better explored. Extensive experiments show that our framework significantly outperforms the current state-of-the-art model and our carefully designed baselines, and the gains are still remarkable using BERT. Source Code of this paper are available on https://github.com/DeepLearnXMU/Cyclic. Jiali Zeng, Linfeng Song, Jinsong Su, Jiebo Luo 0001 |
AAAI | 1 |
| 2020 | Synonym Knowledge Enhanced Reader for Chinese Idiom Reading ComprehensionabstractMachine reading comprehension (MRC) is the task that asks a machine to answer questions based on a given context. For Chinese MRC, due to the non-literal and non-compositional semantic characteristics, Chinese idioms pose unique challenges for machines to understand. Previous studies tend to treat idioms separately without fully exploiting the relationship among them. In this paper, we first define the concept of literal meaning coverage to measure the consistency between semantics and literal meanings for Chinese idioms. With the definition, we prove that the literal meanings of many idioms are far from their semantics, and we also verify that the synonymic relationship can mitigate this inconsistency, which would be beneficial for idiom comprehension. Furthermore, to fully utilize the synonymic relationship, we propose the synonym knowledge enhanced reader. Specifically, for each idiom, we first construct a synonym graph according to the annotations from the high-quality synonym dictionary or the cosine similarity between the pre-trained idiom embeddings and then incorporate the graph attention network and gate mechanism to encode the graph. Experimental results on ChID, a large-scale Chinese idiom reading comprehension dataset, show that our model achieves state-of-the-art performance. Siyu Long, Ran Wang 0010, Kun Tao, Jiali Zeng, Xinyu Dai |
COLING | 4 |
| 2019 | Iterative Dual Domain Adaptation for Neural Machine TranslationabstractJiali Zeng, Yang Liu, Jinsong Su, Yubing Ge, Yaojie Lu, Yongjing Yin, Jiebo Luo. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jiali Zeng, Yang Liu 0005, Jinsong Su, Yubin Ge, Yaojie Lu 0001, Yongjing Yin, Jiebo Luo 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Graph-based Neural Sentence OrderingabstractSentence ordering is to restore the original paragraph from a set of sentences. It involves capturing global dependencies among sentences regardless of their input order. In this paper, we propose a novel and flexible graph-based neural sentence ordering model, which adopts graph recurrent network \citep{Zhang:acl18} to accurately learn semantic representations of the sentences. Instead of assuming connections between all pairs of input sentences, we use entities that are shared among multiple sentences to make more expressive graph representations with less noise. Experimental results show that our proposed model outperforms the existing state-of-the-art systems on several benchmark datasets, demonstrating the effectiveness of our model. We also conduct a thorough analysis on how entities help the performance. Our code is available at https://github.com/DeepLearnXMU/NSEG.git. Yongjing Yin, Linfeng Song, Jinsong Su, Jiali Zeng, Chulun Zhou, Jiebo Luo 0001 |
IJCAI | 4 |
| 2019 | POS Tag-enhanced Coarse-to-fine Attention for Neural Machine TranslationabstractAlthough neural machine translation (NMT) has certain capability to implicitly learn semantic information of sentences, we explore and show that Part-of-Speech (POS) tags can be explicitly incorporated into the attention mechanism of NMT effectively to yield further improvements. In this article, we propose an NMT model with tag-enhanced attention mechanism. In our model, NMT and POS tagging are jointly modeled via multi-task learning. Besides following common practice to enrich encoder annotations by introducing predicted source POS tags, we exploit predicted target POS tags to refine attention model in a coarse-to-fine manner. Specifically, we first implement a coarse attention operation solely on source annotations and target hidden state, where the produced context vector is applied to update target hidden state used for target POS tagging. Then, we perform a fine attention operation that extends the coarse one by further exploiting the predicted target POS tags. Finally, we facilitate word prediction by simultaneously utilizing the context vector from fine attention and the predicted target POS tags. Experimental results and further analyses on Chinese-English and Japanese-English translation tasks demonstrate the superiority of our proposed model over the conventional NMT models. We release our code at https://github.com/middlekisser/PEA-NMT.git. Yongjing Yin, Jinsong Su, Huating Wen, Jiali Zeng, Yang Liu 0005, Yidong Chen 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2018 | Multi-Domain Neural Machine Translation with Word-Level Domain Context DiscriminationabstractWith great practical value, the study of Multidomain Neural Machine Translation (NMT) mainly focuses on using mixed-domain parallel sentences to construct a unified model that allows translation to switch between different domains.Intuitively, words in a sentence are related to its domain to varying degrees, so that they will exert disparate impacts on the multi-domain NMT modeling.Based on this intuition, in this paper, we devote to distinguishing and exploiting word-level domain contexts for multi-domain NMT.To this end, we jointly model NMT with monolingual attention-based domain classification tasks and improve NMT as follows: 1) Based on the sentence representations produced by a domain classifier and an adversarial domain classifier, we generate two gating vectors and use them to construct domain-specific and domain-shared annotations, for later translation predictions via different attention models; 2) We utilize the attention weights derived from target-side domain classifier to adjust the weights of target words in the training objective, enabling domain-related words to have greater impacts during model training.Experimental results on Chinese-English and English-French multi-domain translation tasks demonstrate the effectiveness of the proposed model.Source codes of this paper are available on Github https://github.com/DeepLearnXMU/WDCNMT. Jiali Zeng, Jinsong Su, Huating Wen, Yang Liu 0005, Yongjing Yin, Jianqiang Zhao |
EMNLP | 1 |
| 2018 | A Hierarchy-to-Sequence Attentional Neural Machine Translation ModelabstractAlthough sequence-to-sequence attentional neural machine translation (NMT) has achieved great progress recently, it is confronted with two challenges: learning optimal model parameters for long parallel sentences and well exploiting different scopes of contexts. In this paper, partially inspired by the idea of segmenting a long sentence into short clauses, each of which can be easily translated by NMT, we propose a hierarchy-to-sequence attentional NMT model to handle these two challenges. Our encoder takes the segmented clause sequence as input and explores a hierarchical neural network structure to model words, clauses, and sentences at different levels, particularly with two layers of recurrent neural networks modeling semantic compositionality at the word and clause level. Correspondingly, the decoder sequentially translates segmented clauses and simultaneously applies two types of attention models to capture contexts of interclause and intraclause for translation prediction. In this way, we can not only improve parameter learning, but also well explore different scopes of contexts for translation. Experimental results on Chinese-English and English-German translation demonstrate the superiorities of the proposed model over the conventional NMT model. Jinsong Su, Jiali Zeng, Deyi Xiong, Yang Liu 0005, Mingxuan Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |