VLDB 2026 Research / reviewers in the wild / expert
Yongjing Yin
dblp:228/5564
· DBLP profile ↗
24ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0003-1138-4612ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 9 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval BenchmarksabstractJunhao Ruan, Abudukeyumu Abudula, Bei Li, Yongjing Yin, Xinyu Liu, Kechen Jiao, Xin Chen, Jingang Wang, Xunliang Cai, Tong Xiao, JingBo Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junhao Ruan, Abudukeyumu Abudula, Yongjing Yin, Kechen Jiao, Jingang Wang, Tong Xiao 0001 |
ACL (1) | 4 |
| 2026 | Heterogeneous Text Style Control Using PromptsabstractAdvancements in natural language processing (NLP) have markedly improved paraphrase generation, an essential task for numerous applications. However, current methods face limitations due to model and constraint specificity, which hinder their flexibility and practical deployment. In this work, we introduce a unified prompt-driven approach to paraphrase generation that leverages diverse prompts, enabling fine-grained user control over aspects such as syntax and sentiment. Moreover, we incorporate translation to enable sophisticated cross-lingual text controls. Our system employs a data-centric paradigm which organizes prompts with natural language instructions. The proposed method is compatible with various sequence-to-sequence architectures and utilizes a novel training strategy to address the versatility of prompt combinations. Empirical results show that our approach not only demonstrates its capacity to adhere to multiple user-defined constraints but also maintains high performance in generation tasks without prompts. Moreover, extensive analysis shows that the model exhibits robustness to prompt variance such as language and quantity. Yafu Li, Jiahao Gai, Yongjing Yin, Jianhao Yan, Yue Zhang 0004 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2025 | Lost in Literalism: How Supervised Training Shapes Translationese in LLMsabstractLarge language models (LLMs) have achieved remarkable success in machine translation, demonstrating impressive performance across diverse languages. However, translationese—characterized by overly literal and unnatural translations—remains a persistent challenge in LLM-based translation systems. Despite their pre-training on vast corpora of natural utterances, LLMs exhibit translationese errors and generate unexpected unnatural translations, stemming from biases introduced during supervised fine-tuning (SFT). In this work, we systematically evaluate the prevalence of translationese in LLM-generated translations and investigate its roots during supervised training. We introduce methods to mitigate these biases, including polishing golden references and filtering unnatural training instances. Empirical evaluations demonstrate that these approaches significantly reduce translationese while improving translation naturalness, validated by human evaluations and automatic metrics. Our findings highlight the need for training-aware adjustments to optimize LLM translation outputs, paving the way for more fluent and target-language-consistent translations. Yafu Li, Ronghao Zhang, Zhilin Wang, Leyang Cui, Yongjing Yin, Tong Xiao 0001, Yue Zhang 0004 |
ACL (1) | 6 |
| 2024 | Teaching Large Language Models to Translate with ComparisonabstractOpen-sourced large language models (LLMs) have demonstrated remarkable efficacy in various tasks with instruction tuning. However, these models can sometimes struggle with tasks that require more specialized knowledge such as translation. One possible reason for such deficiency is that instruction tuning aims to generate fluent and coherent text that continues from a given instruction without being constrained by any task-specific requirements. Moreover, it can be more challenging to tune smaller LLMs with lower-quality training data. To address this issue, we propose a novel framework using examples in comparison to teach LLMs to learn translation. Our approach involves output comparison and preference comparison, presenting the model with carefully designed examples of correct and incorrect translations and an additional preference loss for better regularization. Empirical evaluation on four language directions of WMT2022 and FLORES-200 benchmarks shows the superiority of our proposed method over existing methods. Our findings offer a new perspective on fine-tuning LLMs for translation tasks and provide a promising solution for generating high-quality translations. Please refer to Github for more details: https://github.com/lemon0830/TIM. Jiali Zeng, Fandong Meng, Yongjing Yin, Jie Zhou 0016 |
AAAI | 3 |
| 2024 | Semformer: Transformer Language Models with Semantic PlanningabstractNext-token prediction serves as the dominant component in current neural language models.During the training phase, the model employs teacher forcing, which predicts tokens based on all preceding ground truth tokens.However, this approach has been found to create shortcuts, utilizing the revealed prefix to spuriously fit future tokens, potentially compromising the accuracy of the next-token predictor.In this paper, we introduce Semformer, a novel method of training a Transformer language model that explicitly models the semantic planning of response.Specifically, we incorporate a sequence of planning tokens into the prefix, guiding the planning token representations to predict the latent semantic representations of the response, which are induced by an autoencoder.In a minimal planning task (i.e., graph path-finding), our model exhibits near-perfect performance and effectively mitigates shortcut learning, a feat that standard training methods and baseline models have been unable to accomplish.Furthermore, we pretrain Semformer from scratch with 125M parameters, demonstrating its efficacy through measures of perplexity, in-context learning, and fine-tuning on summarization tasks 1 . Yongjing Yin, Junran Ding, Yue Zhang 0004 |
EMNLP | 1 |
| 2023 | Explicit Syntactic Guidance for Neural Text GenerationabstractMost existing text generation models follow the sequence-to-sequence paradigm.Generative Grammar suggests that humans generate natural language texts by learning language grammar.We propose a syntax-guided generation schema, which generates the sequence guided by a constituency parse tree in a topdown direction.The decoding process can be decomposed into two parts: (1) predicting the infilling texts for each constituent in the lexicalized syntax context given the source sentence;(2) mapping and expanding each constituent to construct the next-level syntax context.Accordingly, we propose a structural beam search method to find possible syntax structures hierarchically.Experiments on paraphrase generation and machine translation show that the proposed method outperforms autoregressive baselines, while also demonstrating effectiveness in terms of interpretability, controllability, and diversity. Yafu Li, Leyang Cui, Jianhao Yan, Yongjing Yin, Wei Bi, Shuming Shi 0001, Yue Zhang 0004 |
ACL (1) | 4 |
| 2023 | Consistency Regularization Training for Compositional GeneralizationabstractExisting neural models have difficulty generalizing to unseen combinations of seen components.To achieve compositional generalization, models are required to consistently interpret (sub)expressions across contexts.Without modifying model architectures, we improve the capability of Transformer on compositional generalization through consistency regularization training, which promotes representation consistency across samples and prediction consistency for a single sample.Experimental results on semantic parsing and machine translation benchmarks empirically demonstrate the effectiveness and generality of our method.In addition, we find that the prediction consistency scores on in-distribution validation sets can be an alternative for evaluating models during training, when commonly-used metrics are not informative. Yongjing Yin, Jiali Zeng, Yafu Li, Fandong Meng, Jie Zhou 0016, Yue Zhang 0004 |
ACL (1) | 1 |
| 2023 | Soft Language Clustering for Multilingual Model Pre-trainingabstractJiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou 0016 |
ACL (1) | 3 |
| 2023 | Multi-modal graph contrastive encoding for neural machine translation
Yongjing Yin, Jiali Zeng, Jinsong Su, Chulun Zhou, Fandong Meng, Jie Zhou 0016, Degen Huang, Jiebo Luo 0001 |
Artif. Intell. | 1 |
| 2023 | Meta-learning based instance manipulation for implicit discourse relation recognition
Jiali Zeng, Binbin Xie, Changxing Wu, Yongjing Yin, Hualin Zeng, Jinsong Su |
Knowl. Based Syst. | 4 |
| 2022 | Categorizing Semantic Representations for Neural Machine TranslationabstractModern neural machine translation (NMT) models have achieved competitive performance in standard benchmarks. However, they have recently been shown to suffer limitation in compositional generalization, failing to effectively learn the translation of atoms (e.g., words) and their semantic composition (e.g., modification) from seen compounds (e.g., phrases), and thus suffering from significantly weakened translation performance on unseen compounds during inference. We address this issue by introducing categorization to the source contextualized representations. The main idea is to enhance generalization by reducing sparsity and overfitting, which is achieved by finding prototypes of token representations over the training set and integrating their embeddings into the source encoding. Experiments on a dedicated MT dataset (i.e., CoGnition) show that our method reduces compositional generalization error rates by 24% error reduction. In addition, our conceptually simple method gives consistently better results than the Transformer baseline on a range of general MT datasets. Yongjing Yin, Yafu Li, Fandong Meng, Jie Zhou 0016, Yue Zhang 0004 |
COLING | 1 |
| 2022 | Multi-Granularity Optimization for Non-Autoregressive TranslationabstractDespite low latency, non-autoregressive machine translation (NAT) suffers severe performance deterioration due to the naive independence assumption.This assumption is further strengthened by cross-entropy loss, which encourages a strict match between the hypothesis and the reference token by token.To alleviate this issue, we propose multi-granularity optimization for NAT, which collects model behaviors on translation segments of various granularities and integrates feedback for backpropagation.Experiments on four WMT benchmarks show that the proposed method significantly outperforms the baseline models trained with cross-entropy loss, and achieves the best performance on WMT'16 En⇔Ro and highly competitive results on WMT'14 En⇔De for fully non-autoregressive translation. Yafu Li, Leyang Cui, Yongjing Yin, Yue Zhang 0004 |
EMNLP | 3 |
| 2021 | On Compositional Generalization of Neural Machine TranslationabstractYafu Li, Yongjing Yin, Yulong Chen, Yue Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yafu Li, Yongjing Yin, Yulong Chen 0001, Yue Zhang 0004 |
ACL/IJCNLP (1) | 2 |
| 2021 | Recurrent Attention for Neural Machine TranslationabstractRecent research questions the importance of the dot-product self-attention in Transformer models and shows that most attention heads learn simple positional patterns.In this paper, we push further in this research line and propose a novel substitute mechanism for self-attention: Recurrent AtteNtion (RAN).RAN directly learns attention weights without any token-to-token interaction and further improves their capacity by layer-to-layer interaction.Across an extensive set of experiments on 10 machine translation tasks, we find that RAN models are competitive and outperform their Transformer counterpart in certain scenarios, with fewer parameters and inference time.Particularly, when apply RAN to the decoder of Transformer, there brings consistent improvements by about +0.5 BLEU on 6 translation tasks and +1.0 BLEU on Turkish-English translation task.In addition, we conduct extensive analysis on the attention weights of RAN to confirm their reasonableness.Our RAN is a promising alternative to build more effective and efficient NMT models. Jiali Zeng, Shuangzhi Wu, Yongjing Yin, Yufan Jiang, Mu Li 0001 |
EMNLP (1) | 3 |
| 2021 | An External Knowledge Enhanced Graph-based Neural Network for Sentence OrderingabstractAs an important text coherence modeling task, sentence ordering aims to coherently organize a given set of unordered sentences. To achieve this goal, the most important step is to effectively capture and exploit global dependencies among these sentences. In this paper, we propose a novel and flexible external knowledge enhanced graph-based neural network for sentence ordering. Specifically, we first represent the input sentences as a graph, where various kinds of relations (i.e., entity-entity, sentence-sentence and entity-sentence) are exploited to make the graph representation more expressive and less noisy. Then, we introduce graph recurrent network to learn semantic representations of the sentences. To demonstrate the effectiveness of our model, we conduct experiments on several benchmark datasets. The experimental results and in-depth analysis show our model significantly outperforms the existing state-of-the-art models. Yongjing Yin, Shaopeng Lai, Linfeng Song, Chulun Zhou, Xianpei Han, Junfeng Yao, Jinsong Su |
J. Artif. Intell. Res. | 1 |
| 2021 | Exploring Discriminative Word-Level Domain Contexts for Multi-Domain Neural Machine TranslationabstractOwing to its practical significance, multi-domain Neural Machine Translation (NMT) has attracted much attention recently. Recent studies mainly focus on constructing a unified NMT model with mixed-domain training corpora to switch translation between different domains. In these models, the words in the same sentence are not well distinguished, while intuitively, they are related to the sentence domain to varying degrees and thus should exert different effects on the multi-domain NMT model. In this article, we are committed to distinguishing and exploiting different word-level domain contexts for multi-domain NMT. For this purpose, we adopt multi-task learning to jointly model NMT and monolingual attention-based domain classification tasks, improving the NMT model in two ways: 1) One domain classifier and one adversarial domain classifier are introduced to conduct domain classifications of input sentences. During this process, two generated gating vectors are used to produce domain-specific and domain-shared annotations for decoder; 2) We equip decoder with an attentional domain classifier. Then, the derived attentional weights are utilized to refine the model training via word-level cost weighting, so that the impacts of target words can be discriminated by their relevance to sentence domain. Experimental results on several multi-domain translations demonstrate the effectiveness of our model. Jinsong Su, Jiali Zeng, Huating Wen, Yongjing Yin, Yang Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Domain Adaptive Meta-Learning for Dialogue State TrackingabstractDomain adaptation for low-resource dialogue state tracking (DST) is of significance due to the growing diversity of conversation scenarios. In this paper, we propose a novel domain adaptive model-agnostic meta-learning (DAMAML) framework. Under this framework, we equip the DST model with two domain adaptors and a unified parameter generator. The parameter generator takes a domain embedding as input to produce parameters of domain adaptors, which modulate domain-shared initial parameters to the subspace of each domain. In this way, we simultaneously model multiple individual meta-learners with each covering the distribution of one domain, allowing more efficient adaptation. Compared with the conventional MAML, this framework not only is able to seek domain-shared initial parameters that facilitate fast adaptation, but also has better capability to fit a diversified domain distribution. Experimental results and in-depth analysis demonstrate the effectiveness of the proposed framework. Jiali Zeng, Yongjing Yin, Yang Liu 0005, Yubin Ge, Jinsong Su |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Enhancing Pointer Network for Sentence Ordering with Pairwise Ordering PredictionsabstractDominant sentence ordering models use a pointer network decoder to generate ordering sequences in a left-to-right fashion. However, such a decoder only exploits the noisy left-side encoded context, which is insufficient to ensure correct sentence ordering. To address this deficiency, we propose to enhance the pointer network decoder by using two pairwise ordering prediction modules: The FUTURE module predicts the relative orientations of other unordered sentences with respect to the candidate sentence, and the HISTORY module measures the local coherence between several (e.g., 2) previously ordered sentences and the candidate sentence, without the influence of noisy left-side context. Using the pointer mechanism, we then incorporate this dynamically generated information into the decoder as a supplement to the left-side context for better predictions. On several commonly-used datasets, our model significantly outperforms other baselines, achieving the state-of-the-art performance. Further analyses verify that pairwise ordering predictions indeed provide extra useful context as expected, leading to better sentence ordering. We also evaluate our sentence ordering models on a downstream task, multi-document summarization, and the summaries reordered by our model achieve the best coherence scores. Our code is available at https://github.com/DeepLearnXMU/Pairwise.git. Yongjing Yin, Fandong Meng, Jinsong Su, Yubin Ge, Linfeng Song, Jie Zhou 0016, Jiebo Luo 0001 |
AAAI | 1 |
| 2020 | A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationabstractMulti-modal neural machine translation (NMT) aims to translate source sentences into a target language paired with images. However, dominant multi-modal NMT models do not fully exploit fine-grained semantic correspondences between semantic units of different modalities, which have potential to refine multi-modal representation learning. To deal with this issue, in this paper, we propose a novel graph-based multi-modal fusion encoder for NMT. Specifically, we first represent the input sentence and image using a unified multi-modal graph, which captures various semantic relationships between multi-modal semantic units (words and visual objects). We then stack multiple graph-based multi-modal fusion layers that iteratively perform semantic interactions to learn node representations. Finally, these representations provide an attention-based context vector for the decoder. We evaluate our proposed encoder on the Multi30K datasets. Experimental results and in-depth analysis show the superiority of our multi-modal NMT model. Yongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou, Zhengyuan Yang, Jie Zhou 0016, Jiebo Luo 0001 |
ACL | 1 |
| 2020 | Dynamic Context-guided Capsule Network for Multimodal Machine TranslationabstractMultimodal machine translation (MMT), which mainly focuses on enhancing text-only translation with visual features, has attracted considerable attention from both computer vision and natural language processing communities. Most current MMT models resort to attention mechanism, global context modeling or multimodal joint representation learning to utilize visual features. However, the attention mechanism lacks sufficient semantic interactions between modalities while the other two provide fixed visual context, which is unsuitable for modeling the observed variability when generating translation. To address the above issues, in this paper, we propose a novel Dynamic Context-guided Capsule Network (DCCN) for MMT. Specifically, at each timestep of decoding, we first employ the conventional source-target attention to produce a timestep-specific source-side context vector. Next, DCCN takes this vector as input and uses it to guide the iterative extraction of related visual features via a context-guided dynamic routing mechanism. Particularly, we represent the input image with global and regional visual features, we introduce two parallel DCCNs to model multimodal context vectors with visual features at different granularities. Finally, we obtain two multimodal context vectors, which are fused and incorporated into the decoder for the prediction of the target word. Experimental results on the Multi30K dataset of English-to-German and English-to-French translation demonstrate the superiority of DCCN. Our code is available on https://github.com/DeepLearnXMU/MM-DCCN. Fandong Meng, Jinsong Su, Yongjing Yin, Zhengyuan Yang, Yubin Ge, Jie Zhou 0016, Jiebo Luo 0001 |
ACM Multimedia | 4 |
| 2019 | Iterative Dual Domain Adaptation for Neural Machine TranslationabstractJiali Zeng, Yang Liu, Jinsong Su, Yubing Ge, Yaojie Lu, Yongjing Yin, Jiebo Luo. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jiali Zeng, Yang Liu 0005, Jinsong Su, Yubin Ge, Yaojie Lu 0001, Yongjing Yin, Jiebo Luo 0001 |
EMNLP/IJCNLP (1) | 6 |
| 2019 | Graph-based Neural Sentence OrderingabstractSentence ordering is to restore the original paragraph from a set of sentences. It involves capturing global dependencies among sentences regardless of their input order. In this paper, we propose a novel and flexible graph-based neural sentence ordering model, which adopts graph recurrent network \citep{Zhang:acl18} to accurately learn semantic representations of the sentences. Instead of assuming connections between all pairs of input sentences, we use entities that are shared among multiple sentences to make more expressive graph representations with less noise. Experimental results show that our proposed model outperforms the existing state-of-the-art systems on several benchmark datasets, demonstrating the effectiveness of our model. We also conduct a thorough analysis on how entities help the performance. Our code is available at https://github.com/DeepLearnXMU/NSEG.git. Yongjing Yin, Linfeng Song, Jinsong Su, Jiali Zeng, Chulun Zhou, Jiebo Luo 0001 |
IJCAI | 1 |
| 2019 | POS Tag-enhanced Coarse-to-fine Attention for Neural Machine TranslationabstractAlthough neural machine translation (NMT) has certain capability to implicitly learn semantic information of sentences, we explore and show that Part-of-Speech (POS) tags can be explicitly incorporated into the attention mechanism of NMT effectively to yield further improvements. In this article, we propose an NMT model with tag-enhanced attention mechanism. In our model, NMT and POS tagging are jointly modeled via multi-task learning. Besides following common practice to enrich encoder annotations by introducing predicted source POS tags, we exploit predicted target POS tags to refine attention model in a coarse-to-fine manner. Specifically, we first implement a coarse attention operation solely on source annotations and target hidden state, where the produced context vector is applied to update target hidden state used for target POS tagging. Then, we perform a fine attention operation that extends the coarse one by further exploiting the predicted target POS tags. Finally, we facilitate word prediction by simultaneously utilizing the context vector from fine attention and the predicted target POS tags. Experimental results and further analyses on Chinese-English and Japanese-English translation tasks demonstrate the superiority of our proposed model over the conventional NMT models. We release our code at https://github.com/middlekisser/PEA-NMT.git. Yongjing Yin, Jinsong Su, Huating Wen, Jiali Zeng, Yang Liu 0005, Yidong Chen 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2018 | Multi-Domain Neural Machine Translation with Word-Level Domain Context DiscriminationabstractWith great practical value, the study of Multidomain Neural Machine Translation (NMT) mainly focuses on using mixed-domain parallel sentences to construct a unified model that allows translation to switch between different domains.Intuitively, words in a sentence are related to its domain to varying degrees, so that they will exert disparate impacts on the multi-domain NMT modeling.Based on this intuition, in this paper, we devote to distinguishing and exploiting word-level domain contexts for multi-domain NMT.To this end, we jointly model NMT with monolingual attention-based domain classification tasks and improve NMT as follows: 1) Based on the sentence representations produced by a domain classifier and an adversarial domain classifier, we generate two gating vectors and use them to construct domain-specific and domain-shared annotations, for later translation predictions via different attention models; 2) We utilize the attention weights derived from target-side domain classifier to adjust the weights of target words in the training objective, enabling domain-related words to have greater impacts during model training.Experimental results on Chinese-English and English-French multi-domain translation tasks demonstrate the effectiveness of the proposed model.Source codes of this paper are available on Github https://github.com/DeepLearnXMU/WDCNMT. Jiali Zeng, Jinsong Su, Huating Wen, Yang Liu 0005, Yongjing Yin, Jianqiang Zhao |
EMNLP | 6 |