VLDB 2026 Research / reviewers in the wild / expert
Xinglin Lyu
dblp:305/6466
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-1971-6618ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Machine translation · 61% Language models and text generation · 22% Deep learning architectures and training · 18% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation › document-level machine translation
document-level neural machine translation |
1.1 | 2 | 2022 | Modeling Consistency Preference via Lexical Chains for Document-level Neural Machine Translation · EMNLP 2022 Encouraging Lexical Translation Consistency for Document-Level Neural Machine Translation · EMNLP (1) 2021 |
Machine learning › Deep learning architectures and training
data augmentation |
1.0 | 1 | 2026 | Paraphrasing as Zero-shot Translation with Feature-guided Diversity Enhancement · ACL (1) 2026 |
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation |
1.0 | 1 | 2026 | Paraphrasing as Zero-shot Translation with Feature-guided Diversity Enhancement · ACL (1) 2026 |
Machine learning › Deep learning architectures and training › data augmentation
paraphrase augmentation |
1.0 | 1 | 2026 | Paraphrasing as Zero-shot Translation with Feature-guided Diversity Enhancement · ACL (1) 2026 |
Natural language and speech › Language models and text generation › text generation
paraphrase generation |
1.0 | 1 | 2026 | Paraphrasing as Zero-shot Translation with Feature-guided Diversity Enhancement · ACL (1) 2026 |
Natural language and speech › Machine translation
document-level machine translation |
0.9 | 1 | 2025 | Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement · ACL (1) 2025 |
Natural language and speech › Machine translation
translation refinement |
0.9 | 1 | 2025 | Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement · ACL (1) 2025 |
Natural language and speech › Machine translation › neural machine translation
context-aware neural machine translation |
0.8 | 1 | 2024 | DeMPT: Decoding-enhanced Multi-phase Prompt Tuning for Making LLMs Be Better Context-aware Translators · EMNLP 2024 |
Natural language and speech › Language models and text generation
prompt tuning |
0.8 | 1 | 2024 | DeMPT: Decoding-enhanced Multi-phase Prompt Tuning for Making LLMs Be Better Context-aware Translators · EMNLP 2024 |
Natural language and speech › Language models and text generation
decoding |
0.7 | 1 | 2023 | Refining History for Future-Aware Neural Machine Translation · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Natural language and speech › Machine translation
neural machine translation |
0.7 | 1 | 2023 | Refining History for Future-Aware Neural Machine Translation · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Natural language and speech › Machine translation
low-resource machine translation |
0.3 | 1 | 2026 | Paraphrasing as Zero-shot Translation with Feature-guided Diversity Enhancement · ACL (1) 2026 |
Natural language and speech › Machine translation › neural machine translation
training-inference discrepancy |
0.2 | 1 | 2023 | Refining History for Future-Aware Neural Machine Translation · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Methods — techniques the papers use, named apart from their topics
feature-guided diversity enhancement · 1.0copy tags · 1.0self-refinement · 0.9fine-tuning · 0.9multi-phase prompt tuning · 0.8decoding enhancement · 0.8transformer · 0.7history-refining module · 0.7future-foreseeing module · 0.7attention mechanism · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementabstractParaphrasing uses different words, sentence structures, or expressions to convey similar semantics.It is an effective training data augmentation method to improve low-resource Natural Language Processing (NLP) tasks.Existing studies normally leverage parallel corpora to construct parabanks, regarding the Machine Translation (MT) results of source sentences as the paraphrases of the corresponding target sentences.As MT models are usually trained on the same parallel corpus, translation of the training set may suffer from overfitting, which leads to less diverse paraphrases.Training paraphrasers on the parabank generated via MT may also suffer from the information loss issue, as the parabank is derived from the parallel corpora, and the knowledge inside the parabank is a subset of that inside the parallel corpora.In this paper, we train bidirectional Multilingual Neural Machine Translation (MNMT) on the bi-directional bilingual parallel corpus, and use the MNMT model directly as a paraphrasing model by asking it to generate "translations" of the input language.As some source tokens also appear in the translation in the parallel corpus, we introduce "copy"/"not-copy" tags to indicate the existence/non-existence of source tokens in the target translation during training, and use the "not-copy" tag to encourage paraphrasing during inference.Manual and automatic evaluation results show that our ParaMNMT method can generate paraphrases of higher semantic consistency, literal fluency and sentential diversity compared to existing parabanks and LLMs.Our data augmentation experiments verify the effectiveness of ParaM-NMT on improving low-resource NLP tasks. Ziyue Yan, Hongying Zan, Xinglin Lyu, Hongfei Xu |
ACL (1) | 3 |
| 2026 | MoMKE-DIR: A multimodal sentiment analysis model based on dynamic feature integration and iterative refinement
Quanyi Wang, Xinglin Lyu, Xijie Cheng, Chunzhi Xie, Jia Liu 0033, Zhisheng Gao |
Expert Syst. Appl. | 2 |
| 2026 | Why not transform chat large language models to non-English?
Xiang Geng, Ming Zhu 0010, Jiahuan Li, Zhejian Lai, Shuaijie She, Yinglu Li, Yuang Li, Chang Su 0001, Xinglin Lyu, Min Zhang 0042, Jiajun Chen 0001, Hao Yang 0006, Shujian Huang |
Frontiers Comput. Sci. | 13 |
| 2026 | Multiphase and Multitask Prompt Tuning for LLM-Based Context-Aware Machine TranslationabstractLarge language models (LLMs) are typically adapted for context-aware machine translation (MT) by combining both the source sentence and its surrounding sentences into a single input. This unified input is then processed in one go, with the model producing the target translation step by step. However, this method treats the intrasentence and intersentence contexts similarly, even though they play distinct roles. In this study, we present a novel strategy called multiphase prompt tuning (MPT) to address this issue by enabling LLMs to treat these two context types differently. MPT divides the context-aware translation task into three phases: encoding the intersentence context, encoding the source sentence, and the final decoding phase. Each phase incorporates distinct continuous prompts that help the model focus on the appropriate task for each type of context. We also introduce a multitask fine-tuning approach to emphasize the distinction between intersentence and intrasentence contexts and enhance intersentence dependencies. This includes two auxiliary tasks: context-agnostic translation and cross-lingual next sentence generation, which help extract additional information and improve the model's handling of discourse-related challenges. Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006, Min Zhang 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation RefinementabstractRecent research has shown that large language models (LLMs) can enhance translation quality through self-refinement. In this paper, we build on this idea by extending the refinement from sentence-level to document-level translation, specifically focusing on document-to-document (Doc2Doc) translation refinement. Since sentence-to-sentence (Sent2Sent) and Doc2Doc translation address different aspects of the translation process, we propose fine-tuning LLMs for translation refinement using two intermediate translations, combining the strengths of both Sent2Sent and Doc2Doc. Additionally, recognizing that the quality of intermediate translations varies, we introduce an enhanced fine-tuning method with quality awareness that assigns lower weights to easier translations and higher weights to more difficult ones, enabling the model to focus on challenging translation cases. Experimental results across ten translation tasks with LLaMA-3-8B-Instruct and Mistral-Nemo-Instruct demonstrate the effectiveness of our approach. We will release our code on GitHub. Yichen Dong, Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
ACL (1) | 2 |
| 2025 | Improving LLM-Based Document-Level MT with Multi-Knowledge Fusion
Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
NLPCC (3) | 2 |
| 2024 | DeMPT: Decoding-enhanced Multi-phase Prompt Tuning for Making LLMs Be Better Context-aware TranslatorsabstractGenerally, the decoder-only large language models (LLMs) are adapted to context-aware neural machine translation (NMT) in a concatenating way, where LLMs take the concatenation of the source sentence (i.e., intrasentence context) and the inter-sentence context as the input, and then to generate the target tokens sequentially.This adaptation strategy, i.e., concatenation mode, considers intrasentence and inter-sentence contexts with the same priority, despite an apparent difference between the two kinds of contexts.In this paper, we propose an alternative adaptation approach, named Decoding-enhanced Multiphase Prompt Tuning (DeMPT), to make LLMs discriminately model and utilize the inter-and intra-sentence context and more effectively adapt LLMs to context-aware NMT.First, DeMPT divides the context-aware NMT process into three separate phases.During each phase, different continuous prompts are introduced to make LLMs discriminately model various information.Second, DeMPT employs a heuristic way to further discriminately enhance the utilization of the source-side interand intra-sentence information at the final decoding phase.Experiments show that our approach significantly outperforms the concatenation method, and further improves the performance of LLMs in discourse modeling. Xinglin Lyu, Junhui Li 0001, Min Zhang 0042, Daimeng Wei, Shimin Tao, Hao Yang 0006, Min Zhang 0005 |
EMNLP | 1 |
| 2023 | Refining History for Future-Aware Neural Machine TranslationabstractNeural machine translation uses a decoder to generate target words auto-regressively by predicting the next target word conditioned on a given source sentence and its previously predicted target words, i.e, its translation history, which suffers from two limitations: 1) the prediction of next word depends heavily on the quality of its history information. Moreover, the discrepancy between training and inference exacerbates this limitation; 2) this left-to-right decoding way cannot make full use of the target-side future information, which leads to the issue of unbalanced outputs. On the one hand, we alleviate the first limitation with a history-refining module, which learns to examine the quality of each history word by assigning it a confidence score. The confidence score is further used as a gate to control the amount of its word embedding flowing to the decoder. On the other hand, we attack the second limitation with a future-foreseeing module, which learns the distribution of future translation at each decoding time step. More importantly, we further propose refining history for future-aware NMT since the two modules can be closely incorporated as they focus on different kinds of context. Experimental results on various translation tasks with different scaled datasets, including WMT English$\leftrightarrow${German, French, Romanian}, show that our proposed approach achieves significant improvements over strong Transformer-based NMT baselines. Xinglin Lyu, Junhui Li 0001, Min Zhang 0005, Chenchen Ding, Hideki Tanaka, Masao Utiyama |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Modeling Consistency Preference via Lexical Chains for Document-level Neural Machine TranslationabstractIn this paper we aim to relieve the issue of lexical translation inconsistency for documentlevel neural machine translation (NMT) by modeling consistency preference for lexical chains which consist of repeated words in a source-side document and provide a representation of the lexical consistency structure of the document.Specifically, we first propose lexical-consistency attention to capture consistency context among words in the same lexical chains.Then for each lexical chain we define and learn a consistency-tailored latent variable, which will guide the translation of corresponding sentences to enhance lexical translation consistency.Experimental results on Chinese→English and French→English document-level translation tasks show that our approach not only significantly improves translation performance in BLEU, but also substantially alleviates the problem of the lexical translation inconsistency. Xinglin Lyu, Junhui Li 0001, Shimin Tao, Hao Yang 0006, Min Zhang 0005 |
EMNLP | 1 |
| 2021 | Encouraging Lexical Translation Consistency for Document-Level Neural Machine TranslationabstractRecently a number of approaches have been proposed to improve translation performance for document-level neural machine translation (NMT).However, few are focusing on the subject of lexical translation consistency.In this paper we apply "one translation per discourse" in NMT, and aim to encourage lexical translation consistency for document-level NMT.This is done by first obtaining a word link for each source word in a document, which tells the positions where the source word appears at.Then we encourage the translations of those words within a link to be consistent in two ways.On the one hand, when encoding sentences within a document we properly exchange context information of those words.On the other hand, we propose an auxiliary loss function to better constrain that their translations should be consistent.Experimental results on Chinese↔English and English→French translation tasks show that our approach not only achieves state-of-the-art performance in BLEU scores, but also greatly improves lexical translation consistency. Xinglin Lyu, Junhui Li 0001, Zhengxian Gong, Min Zhang 0005 |
EMNLP (1) | 1 |