VLDB 2026 Research / reviewers in the wild / expert
Daimeng Wei
dblp:166/0470
· DBLP profile ↗
25ranked-venue papers
1as first author
25since 2021 · last 2026
0009-0008-9606-5058ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 1 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PLaST: Towards Paralinguistic-aware Speech TranslationabstractSpeech translation (ST) aims to translate speech from a source language into text in the target language. Naturally, speech signals contain paralinguistic cues beyond linguistic content, which could influence or even alter the interpretation of a lexically identical sentence, thereby yielding distinct translations. However, existing ST models lack direct and sufficient modeling of paralinguistic information, which limits their ability to perceive paralinguistic cues and understand speech comprehensively, leading to degraded translation performance. In response, we propose Paralinguistic-aware Speech Translation (PLaST), a novel dual-branch framework which directly leverages paralinguistic cues beyond the linguistic content. Specifically, PLaST employs a speech encoder and a style extractor to independently generate linguistic and paralinguistic representations, respectively. To obtain a purified linguistic representation aligned with the text representation, a hierarchical Optimal Transport (OT) is applied on the layer-wise outputs from an LLM decoder. Then, the paralinguistic information is retrieved and refined with an Attention-based Retrieval (AR) module, with the linguistic representation serving as queries to enable joint guidance for semantic understanding and translation generation. PLaST outperforms the strong baseline with an average of 5.0 directional and 4.5 global contrastive likelihood scores on the paralinguistic-sensitive benchmark ContraProST, demonstrating its superior capability in paralinguistic perception. Further experiments on the standard speech translation benchmark CoVoST-2 show that PLaST generalizes well to typical ST scenarios. Ruiquan Zhang, Jinsong Su, Daimeng Wei, Min Zhang 0042, Yidong Chen 0001 |
AAAI | 5 |
| 2026 | MIDB: Multilingual Instruction Data Booster for Enhancing Cultural Equality in Multilingual Instruction SynthesisabstractDespite doubts on data quality, instruction synthesis has been widely applied into instruction tuning (IT) of LLMs as an economic and rapid alternative. Recent endeavors focus on improving data quality for synthesized instruction pairs in English and have facilitated IT of English-centric LLMs. However, data quality issues in multilingual synthesized instruction pairs are even more severe, since the common synthesizing practice is to translate English synthesized data into other languages using machine translation (MT). Besides the known content errors in these English synthesized data, multilingual synthesized instruction data are further exposed to defects introduced by MT and face insufficient localization of the target languages, leading to cultural inequality in trained LLMs. In this paper, we propose MIDB, a Multilingual Instruction Data Booster to automatically address the quality issues in multilingual synthesized data. MIDB is trained on around 36.8k revision examples across 16 languages by human linguistic experts, thereby can boost the low-quality data by addressing content errors and MT defects, and improving localization in these synthesized data. Both automatic and human evaluation indicate that not only MIDB steadily improved instruction data quality in 16 languages, but also the instruction-following and cultural-understanding abilities of multilingual LLMs fine-tuned on MIDB-boosted data were significantly enhanced, suggesting an improved linguistic and cultural equality. Yilun Liu 0001, Chunguang Zhao, Xinhua Yang, Hongyong Zeng, Shimin Tao, Weibin Meng, Minggui He, Hongxia Ma, Daimeng Wei, Boxing Chen |
AAAI | 11 |
| 2026 | ELSPR: Evaluator LLM Training Data Self-Purification on Non-Transitive Preferences via Tournament Graph ReconstructionabstractPairwise evaluation of large language models (LLMs) has become the dominant paradigm for benchmarking open-ended tasks, yet non-transitive preferences—where evaluators prefer A over B, B over C, but C over A—fundamentally undermine ranking reliability. We show that this critical issue stems largely from low-quality data that contains inherently ambiguous preference pairs. To address this challenge, we propose ELSPR, a principled graph-theoretic framework that models pairwise preferences as tournament graphs and systematically identifies problematic training data. ELSPR quantifies non-transitivity through strongly connected components (SCCs) analysis and measures overall preference clarity using a novel normalized directed graph structural entropy metric. Our filtering methodology selectively removes preference data that induce non-transitivity while preserving transitive preferences. Extensive experiments on the AlpacaEval benchmark demonstrate that models fine-tuned on ELSPR-filtered data achieve substantial improvements: a 13.8% reduction in non-transitivity, a 0.088 decrease in structural entropy, and significantly enhanced discriminative power in real-world evaluation systems. Human validation confirms that discarded data exhibit dramatically lower inter-annotator agreement (34.4% vs. 52.6%) and model-human consistency (51.2% vs. 80.6%) compared to cleaned data. These findings establish ELSPR as an effective data self-purification approach for developing more robust, consistent, and human-aligned LLM evaluation systems. Yilun Liu 0001, Minggui He, Shimin Tao, Weibin Meng, Xinhua Yang, Hongxia Ma, Dengye Li, Daimeng Wei, Boxing Chen, Fuliang Li |
AAAI | 10 |
| 2026 | DeReA: Improving Idiom Translation with Detect-Retrieve-Arbitrate ReasoningabstractRongqing Jiang, Xuebo Liu, Shengxin Liu, Yutong Wang, Min Zhang, Shimin Tao, Daimeng Wei, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rongqing Jiang, Xuebo Liu 0002, Shengxin Liu, Min Zhang 0005, Shimin Tao, Daimeng Wei |
ACL (1) | 7 |
| 2026 | The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language ModelsabstractYilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui HE, Chenxin Liu, Zhang Li, Mahongxia, Jiaxin Guo, Chen Liu, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yilun Liu 0001, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui He, Chenxin Liu, Hongxia Ma, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao |
ACL (1) | 16 |
| 2026 | M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning DatasetsabstractMultilingual instruction fine-tuning (IFT) empowers large language models to generalize across diverse linguistic and cultural contexts; however, high-quality, systematically curated multilingual IFT datasets remain scarce. To address this gap, we propose M-DaQ (Multilingual Diversity and Quality), a diversity-aware sampling framework that jointly optimizes instruction-response quality and cross-lingual semantic diversity. M-DaQ leverages a fine-tuned Quality Scoring Model alongside a maximal marginal relevance-inspired selection strategy to construct balanced, high-fidelity training data. Furthermore, we present the first systematic investigation of the Superficial Alignment Hypothesis in multilingual settings. Extensive evaluations across 18 languages demonstrate that models trained on M-DaQ-curated data achieve average win rates exceeding 60% against strong baselines on Alpaca-Eval and MT-Bench. Complementary human evaluations corroborate these gains, highlighting significant improvements in cultural relevance, contextual appropriateness, and instruction-following capability. The code are publicly released to facilitate reproducibility and future research. Chunguang Zhao, Yilun Liu 0001, Pufan Zeng, Yuanchang Luo, Shimin Tao, Minggui He, Weibin Meng, Hongxia Ma, Boxing Chen, Daimeng Wei |
SIGIR | 13 |
| 2026 | Multiphase and Multitask Prompt Tuning for LLM-Based Context-Aware Machine TranslationabstractLarge language models (LLMs) are typically adapted for context-aware machine translation (MT) by combining both the source sentence and its surrounding sentences into a single input. This unified input is then processed in one go, with the model producing the target translation step by step. However, this method treats the intrasentence and intersentence contexts similarly, even though they play distinct roles. In this study, we present a novel strategy called multiphase prompt tuning (MPT) to address this issue by enabling LLMs to treat these two context types differently. MPT divides the context-aware translation task into three phases: encoding the intersentence context, encoding the source sentence, and the final decoding phase. Each phase incorporates distinct continuous prompts that help the model focus on the appropriate task for each type of context. We also introduce a multitask fine-tuning approach to emphasize the distinction between intersentence and intrasentence contexts and enhance intersentence dependencies. This includes two auxiliary tasks: context-agnostic translation and cross-lingual next sentence generation, which help extract additional information and improve the model's handling of discourse-related challenges. Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006, Min Zhang 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation RefinementabstractRecent research has shown that large language models (LLMs) can enhance translation quality through self-refinement. In this paper, we build on this idea by extending the refinement from sentence-level to document-level translation, specifically focusing on document-to-document (Doc2Doc) translation refinement. Since sentence-to-sentence (Sent2Sent) and Doc2Doc translation address different aspects of the translation process, we propose fine-tuning LLMs for translation refinement using two intermediate translations, combining the strengths of both Sent2Sent and Doc2Doc. Additionally, recognizing that the quality of intermediate translations varies, we introduce an enhanced fine-tuning method with quality awareness that assigns lower weights to easier translations and higher weights to more difficult ones, enabling the model to focus on challenging translation cases. Experimental results across ten translation tasks with LLaMA-3-8B-Instruct and Mistral-Nemo-Instruct demonstrate the effectiveness of our approach. We will release our code on GitHub. Yichen Dong, Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
ACL (1) | 4 |
| 2025 | Enhancing Large Language Models for Document-Level Translation Post-Editing Using Monolingual DataabstractThe translation capabilities of neural machine translation (NMT) models based on the encoder-decoder framework are extremely potent. Although Large Language Models (LLMs) have achieved remarkable results in many tasks, they have not reached state-of-the-art performance in NMT. However, traditional NMT still faces significant challenges in areas of document translation such as context consistency, tense, and pronoun resolution, where LLMs inherently possess substantial advantages. Instead of directly using LLMs for translation, employing them for Automatic Post-Editing (APE) to post-edit NMT outputs proves to be a viable option. However, document-level bilingual data is extremely scarce. This paper proposes a method that can effectively leverage the capabilities of LLMs to optimize document translation using only monolingual data. By employing two NMT models in opposite directions (Source-to-Target and Target-to-Source), we generate pseudo-document training data for the training of APE. We have identified and resolved the issue between training and inference mode inconsistency brought about by the pseudo-document training data. The final experimental results demonstrate that by using only document-level monolingual data, we can significantly improve the quality of NMT and greatly enhance issues such as reference and contextual consistency in NMT. Zhiqiang Rao, Hengchao Shang, Daimeng Wei, Hao Yang 0006 |
COLING | 6 |
| 2025 | Generative Annotation for ASR Named Entity CorrectionabstractYuanchang Luo, Daimeng Wei, Shaojun Li, Hengchao Shang, Jiaxin Guo, Zongyao Li, Zhanglin Wu, Xiaoyu Chen, Zhiqiang Rao, Jinlong Yang, Hao Yang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yuanchang Luo, Daimeng Wei, Hengchao Shang, Zhanglin Wu, Xiaoyu Chen 0004, Zhiqiang Rao, Hao Yang 0006 |
EMNLP | 2 |
| 2025 | Doc-Guided Sent2Sent++: A Sent2Sent++ Agent with Doc-Guided Memory for Document-Level Machine Translation
Yuanchang Luo, Daimeng Wei, Hengchao Shang, Zhiqiang Rao, Zhanglin Wu, Hao Yang 0006 |
NLPCC (3) | 3 |
| 2025 | Improving LLM-Based Document-Level MT with Multi-Knowledge Fusion
Xinglin Lyu, Junhui Li 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
NLPCC (3) | 4 |
| 2024 | Cross-Domain Audio Deepfake Detection: Dataset and AnalysisabstractAudio deepfake detection (ADD) is essential for preventing the misuse of synthetic voices that may infringe on personal rights and privacy.Recent zero-shot text-to-speech (TTS) models pose higher risks as they can clone voices with a single utterance.However, the existing ADD datasets are outdated, leading to suboptimal generalization of detection models.In this paper, we construct a new cross-domain ADD dataset comprising over 300 hours of speech data that is generated by five advanced zeroshot TTS models.To simulate real-world scenarios, we employ diverse attack methods and audio prompts from different datasets.Experiments show that, through novel attackaugmented training, the Wav2Vec2-large and Whisper-medium models achieve equal error rates of 4.1% and 6.5% respectively.Additionally, we demonstrate our models' outstanding few-shot ADD ability by fine-tuning with just one minute of target-domain data.Nonetheless, neural codec compressors greatly affect the detection accuracy, necessitating further research.Our dataset is publicly available 1 . Yuang Li, Min Zhang 0042, Mengxin Ren, Xiaosong Qiao, Miaomiao Ma, Daimeng Wei, Hao Yang 0006 |
EMNLP | 6 |
| 2024 | DeMPT: Decoding-enhanced Multi-phase Prompt Tuning for Making LLMs Be Better Context-aware TranslatorsabstractGenerally, the decoder-only large language models (LLMs) are adapted to context-aware neural machine translation (NMT) in a concatenating way, where LLMs take the concatenation of the source sentence (i.e., intrasentence context) and the inter-sentence context as the input, and then to generate the target tokens sequentially.This adaptation strategy, i.e., concatenation mode, considers intrasentence and inter-sentence contexts with the same priority, despite an apparent difference between the two kinds of contexts.In this paper, we propose an alternative adaptation approach, named Decoding-enhanced Multiphase Prompt Tuning (DeMPT), to make LLMs discriminately model and utilize the inter-and intra-sentence context and more effectively adapt LLMs to context-aware NMT.First, DeMPT divides the context-aware NMT process into three separate phases.During each phase, different continuous prompts are introduced to make LLMs discriminately model various information.Second, DeMPT employs a heuristic way to further discriminately enhance the utilization of the source-side interand intra-sentence information at the final decoding phase.Experiments show that our approach significantly outperforms the concatenation method, and further improves the performance of LLMs in discourse modeling. Xinglin Lyu, Junhui Li 0001, Min Zhang 0042, Daimeng Wei, Shimin Tao, Hao Yang 0006, Min Zhang 0005 |
EMNLP | 5 |
| 2024 | CSNet: Contrastive Siamese Network for Robust SLUabstractAutomatic speech recognition (ASR) results based on clean references are much more accurate than those based on ASR transcripts in spoken language understanding (SLU). Effective utilization of manually-checked clean transcripts is key to improving SLU performance. This paper proposes a siamese network with contrastive learning to enhance SLU effects. A siamese network on sentence pairs that are composed of ASR transcripts and clean transcripts is used for the SLU task. During training, contrastive learning brings closer the sentence-level semantic representations of ASR transcripts and clean transcripts. During inference, k-nearest neighbors (KNN) semantic search via the siamese network first finds the pseudo clean transcript, then forms a sentence pair based on the ASR transcript and pseudo clean transcript for prediction. Experiments on three benchmark datasets prove the effectiveness of our proposed approach, which improves the Intent Classification (IC) performance by over 1.3% on the SLURP dataset. Hao Yang 0006, Min Zhang 0042, Daimeng Wei |
ICASSP | 3 |
| 2024 | Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
Daimeng Wei, Hengchao Shang, Zhanglin Wu, Zhiqiang Rao, Yuanchang Luo, Xianghui He, Hao Yang 0006 |
INTERSPEECH | 2 |
| 2024 | A Multitask Training Approach to Enhance Whisper with Open-Vocabulary Keyword SpottingabstractThe recognition of rare named entities, such as personal names and terminologies, is challenging for automatic speech recognition (ASR) systems, especially when they are not frequently observed in the training data.In this paper, we introduce keyword spotting enhanced Whisper (KWS-Whisper), a novel ASR system that leverages the Whisper model and performs openvocabulary keyword spotting (OV-KWS) on the hidden states of the Whisper encoder to recognize user-defined named entities.These entities serve as prompts for the Whisper decoder.To optimize the model, we propose a multitask training approach that learns OV-KWS and contextual-ASR tasks.We evaluate our approach on Chinese Aishell hot word subsets and two internal code-switching test sets and show that it significantly improves the entity recall compared to the original Whisper model.Moreover, we demonstrate that the OV-KWS can be a plug-andplay module to enhance the ASR error correction methods and frozen Whisper models. Yuang Li, Min Zhang 0042, Chang Su 0001, Yinglu Li, Xiaosong Qiao, Mengxin Ren, Miaomiao Ma, Daimeng Wei, Shimin Tao, Hao Yang 0006 |
INTERSPEECH | 8 |
| 2024 | An End-to-End Speech Summarization Using Large Language Model
Hengchao Shang, Zhiqiang Rao, Yuanchang Luo, Daimeng Wei |
INTERSPEECH | 7 |
| 2023 | Text Style Transfer Back-TranslationabstractDaimeng Wei, Zhanglin Wu, Hengchao Shang, Zongyao Li, Minghan Wang, Jiaxin Guo, Xiaoyu Chen, Zhengzhe Yu, Hao Yang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Daimeng Wei, Zhanglin Wu, Hengchao Shang, Minghan Wang, Xiaoyu Chen 0004, Zhengzhe Yu, Hao Yang 0006 |
ACL (1) | 1 |
| 2023 | UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error CorrectionabstractError correction techniques have been used to refine the output sentences from automatic speech recognition (ASR) models and achieve a lower word error rate (WER). Previous works usually adopt end-to-end models and has strong dependency on Pseudo Paired Data and Original Paired Data. But when only pre-training on Pseudo Paired Data, previous models have negative effect on correction. While fine-tuning on Original Paired Data, the source side data must be transcribed by a well-trained ASR model, which takes a lot of time and not universal. In this paper, we propose UCorrect, an unsupervised Detector-Generator-Selector framework for ASR Error Correction. UCorrect has no dependency on the training data mentioned before. The whole procedure is first to detect whether the character is erroneous, then to generate some candidate characters and finally to select the most confident one to replace the error character. Experiments on the public AISHELL-1 dataset and WenetSpeech dataset show the effectiveness of UCorrect for ASR error correction: 1) it achieves significant WER reduction, achieves 6.83% even without fine-tuning and 14.29% after fine-tuning; 2) it outperforms the popular NAR correction models by a large margin with a competitive low latency; and 3) it is an universal method, as it reduces all WERs of the ASR model with different decoding strategies and reduces all WERs of ASR models trained on different scale datasets. Minghan Wang, Xiaosong Qiao, Daimeng Wei, Hengchao Shang, Zhengzhe Yu, Yinglu Li, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
ICASSP | 4 |
| 2023 | WhiSLU: End-to-End Spoken Language Understanding with Whisper
Minghan Wang, Yinglu Li, Xiaosong Qiao, Hengchao Shang, Daimeng Wei, Shimin Tao, Min Zhang 0042, Hao Yang 0006 |
INTERSPEECH | 7 |
| 2022 | Diformer: Directional Transformer for Neural Machine TranslationabstractAutoregressive (AR) and Non-autoregressive (NAR) models have their own superiority on the performance and latency, combining them into one model may take advantage of both. Current combination frameworks focus more on the integration of multiple decoding paradigms with a unified generative model, e.g. Masked Language Model. However, the generalization can be harmful on the performance due to the gap between training objective and inference. In this paper, we aim to close the gap by preserving the original objective of AR and NAR under a unified framework. Specifically, we propose the Directional Transformer (Diformer) by jointly modelling AR and NAR into three generation directions (left-to-right, right-to-left and straight) with a newly introduced direction variable, which works by controlling the prediction of each token to have specific dependencies under that direction. The unification achieved by direction successfully preserves the original dependency assumption used in AR and NAR, retaining both generalization and performance. Experiments on 4 WMT benchmarks demonstrate that Diformer outperforms current united-modelling works with more than 1.5 BLEU points for both AR and NAR decoding, and is also competitive to the state-of-the-art independent AR and NAR models. Minghan Wang, Yuxia Wang 0003, Daimeng Wei, Hengchao Shang, Yinglu Li, Chang Su 0001, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
EAMT | 4 |
| 2022 | CCDC: A Chinese-Centric Cross Domain Contrastive Learning Framework
Hao Yang 0006, Shimin Tao, Minghan Wang, Min Zhang 0042, Daimeng Wei, Shuai Zhao 0001, Miaomiao Ma |
KSEM (2) | 5 |
| 2022 | Augmented Topic-Specific Summarization for Domain Dialogue Text
Zhiqiang Rao, Daimeng Wei, Hengchao Shang, Zhengzhe Yu, Zhanglin Wu, Lizhi Lei, Hao Yang 0006 |
NLPCC (2) | 2 |
| 2021 | HI-CMLM: Improve CMLM with Hybrid Decoder InputabstractMask-predict CMLM (Ghazvininejad et al., 2019) has achieved stunning performance among non-autoregressive NMT models, but we find that the mechanism of predicting all of the target words only depending on the hidden state of [MASK] is not effective and efficient in initial iterations of refinement, resulting in ungrammatical repetitions and slow convergence.In this work, we mitigate this problem by combining copied source with embeddings of [MASK] in decoder.Notably.it's not a straightforward copying that is shown to be useless, but a novel heuristic hybrid strategy -fence-mask.Experimental results show that it gains consistent boosts on both WMT14 En↔De and WMT16 En↔Ro corpus by 0.5 BLEU on average, and 1 BLEU for lessinformative short sentences.This reveals that incorporating additional information by proper strategies is beneficial to improve CMLM, particularly translation quality of short texts and speeding up early-stage convergence. Minghan Wang, Yuxia Wang 0003, Chang Su 0001, Daimeng Wei, Min Zhang 0042, Shimin Tao, Hao Yang 0006 |
INLG | 6 |