VLDB 2026 Research / reviewers in the wild / expert
Dong Yu 0003
dblp:71/4598-3
· DBLP profile ↗
17ranked-venue papers
0as first author
9since 2021 · last 2026
0009-0003-5278-8346ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Detection: Evaluating Fallacy Awareness of LLMs in Interactive ScenariosabstractLarge Language Models (LLMs) often fail to recognize fallacious reasoning in real-world interactions, despite strong performance on static fallacy detection tasks. We define this ability as fallacy awareness, the capacity to autonomously perceive and resist fallacies in dynamic, pragmatic contexts. To study this, we introduce ISFallacy, a large-scale Chinese benchmark of 50K interactive scenarios spanning six fallacy types, five social interaction settings, diverse role relationships, and personality traits. We further propose FATE, a two-stage evaluation framework that assesses fallacy awareness without explicit cues, combining natural dialogue responses and reasoning-based decisions. Experiments on five representative LLMs reveal a substantial gap between fallacy classification and awareness, with models particularly vulnerable to emotion-driven fallacies and scenarios involving cooperative or trust-based relationships. Deeper analysis uncovers a cognition–behavior gap and fragile internal representations underlying awareness failures. Our work establishes a foundation for evaluating and enhancing the robustness of LLMs against fallacious reasoning in interactive settings. Conghui Niu, Ningxin Wu, Ziran Zhao, Dong Yu 0003, Chen Kang, Pengyuan Liu 0001 |
ACL (1) | 4 |
| 2026 | SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language ModelsabstractBinxian Su, Haoye Lou, Shucheng Zhu, Weikang Wang, Ying Liu, Dong Yu, Pengyuan Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Binxian Su, Haoye Lou, Shucheng Zhu, Weikang Wang 0003, Dong Yu 0003, Pengyuan Liu 0001 |
ACL (1) | 6 |
| 2023 | Bridging the Gap between Synthetic and Authentic Images for Multimodal Machine TranslationabstractMultimodal machine translation (MMT) simultaneously takes the source sentence and a relevant image as input for translation.Since there is no paired image available for the input sentence in most cases, recent studies suggest utilizing powerful text-to-image generation models to provide image inputs.Nevertheless, synthetic images generated by these models often follow different distributions compared to authentic images.Consequently, using authentic images for training and synthetic images for inference can introduce a distribution shift, resulting in performance degradation during inference.To tackle this challenge, in this paper, we feed synthetic and authentic images to the MMT model, respectively.Then we minimize the gap between the synthetic and authentic images by drawing close the input image representations of the Transformer Encoder and the output distributions of the Transformer Decoder.Therefore, we mitigate the distribution disparity introduced by the synthetic images during inference, thereby freeing the authentic images from the inference process.Experimental results show that our approach achieves state-ofthe-art performance on the Multi30K En-De and En-Fr datasets, while remaining independent of authentic images during inference. Wenyu Guo, Qingkai Fang, Dong Yu 0003, Yang Feng 0004 |
EMNLP | 3 |
| 2022 | From Polarity to Intensity: Mining Morality from Semantic SpaceabstractMost works on computational morality focus on moral polarity recognition, i.e., distinguishing right from wrong. However, a discrete polarity label is not informative enough to reflect morality as it does not contain any degree or intensity information. Existing approaches to compute moral intensity are limited to word-level measurement and heavily rely on human labelling. In this paper, we propose MoralScore, a weakly-supervised framework that can automatically measure moral intensity from text. It only needs moral polarity labels, which are more robust and easier to acquire. Besides, the framework can capture latent moral information not only from words but also from sentence-level semantics which can provide a more comprehensive measurement. To evaluate the performance of our method, we introduce a set of evaluation metrics and conduct extensive experiments. Results show that our method achieves good performance on both automatic and human evaluations. Chunxu Zhao, Pengyuan Liu 0001, Dong Yu 0003 |
COLING | 3 |
| 2022 | Perception and Cognition Matters: A new light on sentiment analysis taskabstractIn this paper, we introduce a new psychology-related sentiment category: cognition-sentiment and perception-sentiment. The motivation comes from the investigation of real-world sentiment datasets where we find that sentiment datasets often get an unequal benefits from same external knowledge. We suggest that the difference can be viewed as a new category basis for sentiment data, with the emphasis on a deeper grasping of sentiment feature from an implicit perspective. To test our idea, we first construct a Chinese sentiment dataset based on our proposed categories. Then we set up three experiments to observe the category performance and verify whether our category is suitable for sentiment analysis. Experimental results demonstrate that our proposed category is feasible for the sentiment analysis. Besides, we further try to exploit how can we use the proposed category to help the sentiment task. Specifically, we suggest a new task: perception-sentiment and cognition-sentiment detection (PCSD). Then we set up a joint learning experimental scene utilize PCSD as an auxiliary task to verify the effect of using the category information. Experiments demonstrate the usefulness of our category. Compare with the current SA, our findings widen previous sentiment analysis studies that ignored the implicit feature and sentiment-related psychology knowledge. Shiya Peng, Dong Yu 0003, Pengyuan Liu 0001 |
IJCNN | 3 |
| 2022 | Contrastive Learning Based Visual Representation Enhancement for Multimodal Machine TranslationabstractMultimodal machine translation (MMT) is a task that incorporates extra image modality with text to translate. Previous works have worked on the interaction between two modalities and investigated the need of visual modality. However, few works focus on the models with better and more effective visual representation as input. We argue that the performance of MMT systems will get improved when better visual representation inputs into the systems. To investigate the thought, we introduce mT-ICL, a multimodal Transformer model with image contrastive learning. The contrastive objective is optimized to enhance the representation ability of the image encoder so that the encoder can generate better and more adaptive visual representation. Experiments show that our mT-ICL significantly outperforms the strong baseline and achieves the new SOTA on most of test sets of English-to-German and English-to-French. Further analysis reveals that visual modality works more than a regularization method under contrastive learning framework. Shike Wang, Wen Zhang 0009, Wenyu Guo, Dong Yu 0003, Pengyuan Liu 0001 |
IJCNN | 4 |
| 2022 | CLGC: A Corpus for Chinese Literary Grace EvaluationabstractIn this paper, we construct a Chinese literary grace corpus, CLGC, with 10,000 texts and more than 1.85 million tokens. Multi-level annotations are provided for each text in our corpus, including literary grace level, sentence category, and figure-of-speech type. Based on the corpus, we dig deep into the correlation between fine-grained features (semantic information, part-of-speech and figure-of-speech, etc.) and literary grace level. We also propose a new Literary Grace Evaluation (LGE) task, which aims at making a comprehensive assessment of the literary grace level according to the text. In the end, we build some classification models with machine learning algorithms (such as SVM, TextCNN) to prove the effectiveness of our features and corpus for LGE. The results of our preliminary classification experiments have achieved 79.71% on the weighted average F1-score. Dong Yu 0003, Pengyuan Liu 0001 |
LREC | 2 |
| 2021 | Importance-based Neuron Allocation for Multilingual Neural Machine TranslationabstractWanying Xie, Yang Feng, Shuhao Gu, Dong Yu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wanying Xie, Yang Feng 0004, Shuhao Gu, Dong Yu 0003 |
ACL/IJCNLP (1) | 4 |
| 2021 | ExperienceGen 1.0: A Text Generation Challenge Which Requires Deduction and Induction Ability
Pengyuan Liu 0001, Dong Yu 0003, Sanle Zhang |
NLPCC (2) | 3 |
| 2020 | Modeling Fluency and Faithfulness for Diverse Neural Machine TranslationabstractNeural machine translation models usually adopt the teacher forcing strategy for training which requires the predicted sequence matches ground truth word by word and forces the probability of each prediction to approach a 0-1 distribution. However, the strategy casts all the portion of the distribution to the ground truth word and ignores other words in the target vocabulary even when the ground truth word cannot dominate the distribution. To address the problem of teacher forcing, we propose a method to introduce an evaluation module to guide the distribution of the prediction. The evaluation module accesses each prediction from the perspectives of fluency and faithfulness to encourage the model to generate the word which has a fluent connection with its past and future translation and meanwhile tends to form a translation equivalent in meaning to the source. The experiments on multiple translation tasks show that our method can achieve significant improvements over strong baselines. Yang Feng 0004, Wanying Xie, Shuhao Gu, Chenze Shao, Wen Zhang 0009, Zhengxin Yang, Dong Yu 0003 |
AAAI | 7 |
| 2020 | Token-level Adaptive Training for Neural Machine TranslationabstractThere exists a token imbalance phenomenon in natural language as different tokens appear with different frequencies, which leads to different learning difficulties for tokens in Neural Machine Translation (NMT).The vanilla NMT model usually adopts trivial equal-weighted objectives for target tokens with different frequencies and tends to generate more high-frequency tokens and less lowfrequency tokens compared with the golden token distribution.However, low-frequency tokens may carry critical semantic information that will affect the translation quality once they are neglected.In this paper, we explored target token-level adaptive objectives based on token frequencies to assign appropriate weights for each target token during training.We aimed that those meaningful but relatively low-frequency words could be assigned with larger weights in objectives to encourage the model to pay more attention to these tokens.Our method yields consistent improvements in translation quality on ZH-EN, EN-RO, and EN-DE translation tasks, especially on sentences that contain more low-frequency tokens where we can get 1.68, 1.02, and 0.52 BLEU increases compared with baseline, respectively.Further analyses show that our method can also improve the lexical diversity of translation. Shuhao Gu, Jinchao Zhang 0001, Fandong Meng, Yang Feng 0004, Wanying Xie, Jie Zhou 0016, Dong Yu 0003 |
EMNLP (1) | 7 |
| 2020 | Task-to-Task Transfer Learning with Parameter-Efficient Adapter
Haiou Zhang, Hanjun Zhao, Chunhua Liu, Dong Yu 0003 |
NLPCC (2) | 4 |
| 2018 | Multi-turn Inference Matching Network for Natural Language Inference
Chunhua Liu, Hainan Yu, Dong Yu 0003 |
NLPCC (2) | 4 |
| 2018 | From Plots to Endings: A Reinforced Pointer Generator for Story Ending Generation
Yan Zhao 0020, Lu Liu 0008, Chunhua Liu, Ruoyao Yang, Dong Yu 0003 |
NLPCC (1) | 5 |
| 2018 | DEMN: Distilled-Exposition Enhanced Matching Network for Story Comprehension
Chunhua Liu, Haiou Zhang, Dong Yu 0003 |
PACLIC | 4 |
| 2015 | A Hybrid Re-ranking Method for Entity Recognition and Linking in Search QueriesabstractIn this paper, we construct an entity recognition and linking system using Chinese Wikipedia and knowledge base. We utilize refined filter rules in entity recognition module, and then generate candidate entities by search engine and attributes in Wikipedia article pages. In entity linking module, we propose a hybrid entity re-ranking method combined with three features: textual and semantic match-degree, the similarity between candidate entity and entity mention, entity frequency. Finally, we get the linking results by the entity’s final score. In the task of entity recognition and linking in search queries at NLPCC 2015, the Average-F1 value of this method achieved 61.1% in 3849 test dataset, which ranks second place in fourteen teams. Gongbo Tang, Dong Yu 0003, Endong Xun |
NLPCC | 3 |
| 2014 | Chinese Microblog Entity Linking System Combining Wikipedia and Search Engine Retrieval Results
Zeyu Meng, Dong Yu 0003, Endong Xun |
NLPCC | 2 |