VLDB 2026 Research / reviewers in the wild / expert
Yangfan Ye
dblp:343/9917
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0000-6845-1068ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Language models and text generation · 64% Transfer learning and domain adaptation · 10% Efficient and distributed learning · 6% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
multilingual language models |
2.7 | 3 | 2026 | LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction Tuning · AAAI 2026 CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-Tuning · ACL (1) 2025 Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-Tuning · AAAI 2025 |
Machine learning › Efficient and distributed learning
data selection |
1.0 | 1 | 2026 | LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction Tuning · AAAI 2026 |
Natural language and speech › Language models and text generation › instruction tuning
multilingual instruction tuning |
1.0 | 1 | 2026 | LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction Tuning · AAAI 2026 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play |
1.0 | 1 | 2026 | Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play · ACL (1) 2026 |
Natural language and speech › Language models and text generation › hallucination detection
context faithfulness |
0.9 | 1 | 2025 | Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced Optimization · ACL (1) 2025 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.9 | 1 | 2025 | CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-Tuning · ACL (1) 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.9 | 1 | 2025 | Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-Tuning · AAAI 2025 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.9 | 1 | 2025 | Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention Learning · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model fine-tuning
multilingual fine-tuning |
0.9 | 1 | 2025 | CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-Tuning · ACL (1) 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.9 | 1 | 2025 | Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced Optimization · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning |
0.9 | 1 | 2025 | CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-Tuning · ACL (1) 2025 |
Natural language and speech › Language models and text generation › text summarization › multilingual summarization
cross-lingual summarization |
0.8 | 1 | 2024 | GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News Summarization · EMNLP 2024 |
Natural language and speech › Language models and text generation › text summarization
multi-document summarization |
0.8 | 1 | 2024 | GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News Summarization · EMNLP 2024 |
Natural language and speech › Language models and text generation › text summarization
multilingual summarization |
0.8 | 1 | 2024 | GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News Summarization · EMNLP 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.3 | 1 | 2026 | Culture-Aware Machine Translation in Large Language Models: Benchmarking and Investigation · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
trajectory modulation · 1.0language separability · 1.0game self-play · 1.0curriculum learning · 1.0benchmarking · 1.0optimization · 0.9logit supervision · 0.9feature supervision · 0.9deep supervision · 0.9abstention learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction TuningabstractJoint multilingual instruction tuning is a widely adopted approach to improve the multilingual instruction-following ability and downstream performance of large language models (LLMs), but the resulting multilingual capability remains highly sensitive to the composition and selection of the training data. Existing selection methods, often based on features like text quality, diversity, or task relevance, typically overlook the intrinsic linguistic structure of multilingual data. In this paper, we propose LangGPS, a lightweight two-stage pre-selection framework guided by language separability—a signal that quantifies how well samples in different languages can be distinguished in the model’s representation space. LangGPS first filters training data based on separability scores and then refines the subset using existing selection methods. Extensive experiments across six benchmarks and 22 languages demonstrate that applying LangGPS on top of existing selection methods improves their effectiveness and generalizability in multilingual training, especially for understanding tasks and low-resource languages. Further analysis reveals that highly separable samples facilitate the formation of clearer language boundaries and support faster adaptation, while low-separability samples tend to function as bridges for cross-lingual alignment. Besides, we also find that language separability can serves as an effective signal for multilingual curriculum learning, where interleaving samples with diverse separability levels yields stable and generalizable gains. Together, we hope our work offers a new perspective on data utility in multilingual contexts and support the development of more linguistically informed LLMs. Yangfan Ye, Xiachong Feng, Lei Huang 0021, Weitao Ma, Qichen Hong, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
AAAI | 1 |
| 2026 | Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-PlayabstractXiachong Feng, Deyi Yin, Xiaocheng Feng, Yi Jiang, Libo Qin, Yangfan Ye, Lei Huang, Weitao Ma, Qiming Li, Yuxuan Gu, Bing Qin, Lingpeng Kong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiachong Feng, Deyi Yin, Libo Qin 0001, Yangfan Ye, Lei Huang 0021, Weitao Ma, Yuxuan Gu 0004, Bing Qin 0001, Lingpeng Kong |
ACL (1) | 6 |
| 2026 | Culture-Aware Machine Translation in Large Language Models: Benchmarking and InvestigationabstractZekun Yuan, Yangfan Ye, Xiaocheng Feng, Baohang Li, Qichen Hong, Yunfei Lu, Dandan Tu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zekun Yuan, Yangfan Ye, Baohang Li, Qichen Hong, Yunfei Lu, Dandan Tu, Bing Qin 0001 |
ACL (1) | 2 |
| 2025 | Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-TuningabstractLarge language models (LLMs) have demonstrated significant progress in multilingual language understanding and generation. However, due to the imbalance in training data, their capabilities in non-English languages are limited. Recent studies revealed the English-pivot multilingual mechanism of LLMs, where LLMs implicitly convert non-English queries into English ones at the bottom layers and adopt English for thinking at the middle layers. However, due to the absence of explicit supervision for cross-lingual alignment in the intermediate layers of LLMs, the internal representations during these stages may become inaccurate. In this work, we introduce a deep supervision fine-tuning method (DFT) that incorporates additional supervision in the internal layers of the model to guide its workflow. Specifically, we introduce two training objectives on different layers of LLMs: one at the bottom layers to constrain the conversion of the target language into English, and another at the middle layers to constrain reasoning in English. To effectively achieve the guiding purpose, we designed two types of supervision signals: logits and feature, which represent a stricter constraint and a relatively more relaxed guidance. Our method guides the model to not only consider the final generated result when processing non-English inputs but also ensure the accuracy of internal representations. We conducted extensive experiments on typical English-centric large models, LLaMA-2 and Gemma-2, and the results on multiple multilingual datasets show that our method significantly outperforms traditional fine-tuning methods. Wenshuai Huo, Yichong Huang, Chengpeng Fu, Baohang Li, Yangfan Ye, Zhirui Zhang, Dandan Tu, Duyu Tang, Yunfei Lu, Hui Wang 0030, Bing Qin 0001 |
AAAI | 6 |
| 2025 | Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention LearningabstractLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu, Yangfan Ye, Liang Zhao, Weihong Zhong, Baoxin Wang, Dayong Wu, Guoping Hu, Lingpeng Kong, Tong Xiao, Ting Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lei Huang 0021, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu 0004, Yangfan Ye, Weihong Zhong, Baoxin Wang, Dayong Wu, Lingpeng Kong, Tong Xiao 0001, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 7 |
| 2025 | Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced OptimizationabstractLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan, Xiachong Feng, Yangfan Ye, Weihong Zhong, Yuxuan Gu, Baoxin Wang, Dayong Wu, Guoping Hu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lei Huang 0021, Weitao Ma, Yuchun Fan, Xiachong Feng, Yangfan Ye, Weihong Zhong, Yuxuan Gu 0004, Baoxin Wang, Dayong Wu, Bing Qin 0001 |
ACL (1) | 6 |
| 2025 | CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-TuningabstractYangfan Ye, Xiaocheng Feng, Zekun Yuan, Xiachong Feng, Libo Qin, Lei Huang, Weitao Ma, Yichong Huang, Zhirui Zhang, Yunfei Lu, Xiaohui Yan, Duyu Tang, Dandan Tu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yangfan Ye, Zekun Yuan, Xiachong Feng, Libo Qin 0001, Lei Huang 0021, Weitao Ma, Yichong Huang, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
ACL (1) | 1 |
| 2025 | Unveiling Entity-Level Unlearning for Large Language Models: A Comprehensive AnalysisabstractLarge language model unlearning has garnered increasing attention due to its potential to address security and privacy concerns, leading to extensive research in the field. However, existing studies have predominantly focused on instance-level unlearning, specifically targeting the removal of predefined instances containing sensitive content. This focus has left a gap in the exploration of removing an entire entity, which is critical in real-world scenarios such as copyright protection. To close this gap, we propose a novel task named Entity-level unlearning, which aims to erase entity-related knowledge from the target model completely. To investigate this task, we systematically evaluate popular unlearning algorithms, revealing that current methods struggle to achieve effective entity-level unlearning. Then, we further explore the factors that influence the performance of unlearning algorithms, identifying that the knowledge coverage of the forget set and its size play pivotal roles. Notably, our analysis also uncovers that entities introduced through fine-tuning are more vulnerable than pre-trained entities during unlearning. We hope these findings can inspire future improvements in entity-level unlearning for LLMs. Weitao Ma, Weihong Zhong, Lei Huang 0021, Yangfan Ye, Xiachong Feng, Bing Qin 0001 |
COLING | 5 |
| 2024 | GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News SummarizationabstractYangfan Ye, Xiachong Feng, Xiaocheng Feng, Weitao Ma, Libo Qin, Dongliang Xu, Qing Yang, Hongtao Liu, Bing Qin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yangfan Ye, Xiachong Feng, Weitao Ma, Libo Qin 0001, Dongliang Xu, Qing Yang 0033, Hongtao Liu 0008, Bing Qin 0001 |
EMNLP | 1 |
| 2023 | Dialogue Context Modelling for Action Item Detection: Solution for ICASSP 2023 Mug Challenge Track 5abstractAction item detection aims at recognizing sentences containing information about actionable tasks, which can help people quickly grasp core tasks in the meeting without going through the redundant meeting contents. Therefore, in this paper, we thoroughly describe our carefully designed solution for the Action Item Detection Track of the General Meeting Understanding and Generation (MUG) challenge in the ICASSP 2023 Signal Processing Grand Challenge. Specifically, we systematically analyze the task instances provided by MUG and find that the key ingredient for successful action item detection is leveraging the dialogue context information into consideration. To this end, we design a simple and effective method for modelling context and utterance information concurrently. The experimental results show our method achieves remarkable improvements over baseline models, with an absolute increase of 0.62 of the F1score on the validation set. The stable generalizability of our method is further verified by our score on the final test set1. Xiachong Feng, Yangfan Ye, Bing Qin 0001, Ting Liu 0001 |
ICASSP | 3 |