VLDB 2026 Research / reviewers in the wild / expert
Yulun Du
dblp:200/8898
· DBLP profile ↗
7ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0002-6171-6750ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 53% Trustworthy machine learning · 26% Transfer learning and domain adaptation · 10% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
instruction tuning |
0.7 | 1 | 2023 | StoryWars: A Dataset and Instruction Tuning Baselines for Collaborative Story Understanding and Generation · ACL (1) 2023 |
Natural language and speech › Language models and text generation › instruction tuning
multi-task instruction tuning |
0.7 | 1 | 2023 | StoryWars: A Dataset and Instruction Tuning Baselines for Collaborative Story Understanding and Generation · ACL (1) 2023 |
Natural language and speech › Language models and text generation › prompting
automatic prompt generation |
0.6 | 1 | 2022 | GPS: Genetic Prompt Search for Efficient Few-Shot Learning · EMNLP 2022 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.6 | 1 | 2022 | GPS: Genetic Prompt Search for Efficient Few-Shot Learning · EMNLP 2022 |
Natural language and speech › Language models and text generation
prompting |
0.6 | 1 | 2022 | GPS: Genetic Prompt Search for Efficient Few-Shot Learning · EMNLP 2022 |
Natural language and speech › Language models and text generation › prompting
prompt search |
0.6 | 1 | 2022 | GPS: Genetic Prompt Search for Efficient Few-Shot Learning · EMNLP 2022 |
Machine learning › Trustworthy machine learning
interpretability |
0.5 | 1 | 2021 | Distribution Matching for Rationalization · AAAI 2021 |
Machine learning › Trustworthy machine learning › interpretability
rationalization |
0.5 | 1 | 2021 | Distribution Matching for Rationalization · AAAI 2021 |
Machine learning › Trustworthy machine learning
fairness |
0.3 | 1 | 2017 | Controllable Invariance through Adversarial Feature Learning · NIPS 2017 |
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning |
0.3 | 1 | 2017 | Controllable Invariance through Adversarial Feature Learning · NIPS 2017 |
Natural language and speech › Information extraction and text analysis
narrative understanding |
0.2 | 1 | 2023 | StoryWars: A Dataset and Instruction Tuning Baselines for Collaborative Story Understanding and Generation · ACL (1) 2023 |
Machine learning › Trustworthy machine learning › interpretability › rationalization
rationale extraction |
0.1 | 1 | 2021 | Distribution Matching for Rationalization · AAAI 2021 |
Natural language and speech › Information extraction and text analysis
text classification |
0.1 | 1 | 2021 | Distribution Matching for Rationalization · AAAI 2021 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.1 | 1 | 2017 | Controllable Invariance through Adversarial Feature Learning · NIPS 2017 |
Methods — techniques the papers use, named apart from their topics
instruction tuning · 0.7gradient-free optimization · 0.6genetic algorithm · 0.6mutual information maximization · 0.5distribution matching · 0.5minimax optimization · 0.3adversarial feature learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MoBA: Mixture of Block Attention for Long-Context LLMsabstractScaling the effective context length is essential for advancing large language models (LLMs) toward artificial general intelligence (AGI). However, the quadratic increase in computational complexity inherent in traditional attention mechanisms presents a prohibitive overhead. Existing approaches either impose strongly biased structures, such as sink or window attention which are task-specific, or radically modify the attention mechanism into linear approximations, whose performance in complex reasoning tasks remains inadequately explored.
In this work, we propose a solution that adheres to the ``less structure'' principle, allowing the model to determine where to attend autonomously, rather than introducing predefined biases. We introduce Mixture of Block Attention (MoBA), an innovative approach that applies the principles of Mixture of Experts (MoE) to the attention mechanism. This novel architecture demonstrates superior performance on long-context tasks while offering a key advantage: the ability to seamlessly transition between full and sparse attention, enhancing efficiency without the risk of compromising performance. MoBA has already been deployed to handle actual production workloads with long-context requirements, demonstrating significant advancements in efficient attention computation for LLMs. Our code is available at https://github.com/MoonshotAI/MoBA. Enzhe Lu, Zhejun Jiang, Yulun Du, Chao Hong, Weiran He, Enming Yuan, Yuzhi Wang, Huan Yuan, Suting Xu, Xinran Xu, Guokun Lai, Huabin Zheng, Jianlin Su, Yuxin Wu 0006, Jiezhong Qiu |
NeurIPS | 4 |
| 2023 | StoryWars: A Dataset and Instruction Tuning Baselines for Collaborative Story Understanding and GenerationabstractCollaborative stories, which are texts created through the collaborative efforts of multiple authors with different writing styles and intentions, pose unique challenges for NLP models.Understanding and generating such stories remains an underexplored area due to the lack of open-domain corpora.To address this, we introduce STORYWARS, a new dataset of over 40,000 collaborative stories written by 9,400 different authors from an online platform.We design 12 task types, comprising 7 understanding and 5 generation task types, on STORY-WARS, deriving 101 diverse story-related tasks in total as a multi-task benchmark covering all fully-supervised, few-shot, and zero-shot scenarios.Furthermore, we present our instructiontuned model, INSTRUCTSTORY, for the story tasks showing that instruction tuning, in addition to achieving superior results in zero-shot and few-shot scenarios, can also obtain the best performance on the fully-supervised tasks in STORYWARS, establishing strong multi-task benchmark performances on STORYWARS. 1 Yulun Du, Lydia B. Chilton |
ACL (1) | 1 |
| 2022 | GPS: Genetic Prompt Search for Efficient Few-Shot LearningabstractPrompt-based techniques have demostrated great potential for improving the few-shot generalization of pretrained language models.However, their performance heavily relies on the manual design of prompts and thus requires a lot of human efforts.In this paper, we introduce Genetic Prompt Search (GPS) to improve few-shot learning with prompts, which utilizes a genetic algorithm to automatically search for high-performing prompts.GPS is gradient-free and requires no update of model parameters but only a small validation set.Experiments on diverse datasets proved the effectiveness of GPS, which outperforms manual prompts by a large margin of 2.6 points.Our method is also better than other parameter-efficient tuning methods such as prompt tuning. Hanwei Xu, Yujun Chen, Yulun Du, Yanggang Wang |
EMNLP | 3 |
| 2021 | Distribution Matching for RationalizationabstractThe task of rationalization aims to extract pieces of input text as rationales to justify neural network predictions on text classification tasks. By definition, rationales represent key text pieces used for prediction and thus should have similar classification feature distribution compared to the original input text. However, previous methods mainly focused on maximizing the mutual information between rationales and labels while neglecting the relationship between rationales and input text. To address this issue, we propose a novel rationalization method that matches the distributions of rationales and input text in both the feature space and output space. Empirically, the proposed distribution matching approach consistently outperforms previous methods by a large margin. Our data and code are available. Yongfeng Huang 0001, Yujun Chen, Yulun Du |
AAAI | 3 |
| 2018 | Multimodal Polynomial Fusion for Detecting Driver DistractionabstractDistracted driving is deadly, claiming 3,477 lives in the U.S. in 2015 alone. Although there has been a considerable amount of research on modeling the distracted behavior of drivers under various conditions, accurate automatic detection using multiple modalities and especially the contribution of using the speech modality to improve accuracy has received little attention. This paper introduces a new multimodal dataset for distracted driving behavior and discusses automatic distraction detection using features from three modalities: facial expression, speech and car signals. Detailed multimodal feature analysis shows that adding more modalities monotonically increases the predictive accuracy of the model. Finally, a simple and effective multimodal fusion technique using a polynomial fusion layer shows superior distraction detection results compared to the baseline SVM and neural network models. Yulun Du, Alan W. Black, Louis-Philippe Morency, Maxine Eskénazi |
INTERSPEECH | 1 |
| 2017 | Controllable Invariance through Adversarial Feature LearningabstractLearning meaningful representations that maintain the content necessary for a particular task while filtering away detrimental variations is a problem of great interest in machine learning. In this paper, we tackle the problem of learning representations invariant to a specific factor or trait of data. The representation learning process is formulated as an adversarial minimax game. We analyze the optimal equilibrium of such a game and find that it amounts to maximizing the uncertainty of inferring the detrimental factor given the representation while maximizing the certainty of making task-specific predictions. On three benchmark tasks, namely fair and bias-free classification, language-independent generation, and lighting-independent image classification, we show that the proposed framework induces an invariant representation, and leads to better generalization evidenced by the improved performance. Qizhe Xie, Zihang Dai, Yulun Du, Eduard H. Hovy, Graham Neubig |
NIPS | 3 |
| 2017 | DialPort, Gone Live: An Update After A Year of DevelopmentabstractKyusong Lee, Tiancheng Zhao, Yulun Du, Edward Cai, Allen Lu, Eli Pincus, David Traum, Stefan Ultes, Lina M. Rojas-Barahona, Milica Gasic, Steve Young, Maxine Eskenazi. Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue. 2017. Kyusong Lee, Yulun Du, Edward Cai, Allen Lu, Eli Pincus, David R. Traum, Stefan Ultes, Lina Maria Rojas-Barahona, Milica Gasic, Steve J. Young, Maxine Eskénazi |
SIGDIAL Conference | 3 |