EDBT 2026 Demo / reviewers in the wild / expert
Kaiyue Wen
dblp:322/0395
· DBLP profile ↗
16ranked-venue papers
7as first author
16since 2021 · last 2026
0009-0002-9168-1851ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing underwater images: a dual-constraint latent diffusion approach with multi-view contrastive learning
Peng Liu 0036, Kaiyue Wen, Yong-Chao Li, Junyu Dong |
Expert Syst. Appl. | 2 |
| 2025 | Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert ModelsabstractZihan Qiu, Zeyu Huang, Bo Zheng, Kaiyue Wen, Zekun Wang, Rui Men, Ivan Titov, Dayiheng Liu, Jingren Zhou, Junyang Lin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zihan Qiu, Bo Zheng 0007, Kaiyue Wen, Rui Men, Ivan Titov 0001, Dayiheng Liu, Jingren Zhou 0001, Junyang Lin |
ACL (1) | 4 |
| 2025 | Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive ImagesabstractRecent studies have shown that Large Vision-Language Models (VLMs) tend to neglect image content and over-rely on language-model priors, resulting in errors in visually grounded tasks and hallucinations.We hypothesize that this issue arises because existing VLMs are not explicitly trained to generate texts that are accurately grounded in fine-grained image details.To enhance visual feedback during VLM training, we propose S-VCO (Symmetrical Visual Contrastive Optimization), a novel finetuning objective that steers the model toward capturing important visual details and aligning them with corresponding text tokens.To further facilitate this detailed alignment, we introduce MVC, a paired image-text dataset built by automatically filtering and augmenting visual counterfactual data to challenge the model with hard contrastive cases involving Minimal Visual Contrasts.Experiments show that our method consistently improves VLM performance across diverse benchmarks covering various abilities and domains, achieving up to a 22% reduction in hallucinations, and significant gains in vision-centric and general tasks.Notably, these improvements become increasingly pronounced in benchmarks with higher visual dependency.In short, S-VCO offers a significant enhancement of VLM's visuallydependent task performance while retaining or even improving the model's general abilities. Shengguang Wu, Fan-Yun Sun, Kaiyue Wen, Nick Haber |
ACL (1) | 3 |
| 2025 | Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape ViewabstractTraining language models currently requires pre-determining a fixed compute budget because the typical cosine learning rate schedule depends on the total number of steps. In contrast, the Warmup-Stable-Decay (WSD) schedule uses a constant learning rate to produce a main branch of iterates that can in principle continue indefinitely without a pre-specified compute budget. Then, given any compute budget, one can branch out from the main branch at a proper time with a rapidly decaying learning rate to produce a strong model. Empirically, WSD generates an intriguing, non-traditional loss curve: the loss remains elevated during the stable phase but sharply declines during the decay phase. Towards explaining this phenomenon, we conjecture that pretraining loss exhibits a river valley landscape, which resembles a deep valley with a river at its bottom. Under this assumption, we show that during the stable phase, the iterate undergoes large oscillations due to the high learning rate, yet it progresses swiftly along the river. During the decay phase, the rapidly dropping learning rate minimizes the iterate’s oscillations, moving it closer to the river and revealing true optimization progress. Therefore, the sustained high learning rate phase and fast decaying phase are responsible for progress in the river and the mountain directions, respectively, and are both critical. Our analysis predicts phenomenons consistent with empirical observations and shows that this landscape can naturally emerge from pretraining on a simple bi-gram dataset. Inspired by the theory, we introduce WSD-S, a variant of WSD that reuses previous checkpoints’ decay phases and keeps only one main branch, where we resume from a decayed checkpoint. WSD-S empirically outperforms WSD and Cyclic-Cosine in obtaining multiple pretrained language model checkpoints across various compute budgets in a single run for parameters scaling from 0.1B to 1.2B. Kaiyue Wen, Zhiyuan Li 0005, Jason S. Wang, David Hall 0006, Percy Liang, Tengyu Ma 0001 |
ICLR | 1 |
| 2025 | RNNs are not Transformers (Yet): The Key Bottleneck on In-Context RetrievalabstractThis paper investigates the gap in representation powers of Transformers and Recurrent Neural Networks (RNNs), which are more memory efficient than Transformers. We aim to understand whether RNNs can match the performance of Transformers, particularly when enhanced with Chain-of-Thought (CoT) prompting. Our theoretical analysis reveals that CoT improves RNNs but is insufficient to close the gap with Transformers. A key bottleneck lies in the inability of RNNs to perfectly retrieve information from the context, even with CoT:
for several tasks that explicitly or implicitly require this capability, such as associative recall and determining if a graph is a tree, we prove that RNNs are not expressive enough to solve the tasks while Transformers can solve them with ease.
Conversely, we prove that adopting techniques to enhance the in-context retrieval capability of RNNs, including Retrieval-Augmented Generation (RAG) and adding a single Transformer layer, can elevate RNNs to be capable of solving all polynomial-time solvable problems with CoT, hence closing the representation gap with Transformers. We validate our theory on synthetic and natural language experiments. Kaiyue Wen, Xingyu Dang, Kaifeng Lyu |
ICLR | 1 |
| 2025 | From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample EfficiencyabstractChain-of-thought (CoT) significantly enhances the reasoning performance of large language models (LLM). While current theoretical studies often attribute this improvement to increased expressiveness and computational capacity, we argue that expressiveness is not the primary limitation in the LLM regime, as current large models will fail on simple tasks. Using a parity-learning setup, we demonstrate that CoT can substantially improve sample efficiency even when the representation power is sufficient. Specifically, with CoT, a transformer can learn the function within polynomial samples, whereas without CoT, the required sample size is exponential. Additionally, we show that CoT simplifies the learning process by introducing sparse sequential dependencies among input tokens, and leads to a sparse and interpretable attention. We validate our theoretical analysis with both synthetic and real-world experiments, confirming that sparsity in attention layers is a key factor of the improvement induced by CoT. Kaiyue Wen, Huaqing Zhang 0005, Hongzhou Lin, Jingzhao Zhang |
ICLR | 1 |
| 2025 | Task Generalization with Autoregressive Compositional Structure: Can Learning from D Tasks Generalize to DT Tasks?abstractLarge language models (LLMs) exhibit remarkable task generalization, solving tasks they were never explicitly trained on with only a few demonstrations. This raises a fundamental question: When can learning from a small set of tasks generalize to a large task family? In this paper, we investigate task generalization through the lens of autoregressive compositional structure, where each task is a composition of T operations, and each operation is among a finite family of D subtasks. This yields a total class of size D^T. We first show that generalization to all D^T tasks is theoretically achievable by training on only Õ(D) tasks. Empirically, we demonstrate that Transformers achieve such exponential task generalization on sparse parity functions via In-context Learning (ICL) and chain-of-thought (CoT) reasoning. We further demonstrate this exponential generalization in arithmetic and language translation, extending beyond parity functions. Amirhesam Abedsoltan, Huaqing Zhang 0005, Kaiyue Wen, Hongzhou Lin, Jingzhao Zhang, Mikhail Belkin |
ICML | 3 |
| 2025 | Overtrained Language Models Are Harder to Fine-TuneabstractLarge language models are pre-trained on ever-growing token budgets under the assumption that better pre-training performance translates to improved downstream models. In this work, we challenge this assumption and show that extended pre-training can make models harder to fine-tune, leading to degraded final performance. We term this phenomenon \textbf{catastrophic overtraining}. For example, the instruction-tuned OLMo-1B model pre-trained on 3T tokens leads to over 2\% worse performance on multiple standard LLM benchmarks than its 2.3T token counterpart. Through controlled experiments and theoretical analysis, we show that catastrophic overtraining arises from a systematic increase in the broad sensitivity of pre-trained parameters to modifications, including but not limited to fine-tuning. Our findings call for a critical reassessment of pre-training design that considers the downstream adaptability of the model. Jacob Mitchell Springer, Sachin Goyal, Kaiyue Wen, Tanishq Kumar, Xiang Yue, Sadhika Malladi, Graham Neubig, Aditi Raghunathan |
ICML | 3 |
| 2025 | Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeabstractGating mechanisms have been widely utilized, from early models like LSTMs and Highway Networks to recent state space models, linear attention, and also softmax attention.
Yet, existing literature rarely examines the specific effects of gating.
In this work, we conduct comprehensive experiments to systematically investigate gating-augmented softmax attention variants.
Specifically, we perform a comprehensive comparison over 30 variants of 15B Mixture-of-Experts (MoE) models and 1.7B dense models trained on a 3.5 trillion token dataset.
Our central finding is that a simple modification—applying a head-specific sigmoid gate after the Scaled Dot-Product Attention (SDPA)—consistently improves performance.
This modification also enhances training stability, tolerates larger learning rates, and improves scaling properties.
By comparing various gating positions and computational variants, we attribute this effectiveness to two key factors: (1) introducing non-linearity upon the low-rank mapping in the softmax attention, and (2) applying query-dependent sparse gating scores to modulate the SDPA output.
Notably, we find this sparse gating mechanism mitigates `massive activation`, `attention sink` and enhances long-context extrapolation performance.
We also release related codes (https://github.com/qiuzh20/gated_attention}) and models (https://huggingface.co/QwQZh/gated_attention) to facilitate future research.
Furthermore, the most effective SDPA output gating is used in the Qwen3-Next models (https://huggingface.co/collections/Qwen/qwen3-next). Zihan Qiu, Bo Zheng 0007, Kaiyue Wen, Rui Men, Suozhi Huang, Dayiheng Liu, Jingren Zhou 0001, Junyang Lin |
NeurIPS | 5 |
| 2025 | PaTH Attention: Position Encoding via Accumulating Householder TransformationsabstractThe attention mechanism is a core primitive in modern large language models (LLMs) and AI more broadly. Since attention by itself is permutation-invariant, position encoding is essential for modeling structured domains such as language. Rotary position encoding (RoPE) has emerged as the de facto standard approach for position encoding and is part of many modern LLMs. However, in RoPE the key/query transformation between two elements in a sequence is only a function of their relative position and otherwise independent of the actual input. This limits the expressivity of RoPE-based transformers.
This paper describes PaTH, a flexible data-dependent position encoding scheme based on accumulated products of Householder(like) transformations, where each transformation is data-dependent, i.e., a function of the input. We derive an efficient parallel algorithm for training through exploiting a compact representation of products of Householder matrices, and implement a FlashAttention-style blockwise algorithm. Across both targeted synthetic benchmarks and moderate-scale real-world language modeling experiments, we find that PaTH improves upon RoPE and other recent baselines. Finally, we show that we can convert pretrained RoPE transformers into PaTH with continued pretraining. Yikang Shen, Kaiyue Wen, Shawn Tan, Liliang Ren, Rameswar Panda |
NeurIPS | 3 |
| 2023 | How Sharpness-Aware Minimization Minimizes Sharpness?
Kaiyue Wen, Tengyu Ma 0001, Zhiyuan Li 0005 |
ICLR | 1 |
| 2023 | Benign Overfitting in Classification: Provably Counter Label Noise with Larger Models
Kaiyue Wen, Jiaye Teng, Jingzhao Zhang |
ICLR | 1 |
| 2023 | Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better GeneralizationabstractDespite extensive studies, the underlying reason as to why overparameterized
neural networks can generalize remains elusive. Existing theory shows that common stochastic optimizers prefer flatter minimizers of the training loss, and thus
a natural potential explanation is that flatness implies generalization. This work
critically examines this explanation. Through theoretical and empirical investigation, we identify the following three scenarios for two-layer ReLU networks: (1)
flatness provably implies generalization; (2) there exist non-generalizing flattest
models and sharpness minimization algorithms fail to generalize poorly, and (3)
perhaps most strikingly, there exist non-generalizing flattest models, but sharpness
minimization algorithms still generalize. Our results suggest that the relationship
between sharpness and generalization subtly depends on the data distributions
and the model architectures and sharpness minimization algorithms do not only
minimize sharpness to achieve better generalization. This calls for the search for
other explanations for the generalization of over-parameterized neural networks Kaiyue Wen, Zhiyuan Li 0005, Tengyu Ma 0001 |
NeurIPS | 1 |
| 2023 | Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammarsabstractTransformer interpretability aims to understand the algorithm implemented by a learned Transformer by examining various aspects of the model, such as the weight matrices or the attention patterns.
In this work, through a combination of theoretical results and carefully controlled experiments on synthetic data, we take a critical view
of methods that exclusively focus on individual parts of the model, rather than consider the network as a whole.
We consider a simple synthetic setup of learning a (bounded) Dyck language. Theoretically, we show that the set of models that (exactly or approximately) solve this task satisfy a structural characterization derived from ideas in formal languages (the pumping lemma).
We use this characterization to show that the set of optima is qualitatively rich; in particular, the attention pattern of a single layer can be "nearly randomized", while preserving the functionality of the network.
We also show via extensive experiments that these constructions are not merely a theoretical artifact: even with severe constraints to the architecture of the model, vastly different solutions can be reached via standard training. Thus, interpretability claims based on inspecting individual heads or weight matrices in the Transformer can be misleading. Kaiyue Wen, Yuchen Li 0007, Bingbin Liu, Andrej Risteski |
NeurIPS | 1 |
| 2022 | Finding Skill Neurons in Pre-trained Transformer-based Language ModelsabstractTransformer-based pre-trained language models have demonstrated superior performance on various natural language processing tasks.However, it remains unclear how the skills required to handle these tasks distribute among model parameters.In this paper, we find that after prompt tuning for specific tasks, the activations of some neurons within pre-trained Transformers 1 are highly predictive of the task labels.We dub these neurons skill neurons and confirm they encode task-specific skills by finding that: (1) Skill neurons are crucial for handling tasks.Performances of pre-trained Transformers on a task significantly drop when corresponding skill neurons are perturbed.(2) Skill neurons are task-specific.Similar tasks tend to have similar distributions of skill neurons.Furthermore, we demonstrate the skill neurons are most likely generated in pre-training rather than finetuning by showing that the skill neurons found with prompt tuning are also crucial for other fine-tuning methods freezing neuron weights, such as the adapter-based tuning and BitFit.We also explore the applications of skill neurons, including accelerating Transformers with network pruning and building better transferability indicators.These findings may promote further research on understanding Transformers.The source code can be obtained from https: //github.com/THU-KEG/Skill-Neuron. Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou 0001, Zhiyuan Liu 0001, Juan-Zi Li |
EMNLP | 2 |
| 2022 | On Transferability of Prompt Tuning for Natural Language ProcessingabstractYusheng Su, Xiaozhi Wang, Yujia Qin, Chi-Min Chan, Yankai Lin, Huadong Wang, Kaiyue Wen, Zhiyuan Liu, Peng Li, Juanzi Li, Lei Hou, Maosong Sun, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yusheng Su, Xiaozhi Wang, Yujia Qin, Chi-Min Chan, Yankai Lin 0001, Kaiyue Wen, Zhiyuan Liu 0001, Peng Li 0030, Juan-Zi Li, Lei Hou 0001, Maosong Sun 0001, Jie Zhou 0016 |
NAACL-HLT | 7 |