VLDB 2026 Research / reviewers in the wild / expert
Xin Men
dblp:263/9110
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 50% Deep learning architectures and training · 50% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
positional encoding |
1.5 | 2 | 2024 | Base of RoPE Bounds Context Length · NeurIPS 2024 Exploring Context Window of Large Language Models via Decomposed Positional Vectors · NeurIPS 2024 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling
context window extension |
0.8 | 1 | 2024 | Exploring Context Window of Large Language Models via Decomposed Positional Vectors · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | Exploring Context Window of Large Language Models via Decomposed Positional Vectors · NeurIPS 2024 |
Natural language and speech › Language models and text generation › compositional generalization › length generalization
length extrapolation |
0.8 | 1 | 2024 | Exploring Context Window of Large Language Models via Decomposed Positional Vectors · NeurIPS 2024 |
Natural language and speech › Language models and text generation
long context |
0.8 | 1 | 2024 | Base of RoPE Bounds Context Length · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › positional encoding
rotary position embedding |
0.8 | 1 | 2024 | Base of RoPE Bounds Context Length · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Exploring Context Window of Large Language Models via Decomposed Positional Vectors · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
theoretical analysis · 0.8positional vector decomposition · 0.8attention window extension · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Exploring Context Window of Large Language Models via Decomposed Positional VectorsabstractTransformer-based large language models (LLMs) typically have a limited context window, resulting in significant performance degradation when processing text beyond the length of the context window. Extensive studies have been proposed to extend the context window and achieve length extrapolation of LLMs, but there is still a lack of in-depth interpretation of these approaches. In this study, we explore the positional information within and beyond the context window for deciphering the underlying mechanism of LLMs. By using a mean-based decomposition method, we disentangle positional vectors from hidden states of LLMs and analyze their formation and effect on attention. Furthermore, when texts exceed the context window, we analyze the change of positional vectors in two settings, i.e., direct extrapolation and context window extension. Based on our findings, we design two training-free context window extension methods, positional vector replacement and attention window extension. Experimental results show that our methods can effectively extend the context window length. Zican Dong, Junyi Li 0001, Xin Men, Wayne Xin Zhao, Bingning Wang, Zhen Tian 0001, Weipeng Chen, Ji-Rong Wen |
NeurIPS | 3 |
| 2024 | Base of RoPE Bounds Context LengthabstractPosition embedding is a core component of current Large Language Models (LLMs). Rotary position embedding (RoPE), a technique that encodes the position information with a rotation matrix, has been the de facto choice for position embedding in many LLMs, such as the Llama series. RoPE has been further utilized to extend long context capability, which is roughly based on adjusting the \textit{base} parameter of RoPE to mitigate out-of-distribution (OOD) problems in position embedding. However, in this paper, we find that LLMs may obtain a superficial long-context ability based on the OOD theory. We revisit the role of RoPE in LLMs and propose a novel property of long-term decay, we derive that the \textit{base of RoPE bounds context length}: there is an absolute lower bound for the base value to obtain certain context length capability. Our work reveals the relationship between context length and RoPE base both theoretically and empirically, which may shed light on future long context training. Xin Men, Bingning Wang, Xianpei Han, Weipeng Chen |
NeurIPS | 2 |