Lizhe Fang

dblp:391/7583 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 67% Deep learning architectures and training · 22% Generative modeling · 11%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
autoregressive model
0.912025
Rethinking Invariance in In-context Learning · ICLR 2025
Natural language and speech › Language models and text generation
in-context learning
0.912025
Rethinking Invariance in In-context Learning · ICLR 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
What is Wrong with Perplexity for Long-context Language Modeling? · ICLR 2025
Natural language and speech › Language models and text generation
large language model fine-tuning
0.912025
What is Wrong with Perplexity for Long-context Language Modeling? · ICLR 2025
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
long-context evaluation
0.912025
What is Wrong with Perplexity for Long-context Language Modeling? · ICLR 2025
Natural language and speech › Language models and text generation › language modeling
long-context language modeling
0.912025
What is Wrong with Perplexity for Long-context Language Modeling? · ICLR 2025
Machine learning › Deep learning architectures and training › symmetry-aware learning
permutation invariance
0.912025
Rethinking Invariance in In-context Learning · ICLR 2025
Natural language and speech › Language models and text generation › large language model evaluation
perplexity
0.912025
What is Wrong with Perplexity for Long-context Language Modeling? · ICLR 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Rethinking Invariance in In-context Learning · ICLR 2025

Methods — techniques the papers use, named apart from their topics

invariant in-context learning · 0.9contrastive learning · 0.9
YearPublicationVenuePosition
2025 Rethinking Invariance in In-context Learning
abstract
In-Context Learning (ICL) has emerged as a pivotal capability of auto-regressive large language models, yet it is hindered by a notable sensitivity to the ordering of context examples regardless of their mutual independence. To address this issue, recent studies have introduced several variant algorithms of ICL that achieve permutation invariance. However, many of these do not exhibit comparable performance with the standard auto-regressive ICL algorithm. In this work, we identify two crucial elements in the design of an invariant ICL algorithm: information non-leakage and context interdependence, which are not simultaneously achieved by any of the existing methods. These investigations lead us to the proposed \emph{Invariant ICL (InvICL)}, a methodology designed to achieve invariance in ICL while ensuring the two properties. Empirically, our findings reveal that InvICL surpasses previous models, both invariant and non-invariant, in most benchmark datasets, showcasing superior generalization capabilities across varying input lengths. Code is available at https://github.com/PKU-ML/InvICL.
Lizhe Fang, Yifei Wang 0001, Khashayar Gatmiry, Yisen Wang 0001
ICLR1
2025 What is Wrong with Perplexity for Long-context Language Modeling?
abstract
Handling long-context inputs is crucial for large language models (LLMs) in tasks such as extended conversations, document summarization, and many-shot in-context learning. While recent approaches have extended the context windows of LLMs and employed perplexity (PPL) as a standard evaluation metric, PPL has proven unreliable for assessing long-context capabilities. The underlying cause of this limitation has remained unclear. In this work, we provide a comprehensive explanation for this issue. We find that PPL overlooks key tokens, which are essential for long-context understanding, by averaging across all tokens and thereby obscuring the true performance of models in long-context scenarios. To address this, we propose \textbf{LongPPL}, a novel metric that focuses on key tokens by employing a long-short context contrastive method to identify them. Our experiments demonstrate that LongPPL strongly correlates with performance on various long-context benchmarks (e.g., Pearson correlation of -0.96), significantly outperforming traditional PPL in predictive accuracy. Additionally, we introduce \textbf{LongCE} (Long-context Cross-Entropy) loss, a re-weighting strategy for fine-tuning that prioritizes key tokens, leading to consistent improvements across diverse benchmarks. In summary, these contributions offer deeper insights into the limitations of PPL and present effective solutions for accurately evaluating and enhancing the long-context capabilities of LLMs. Code is available at https://github.com/PKU-ML/LongPPL.
Lizhe Fang, Yifei Wang 0001, Zhaoyang Liu 0003, Chenheng Zhang, Stefanie Jegelka, Jinyang Gao, Bolin Ding, Yisen Wang 0001
ICLR1