VLDB 2026 Research / reviewers in the wild / expert
Fredrik Carlsson
dblp:36/1766
· DBLP profile ↗
7ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 48% Generative modeling · 24% Learning theory · 14% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › text generation
open-ended text generation |
1.6 | 2 | 2025 | The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation · ICLR 2025 Branch-GAN: Improving Text Generation with (not so) Large Language Models · ICLR 2024 |
Natural language and speech › Language models and text generation
decoding |
0.9 | 1 | 2025 | The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation · ICLR 2025 |
Machine learning › Learning theory
overfitting |
0.9 | 1 | 2025 | The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation · ICLR 2025 |
Machine learning › Generative modeling
generative adversarial network |
0.8 | 1 | 2024 | Branch-GAN: Improving Text Generation with (not so) Large Language Models · ICLR 2024 |
Machine learning › Generative modeling › generative adversarial network
text GAN |
0.8 | 1 | 2024 | Branch-GAN: Improving Text Generation with (not so) Large Language Models · ICLR 2024 |
Natural language and speech › Language models and text generation
controllable text generation |
0.6 | 1 | 2022 | Fine-Grained Controllable Text Generation Using Non-Residual Prompting · ACL (1) 2022 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.5 | 1 | 2021 | Semantic Re-tuning with Contrastive Tension · ICLR 2021 |
Methods — techniques the papers use, named apart from their topics
greedy decoding · 0.9fine-tuning · 0.9next-token prediction · 0.8adversarial training · 0.8non-residual prompting · 0.6contrastive learning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text GenerationabstractThis paper introduces the counter-intuitive generalization results of overfitting pre-trained large language models (LLMs) on very small datasets. In the setting of open-ended text generation, it is well-documented that LLMs tend to generate repetitive and dull sequences, a phenomenon that is especially apparent when generating using greedy decoding. This issue persists even with state-of-the-art LLMs containing billions of parameters, trained via next-token prediction on large datasets. We find that by further fine-tuning these models to achieve a near-zero training loss on a small set of samples -- a process we refer to as hyperfitting -- the long-sequence generative capabilities are greatly enhanced.
Greedy decoding with these Hyperfitted models even outperform Top-P sampling over long-sequences, both in terms of diversity and human preferences. This phenomenon extends to LLMs of various sizes, different domains, and even autoregressive image generation. We further find this phenomena to be distinctly different from that of Grokking and double descent. Surprisingly, our experiments indicate that hyperfitted models rarely fall into repeating sequences they were trained on, and even explicitly blocking these sequences results in high-quality output. All hyperfitted models produce extremely low-entropy predictions, often allocating nearly all probability to a single token. Fredrik Carlsson, Fangyu Liu 0001, Daniel Ward, Murathan Kurfali, Joakim Nivre |
ICLR | 1 |
| 2024 | GPT-SW3: An Autoregressive Language Model for the Scandinavian LanguagesabstractThis paper details the process of developing the first native large generative language model for the North Germanic languages, GPT-SW3. We cover all parts of the development process, from data collection and processing, training configuration and instruction finetuning, to evaluation, applications, and considerations for release strategies. We discuss pros and cons of developing large language models for smaller languages and in relatively peripheral regions of the globe, and we hope that this paper can serve as a guide and reference for other researchers that undertake the development of large generative models for smaller languages. Ariel Ekgren, Amaru Cuba Gyllensten, Felix Stollenwerk, Joey Öhman, Tim Isbister, Evangelia Gogoulou, Fredrik Carlsson, Judit Casademont, Magnus Sahlgren |
LREC/COLING | 7 |
| 2024 | Branch-GAN: Improving Text Generation with (not so) Large Language ModelsabstractThe current advancements in open domain text generation have been spearheaded by Transformer-based large language models. Leveraging efficient parallelization and vast training datasets, these models achieve unparalleled text generation capabilities. Even so, current models are known to suffer from deficiencies such as repetitive texts, looping issues, and lack of robustness. While adversarial training through generative adversarial networks (GAN) is a proposed solution, earlier research in this direction has predominantly focused on older architectures, or narrow tasks. As a result, this approach is not yet compatible with modern language models for open-ended text generation, leading to diminished interest within the broader research community. We propose a computationally efficient GAN approach for sequential data that utilizes the parallelization capabilities of Transformer models. Our method revolves around generating multiple branching sequences from each training sample, while also incorporating the typical next-step prediction loss on the original data. In this way, we achieve a dense reward and loss signal for both the generator and the discriminator, resulting in a stable training dynamic. We apply our training method to pre-trained language models, using data from their original training set but less than 0.01% of the available data. A comprehensive human evaluation shows that our method significantly improves the quality of texts generated by the model while avoiding the previously reported sparsity problems of GAN approaches. Even our smaller models outperform larger original baseline models with more than 16 times the number of parameters. Finally, we corroborate previous claims that perplexity on held-out data is not a sufficient metric for measuring the quality of generated texts. Fredrik Carlsson, Johan Broberg, Erik Hillbom, Magnus Sahlgren, Joakim Nivre |
ICLR | 1 |
| 2022 | Fine-Grained Controllable Text Generation Using Non-Residual PromptingabstractFredrik Carlsson, Joey Öhman, Fangyu Liu, Severine Verlinden, Joakim Nivre, Magnus Sahlgren. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Fredrik Carlsson, Joey Öhman, Fangyu Liu 0001, Severine Verlinden, Joakim Nivre, Magnus Sahlgren |
ACL (1) | 1 |
| 2022 | Cross-lingual and Multilingual CLIPabstractThe long-standing endeavor of relating the textual and the visual domain recently underwent a pivotal breakthrough, as OpenAI released CLIP. This model distinguishes how well an English text corresponds with a given image with unprecedented accuracy. Trained via a contrastive learning objective over a huge dataset of 400M of images and captions, it is a work that is not easily replicated, especially for low resource languages. Capitalizing on the modularization of the CLIP architecture, we propose to use cross-lingual teacher learning to re-train the textual encoder for various non-English languages. Our method requires no image data and relies entirely on machine translation which removes the need for data in the target language. We find that our method can efficiently train a new textual encoder with relatively low computational cost, whilst still outperforming previous baselines on multilingual image-text retrieval. Fredrik Carlsson, Philipp Eisen, Faton Rekathati, Magnus Sahlgren |
LREC | 1 |
| 2022 | Lessons Learned from GPT-SW3: Building the First Large-Scale Generative Language Model for SwedishabstractWe present GTP-SW3, a 3.5 billion parameter autoregressive language model, trained on a newly created 100 GB Swedish corpus. This paper provides insights with regards to data collection and training, while highlights the challenges of proper model evaluation. The results of quantitive evaluation through perplexity indicate that GPT-SW3 is a competent model in comparison with existing autoregressive models of similar size. Additionally, we perform an extensive prompting study which reveals the good text generation capabilities of GTP-SW3. Ariel Ekgren, Amaru Cuba Gyllensten, Evangelia Gogoulou, Alice Heiman, Severine Verlinden, Joey Öhman, Fredrik Carlsson, Magnus Sahlgren |
LREC | 7 |
| 2021 | Semantic Re-tuning with Contrastive Tension
Fredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää, Magnus Sahlgren |
ICLR | 1 |