VLDB 2026 Research / reviewers in the wild / expert
Magnus Sahlgren
dblp:76/3617
· DBLP profile ↗
31ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0001-5100-0535ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 7 first-author · 10 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ELOQUENT CLEF Shared Tasks for Evaluation of Generative Language Model Quality, 2025 Edition
Jussi Karlgren, Ekaterina Artemova, Ondrej Bojar, Vladislav Mikhailov, Magnus Sahlgren, Erik Velldal, Lilja Øvrelid |
ECIR (5) | 5 |
| 2025 | SWEb: A Large Web Dataset for the Scandinavian LanguagesabstractThis paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the collection and processing pipeline, and introduces a novel model-based text extractor that significantly reduces complexity in comparison with rule-based approaches. We also introduce a new cloze-style benchmark for evaluating language models in Swedish, and use this test to compare models trained on the SWEb data to models trained on FineWeb, with competitive results. All data, models and code are shared openly. Tobias Norlund, Tim Isbister, Amaru Cuba Gyllensten, Paul Gabriel dos Santos, Danila Petrelli, Ariel Ekgren, Magnus Sahlgren |
ICLR | 7 |
| 2024 | GPT-SW3: An Autoregressive Language Model for the Scandinavian LanguagesabstractThis paper details the process of developing the first native large generative language model for the North Germanic languages, GPT-SW3. We cover all parts of the development process, from data collection and processing, training configuration and instruction finetuning, to evaluation, applications, and considerations for release strategies. We discuss pros and cons of developing large language models for smaller languages and in relatively peripheral regions of the globe, and we hope that this paper can serve as a guide and reference for other researchers that undertake the development of large generative models for smaller languages. Ariel Ekgren, Amaru Cuba Gyllensten, Felix Stollenwerk, Joey Öhman, Tim Isbister, Evangelia Gogoulou, Fredrik Carlsson, Judit Casademont, Magnus Sahlgren |
LREC/COLING | 9 |
| 2024 | ELOQUENT CLEF Shared Tasks for Evaluation of Generative Language Model Quality
Jussi Karlgren, Luise Dürlich, Evangelia Gogoulou, Liane Guillou, Joakim Nivre, Magnus Sahlgren, Aarne Talman |
ECIR (5) | 6 |
| 2024 | Branch-GAN: Improving Text Generation with (not so) Large Language ModelsabstractThe current advancements in open domain text generation have been spearheaded by Transformer-based large language models. Leveraging efficient parallelization and vast training datasets, these models achieve unparalleled text generation capabilities. Even so, current models are known to suffer from deficiencies such as repetitive texts, looping issues, and lack of robustness. While adversarial training through generative adversarial networks (GAN) is a proposed solution, earlier research in this direction has predominantly focused on older architectures, or narrow tasks. As a result, this approach is not yet compatible with modern language models for open-ended text generation, leading to diminished interest within the broader research community. We propose a computationally efficient GAN approach for sequential data that utilizes the parallelization capabilities of Transformer models. Our method revolves around generating multiple branching sequences from each training sample, while also incorporating the typical next-step prediction loss on the original data. In this way, we achieve a dense reward and loss signal for both the generator and the discriminator, resulting in a stable training dynamic. We apply our training method to pre-trained language models, using data from their original training set but less than 0.01% of the available data. A comprehensive human evaluation shows that our method significantly improves the quality of texts generated by the model while avoiding the previously reported sparsity problems of GAN approaches. Even our smaller models outperform larger original baseline models with more than 16 times the number of parameters. Finally, we corroborate previous claims that perplexity on held-out data is not a sufficient metric for measuring the quality of generated texts. Fredrik Carlsson, Johan Broberg, Erik Hillbom, Magnus Sahlgren, Joakim Nivre |
ICLR | 4 |
| 2023 | Superlim: A Swedish Language Understanding Evaluation BenchmarkabstractAleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger, Joey Öhman, Yvonne Adesam, Lars Borin, Dana Dannélls, Markus Forsberg, Tim Isbister, Anna Lindahl, Martin Malmsten, Faton Rekathati, Magnus Sahlgren, Elena Volodina, Love Börjeson, Simon Hengchen, Nina Tahmasebi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Aleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger, Joey Öhman, Yvonne Adesam, Lars Borin, Dana Dannélls, Markus Forsberg, Tim Isbister, Anna Lindahl, Martin Malmsten, Faton Rekathati, Magnus Sahlgren, Elena Volodina, Love Börjeson, Simon Hengchen, Nina Tahmasebi |
EMNLP | 14 |
| 2022 | Fine-Grained Controllable Text Generation Using Non-Residual PromptingabstractFredrik Carlsson, Joey Öhman, Fangyu Liu, Severine Verlinden, Joakim Nivre, Magnus Sahlgren. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Fredrik Carlsson, Joey Öhman, Fangyu Liu 0001, Severine Verlinden, Joakim Nivre, Magnus Sahlgren |
ACL (1) | 6 |
| 2022 | Cross-lingual and Multilingual CLIPabstractThe long-standing endeavor of relating the textual and the visual domain recently underwent a pivotal breakthrough, as OpenAI released CLIP. This model distinguishes how well an English text corresponds with a given image with unprecedented accuracy. Trained via a contrastive learning objective over a huge dataset of 400M of images and captions, it is a work that is not easily replicated, especially for low resource languages. Capitalizing on the modularization of the CLIP architecture, we propose to use cross-lingual teacher learning to re-train the textual encoder for various non-English languages. Our method requires no image data and relies entirely on machine translation which removes the need for data in the target language. We find that our method can efficiently train a new textual encoder with relatively low computational cost, whilst still outperforming previous baselines on multilingual image-text retrieval. Fredrik Carlsson, Philipp Eisen, Faton Rekathati, Magnus Sahlgren |
LREC | 4 |
| 2022 | Lessons Learned from GPT-SW3: Building the First Large-Scale Generative Language Model for SwedishabstractWe present GTP-SW3, a 3.5 billion parameter autoregressive language model, trained on a newly created 100 GB Swedish corpus. This paper provides insights with regards to data collection and training, while highlights the challenges of proper model evaluation. The results of quantitive evaluation through perplexity indicate that GPT-SW3 is a competent model in comparison with existing autoregressive models of similar size. Additionally, we perform an extensive prompting study which reveals the good text generation capabilities of GTP-SW3. Ariel Ekgren, Amaru Cuba Gyllensten, Evangelia Gogoulou, Alice Heiman, Severine Verlinden, Joey Öhman, Fredrik Carlsson, Magnus Sahlgren |
LREC | 8 |
| 2022 | Cross-lingual Transfer of Monolingual ModelsabstractRecent studies in cross-lingual learning using multilingual models have cast doubt on the previous hypothesis that shared vocabulary and joint pre-training are the keys to cross-lingual generalization. We introduce a method for transferring monolingual models to other languages through continuous pre-training and study the effects of such transfer from four different languages to English. Our experimental results on GLUE show that the transferred models outperform an English model trained from scratch, independently of the source language. After probing the model representations, we find that model knowledge from the source language enhances the learning of syntactic and semantic knowledge in English. Evangelia Gogoulou, Ariel Ekgren, Tim Isbister, Magnus Sahlgren |
LREC | 4 |
| 2021 | Predicting Treatment Outcome from Patient Texts: The Case of Internet-Based Cognitive Behavioural TherapyabstractEvangelia Gogoulou, Magnus Boman, Fehmi Ben Abdesslem, Nils Hentati Isacsson, Viktor Kaldo, Magnus Sahlgren. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Evangelia Gogoulou, Magnus Boman, Fehmi Ben Abdesslem, Nils Hentati Isacsson, Viktor Kaldo, Magnus Sahlgren |
EACL | 6 |
| 2021 | Semantic Re-tuning with Contrastive Tension
Fredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää, Magnus Sahlgren |
ICLR | 5 |
| 2019 | Enriching Word Embeddings with a Regressor Instead of Labeled CorporaabstractWe propose a novel method for enriching word-embeddings without the need of a labeled corpus. Instead, we show that relying on a regressor – trained with a small lexicon to predict pseudo-labels – significantly improves performance over current techniques that rely on human-derived sentence-level labels for an entire corpora. Our approach enables enrichment for corpora that have no labels (such as Wikipedia). Exploring the utility of this general approach in both sentiment and non-sentiment-focused tasks, we show how enriching embeddings, for both Twitter and Wikipedia-based embeddings, provide notable improvements in performance for: binary sentiment classification, SemEval Tasks, embedding analogy task, and, document classification. Importantly, our approach is notably better and more generalizable than other state-of-the-art approaches for enriching both labeled and unlabeled corpora. Mohamed Abdalla 0001, Magnus Sahlgren, Graeme Hirst |
AAAI | 2 |
| 2018 | Analysis of Open Answers to Survey Questions through Interactive Clustering and Theme ExtractionabstractThis paper describes design principles for and the implementation of Gavagai Explorer---a new application which builds on interactive text clustering to extract themes from topically coherent text sets such as open text answers to surveys or questionnaires. An automated system is quick, consistent, and has full coverage over the study material. A system allows an analyst to analyze more answers in a given time period; provides the same initial results regardless of who does the analysis, reducing the risks of inter-rater discrepancy; and does not risk miss responses due to fatige or boredom. These factors reduce the cost and increase the reliability of the service. The most important feature, however, is relieving the human analyst from the frustrating aspects of the coding task, freeing the effort to the central challenge of understanding themes. Gavagai Explorer is available on-line. Fredrik Espinoza, Ola Hamfors, Jussi Karlgren, Fredrik Olsson, Per Persson, Lars Hamberg, Magnus Sahlgren |
CHIIR | 7 |
| 2018 | Distributional Term Set Expansion
Amaru Cuba Gyllensten, Magnus Sahlgren |
LREC | 2 |
| 2017 | Random indexing of multidimensional dataabstractRandom indexing (RI) is a lightweight dimension reduction method, which is used, for example, to approximate vector semantic relationships in online natural language processing systems. Here we generalise RI to multidimensional arrays and therefore enable approximation of higher-order statistical relationships in data. The generalised method is a sparse implementation of random projections, which is the theoretical basis also for ordinary RI and other randomisation approaches to dimensionality reduction and data representation. We present numerical experiments which demonstrate that a multidimensional generalisation of RI is feasible, including comparisons with ordinary RI and principal component analysis. The RI method is well suited for online processing of data streams because relationship weights can be updated incrementally in a fixed-size distributed representation, and inner products can be approximated on the fly at low computational cost. An open source implementation of generalised RI is provided. Fredrik Sandin, Blerim Emruli, Magnus Sahlgren |
Knowl. Inf. Syst. | 3 |
| 2017 | Active Learning and Visual Analytics for Stance Classification with ALVAabstractThe automatic detection and classification of stance (e.g., certainty or agreement) in text data using natural language processing and machine-learning methods creates an opportunity to gain insight into the speakers’ attitudes toward their own and other people’s utterances. However, identifying stance in text presents many challenges related to training data collection and classifier training. To facilitate the entire process of training a stance classifier, we propose a visual analytics approach, called ALVA, for text data annotation and visualization. ALVA’s interplay with the stance classifier follows an active learning strategy to select suitable candidate utterances for manual annotaion. Our approach supports annotation process management and provides the annotators with a clean user interface for labeling utterances with multiple stance categories. ALVA also contains a visualization method to help analysts of the annotation and training process gain a better understanding of the categories used by the annotators. The visualization uses a novel visual representation, called CatCombos, which groups individual annotation items by the combination of stance categories. Additionally, our system makes a visualization of a vector space model available that is itself based on utterances. ALVA is already being used by our domain experts in linguistics and computational linguistics to improve the understanding of stance phenomena and to build a stance classifier for applications such as social media monitoring. Kostiantyn Kucher, Carita Paradis, Magnus Sahlgren, Andreas Kerren |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2016 | The Effects of Data Size and Frequency Range on Distributional Semantic ModelsabstractThis paper investigates the effects of data size and frequency range on distributional semantic models. We compare the performance of a number of representative models for several test settings over data of varying sizes, and over test items of various frequency. Our results show that neural network-based models underperform when the data is small, and that the most reliable model over data of varying sizes and frequency ranges is the inverted factorized model. Magnus Sahlgren, Alessandro Lenci |
EMNLP | 1 |
| 2016 | The Gavagai Living Lexicon
Magnus Sahlgren, Amaru Cuba Gyllensten, Fredrik Espinoza, Ola Hamfors, Jussi Karlgren, Fredrik Olsson, Per Persson, Akshay Viswanathan, Anders Holst |
LREC | 1 |
| 2015 | Navigating the Semantic Horizon using Relative Neighborhood GraphsabstractThis paper introduces a novel way to navigate neighborhoods in distributional semantic models.The approach is based on relative neighborhood graphs, which uncover the topological structure of local neighborhoods in semantic space.This has the potential to overcome both the problem with selecting a proper k in k-NN search, and the problem that a ranked list of neighbors may conflate several different senses.We provide both qualitative and quantitative results that support the viability of the proposed method. Amaru Cuba Gyllensten, Magnus Sahlgren |
EMNLP | 2 |
| 2015 | Factorization of Latent Variables in Distributional Semantic ModelsabstractThis paper discusses the use of factorization techniques in distributional semantic models.We focus on a method for redistributing the weight of latent variables, which has previously been shown to improve the performance of distributional semantic models.However, this result has not been replicated and remains poorly understood.We refine the method, and provide additional theoretical justification, as well as empirical results that demonstrate the viability of the proposed approach. Arvid Österlund, David Ödling, Magnus Sahlgren |
EMNLP | 3 |
| 2012 | Usefulness of Sentiment Analysis
Jussi Karlgren, Magnus Sahlgren, Fredrik Olsson, Fredrik Espinoza, Ola Hamfors |
ECIR | 2 |
| 2010 | Between Bags and Trees - Constructional Patterns in Text Used for Attitude Identification
Jussi Karlgren, Gunnar Eriksson, Magnus Sahlgren, Oscar Täckström |
ECIR | 3 |
| 2009 | Terminology mining in social mediaabstractThe highly variable and dynamic word usage in social media presents serious challenges for both research and those commercial applications that are geared towards blogs or other user-generated non-editorial texts. This paper discusses and exemplifies a terminology mining approach for dealing with the productive character of the textual environment in social media. We explore the challenges of practically acquiring new terminology, and of modeling similarity and relatedness of terms from observing realistic amounts of data. We also discuss semantic evolution and density, and investigate novel measures for characterizing the preconditions for terminology mining. Magnus Sahlgren, Jussi Karlgren |
CIKM | 1 |
| 2008 | Filaments of Meaning in Word Space
Jussi Karlgren, Anders Holst, Magnus Sahlgren |
ECIR | 3 |
| 2006 | Towards pertinent evaluation methodologies for word-space models
Magnus Sahlgren |
LREC | 1 |
| 2005 | Unsupervised Evaluation of Parser Robustness
Johnny Bigert, Jonas Sjöbergh, Ola Knutsson, Magnus Sahlgren |
CICLing | 4 |
| 2005 | Counting Lumps in Word Space: Density as a Measure of Corpus Homogeneity
Magnus Sahlgren, Jussi Karlgren |
SPIRE | 1 |
| 2005 | Automatic bilingual lexicon acquisition using random indexing of parallel corporaabstractThis paper presents a very simple and effective approach to using parallel corpora for automatic bilingual lexicon acquisition. The approach, which uses the Random Indexing vector space methodology, is based on finding correlations between terms based on their distributional characteristics. The approach requires a minimum of preprocessing and linguistic knowledge, and is efficient, fast and scalable. In this paper, we explain how our approach differs from traditional cooccurrence-based word alignment algorithms, and we demonstrate how to extract bilingual lexica using the Random Indexing approach applied to aligned parallel data. The acquired lexica are evaluated by comparing them to manually compiled gold standards, and we report overlap of around 60%. We also discuss methodological problems with evaluating lexical resources of this kind. Magnus Sahlgren, Jussi Karlgren |
Nat. Lang. Eng. | 1 |
| 2004 | Using Bag-of-Concepts to Improve the Performance of Support Vector Machines in Text Categorization
Magnus Sahlgren, Rickard Cöster |
COLING | 1 |
| 2004 | Automatic Bilingual Lexicon Acquisition Using Random Indexing of Aligned Bilingual Data
Magnus Sahlgren |
LREC | 1 |