Magnus Sahlgren

dblp:76/3617 · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0001-5100-0535ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 7 first-author · 10 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 ELOQUENT CLEF Shared Tasks for Evaluation of Generative Language Model Quality, 2025 Edition
Jussi Karlgren, Ekaterina Artemova, Ondrej Bojar, Vladislav Mikhailov, Magnus Sahlgren, Erik Velldal, Lilja Øvrelid
ECIR (5)5
2025 SWEb: A Large Web Dataset for the Scandinavian Languages
abstract
This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the collection and processing pipeline, and introduces a novel model-based text extractor that significantly reduces complexity in comparison with rule-based approaches. We also introduce a new cloze-style benchmark for evaluating language models in Swedish, and use this test to compare models trained on the SWEb data to models trained on FineWeb, with competitive results. All data, models and code are shared openly.
Tobias Norlund, Tim Isbister, Amaru Cuba Gyllensten, Paul Gabriel dos Santos, Danila Petrelli, Ariel Ekgren, Magnus Sahlgren
ICLR7
2024 GPT-SW3: An Autoregressive Language Model for the Scandinavian Languages
abstract
This paper details the process of developing the first native large generative language model for the North Germanic languages, GPT-SW3. We cover all parts of the development process, from data collection and processing, training configuration and instruction finetuning, to evaluation, applications, and considerations for release strategies. We discuss pros and cons of developing large language models for smaller languages and in relatively peripheral regions of the globe, and we hope that this paper can serve as a guide and reference for other researchers that undertake the development of large generative models for smaller languages.
Ariel Ekgren, Amaru Cuba Gyllensten, Felix Stollenwerk, Joey Öhman, Tim Isbister, Evangelia Gogoulou, Fredrik Carlsson, Judit Casademont, Magnus Sahlgren
LREC/COLING9
2024 ELOQUENT CLEF Shared Tasks for Evaluation of Generative Language Model Quality
Jussi Karlgren, Luise Dürlich, Evangelia Gogoulou, Liane Guillou, Joakim Nivre, Magnus Sahlgren, Aarne Talman
ECIR (5)6
2024 Branch-GAN: Improving Text Generation with (not so) Large Language Models
abstract
The current advancements in open domain text generation have been spearheaded by Transformer-based large language models. Leveraging efficient parallelization and vast training datasets, these models achieve unparalleled text generation capabilities. Even so, current models are known to suffer from deficiencies such as repetitive texts, looping issues, and lack of robustness. While adversarial training through generative adversarial networks (GAN) is a proposed solution, earlier research in this direction has predominantly focused on older architectures, or narrow tasks. As a result, this approach is not yet compatible with modern language models for open-ended text generation, leading to diminished interest within the broader research community. We propose a computationally efficient GAN approach for sequential data that utilizes the parallelization capabilities of Transformer models. Our method revolves around generating multiple branching sequences from each training sample, while also incorporating the typical next-step prediction loss on the original data. In this way, we achieve a dense reward and loss signal for both the generator and the discriminator, resulting in a stable training dynamic. We apply our training method to pre-trained language models, using data from their original training set but less than 0.01% of the available data. A comprehensive human evaluation shows that our method significantly improves the quality of texts generated by the model while avoiding the previously reported sparsity problems of GAN approaches. Even our smaller models outperform larger original baseline models with more than 16 times the number of parameters. Finally, we corroborate previous claims that perplexity on held-out data is not a sufficient metric for measuring the quality of generated texts.
Fredrik Carlsson, Johan Broberg, Erik Hillbom, Magnus Sahlgren, Joakim Nivre
ICLR4
2023 Superlim: A Swedish Language Understanding Evaluation Benchmark
abstract
Aleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger, Joey Öhman, Yvonne Adesam, Lars Borin, Dana Dannélls, Markus Forsberg, Tim Isbister, Anna Lindahl, Martin Malmsten, Faton Rekathati, Magnus Sahlgren, Elena Volodina, Love Börjeson, Simon Hengchen, Nina Tahmasebi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Aleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger, Joey Öhman, Yvonne Adesam, Lars Borin, Dana Dannélls, Markus Forsberg, Tim Isbister, Anna Lindahl, Martin Malmsten, Faton Rekathati, Magnus Sahlgren, Elena Volodina, Love Börjeson, Simon Hengchen, Nina Tahmasebi
EMNLP14
2022 Fine-Grained Controllable Text Generation Using Non-Residual Prompting
abstract
Fredrik Carlsson, Joey Öhman, Fangyu Liu, Severine Verlinden, Joakim Nivre, Magnus Sahlgren. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Fredrik Carlsson, Joey Öhman, Fangyu Liu 0001, Severine Verlinden, Joakim Nivre, Magnus Sahlgren
ACL (1)6
2022 Cross-lingual and Multilingual CLIP
abstract
The long-standing endeavor of relating the textual and the visual domain recently underwent a pivotal breakthrough, as OpenAI released CLIP. This model distinguishes how well an English text corresponds with a given image with unprecedented accuracy. Trained via a contrastive learning objective over a huge dataset of 400M of images and captions, it is a work that is not easily replicated, especially for low resource languages. Capitalizing on the modularization of the CLIP architecture, we propose to use cross-lingual teacher learning to re-train the textual encoder for various non-English languages. Our method requires no image data and relies entirely on machine translation which removes the need for data in the target language. We find that our method can efficiently train a new textual encoder with relatively low computational cost, whilst still outperforming previous baselines on multilingual image-text retrieval.
Fredrik Carlsson, Philipp Eisen, Faton Rekathati, Magnus Sahlgren
LREC4
2022 Lessons Learned from GPT-SW3: Building the First Large-Scale Generative Language Model for Swedish
abstract
We present GTP-SW3, a 3.5 billion parameter autoregressive language model, trained on a newly created 100 GB Swedish corpus. This paper provides insights with regards to data collection and training, while highlights the challenges of proper model evaluation. The results of quantitive evaluation through perplexity indicate that GPT-SW3 is a competent model in comparison with existing autoregressive models of similar size. Additionally, we perform an extensive prompting study which reveals the good text generation capabilities of GTP-SW3.
Ariel Ekgren, Amaru Cuba Gyllensten, Evangelia Gogoulou, Alice Heiman, Severine Verlinden, Joey Öhman, Fredrik Carlsson, Magnus Sahlgren
LREC8
2022 Cross-lingual Transfer of Monolingual Models
abstract
Recent studies in cross-lingual learning using multilingual models have cast doubt on the previous hypothesis that shared vocabulary and joint pre-training are the keys to cross-lingual generalization. We introduce a method for transferring monolingual models to other languages through continuous pre-training and study the effects of such transfer from four different languages to English. Our experimental results on GLUE show that the transferred models outperform an English model trained from scratch, independently of the source language. After probing the model representations, we find that model knowledge from the source language enhances the learning of syntactic and semantic knowledge in English.
Evangelia Gogoulou, Ariel Ekgren, Tim Isbister, Magnus Sahlgren
LREC4
2021 Predicting Treatment Outcome from Patient Texts: The Case of Internet-Based Cognitive Behavioural Therapy
abstract
Evangelia Gogoulou, Magnus Boman, Fehmi Ben Abdesslem, Nils Hentati Isacsson, Viktor Kaldo, Magnus Sahlgren. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Evangelia Gogoulou, Magnus Boman, Fehmi Ben Abdesslem, Nils Hentati Isacsson, Viktor Kaldo, Magnus Sahlgren
EACL6
2021 Semantic Re-tuning with Contrastive Tension
Fredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää, Magnus Sahlgren
ICLR5
2019 Enriching Word Embeddings with a Regressor Instead of Labeled Corpora
abstract
We propose a novel method for enriching word-embeddings without the need of a labeled corpus. Instead, we show that relying on a regressor – trained with a small lexicon to predict pseudo-labels – significantly improves performance over current techniques that rely on human-derived sentence-level labels for an entire corpora. Our approach enables enrichment for corpora that have no labels (such as Wikipedia). Exploring the utility of this general approach in both sentiment and non-sentiment-focused tasks, we show how enriching embeddings, for both Twitter and Wikipedia-based embeddings, provide notable improvements in performance for: binary sentiment classification, SemEval Tasks, embedding analogy task, and, document classification. Importantly, our approach is notably better and more generalizable than other state-of-the-art approaches for enriching both labeled and unlabeled corpora.
Mohamed Abdalla 0001, Magnus Sahlgren, Graeme Hirst
AAAI2
2018 Analysis of Open Answers to Survey Questions through Interactive Clustering and Theme Extraction
abstract
This paper describes design principles for and the implementation of Gavagai Explorer---a new application which builds on interactive text clustering to extract themes from topically coherent text sets such as open text answers to surveys or questionnaires. An automated system is quick, consistent, and has full coverage over the study material. A system allows an analyst to analyze more answers in a given time period; provides the same initial results regardless of who does the analysis, reducing the risks of inter-rater discrepancy; and does not risk miss responses due to fatige or boredom. These factors reduce the cost and increase the reliability of the service. The most important feature, however, is relieving the human analyst from the frustrating aspects of the coding task, freeing the effort to the central challenge of understanding themes. Gavagai Explorer is available on-line.
Fredrik Espinoza, Ola Hamfors, Jussi Karlgren, Fredrik Olsson, Per Persson, Lars Hamberg, Magnus Sahlgren
CHIIR7
2018 Distributional Term Set Expansion
Amaru Cuba Gyllensten, Magnus Sahlgren
LREC2
2017 Random indexing of multidimensional data
abstract
Random indexing (RI) is a lightweight dimension reduction method, which is used, for example, to approximate vector semantic relationships in online natural language processing systems. Here we generalise RI to multidimensional arrays and therefore enable approximation of higher-order statistical relationships in data. The generalised method is a sparse implementation of random projections, which is the theoretical basis also for ordinary RI and other randomisation approaches to dimensionality reduction and data representation. We present numerical experiments which demonstrate that a multidimensional generalisation of RI is feasible, including comparisons with ordinary RI and principal component analysis. The RI method is well suited for online processing of data streams because relationship weights can be updated incrementally in a fixed-size distributed representation, and inner products can be approximated on the fly at low computational cost. An open source implementation of generalised RI is provided.
Fredrik Sandin, Blerim Emruli, Magnus Sahlgren
Knowl. Inf. Syst.3
2017 Active Learning and Visual Analytics for Stance Classification with ALVA
abstract
The automatic detection and classification of stance (e.g., certainty or agreement) in text data using natural language processing and machine-learning methods creates an opportunity to gain insight into the speakers’ attitudes toward their own and other people’s utterances. However, identifying stance in text presents many challenges related to training data collection and classifier training. To facilitate the entire process of training a stance classifier, we propose a visual analytics approach, called ALVA, for text data annotation and visualization. ALVA’s interplay with the stance classifier follows an active learning strategy to select suitable candidate utterances for manual annotaion. Our approach supports annotation process management and provides the annotators with a clean user interface for labeling utterances with multiple stance categories. ALVA also contains a visualization method to help analysts of the annotation and training process gain a better understanding of the categories used by the annotators. The visualization uses a novel visual representation, called CatCombos, which groups individual annotation items by the combination of stance categories. Additionally, our system makes a visualization of a vector space model available that is itself based on utterances. ALVA is already being used by our domain experts in linguistics and computational linguistics to improve the understanding of stance phenomena and to build a stance classifier for applications such as social media monitoring.
Kostiantyn Kucher, Carita Paradis, Magnus Sahlgren, Andreas Kerren
ACM Trans. Interact. Intell. Syst.3
2016 The Effects of Data Size and Frequency Range on Distributional Semantic Models
abstract
This paper investigates the effects of data size and frequency range on distributional semantic models. We compare the performance of a number of representative models for several test settings over data of varying sizes, and over test items of various frequency. Our results show that neural network-based models underperform when the data is small, and that the most reliable model over data of varying sizes and frequency ranges is the inverted factorized model.
Magnus Sahlgren, Alessandro Lenci
EMNLP1
2016 The Gavagai Living Lexicon
Magnus Sahlgren, Amaru Cuba Gyllensten, Fredrik Espinoza, Ola Hamfors, Jussi Karlgren, Fredrik Olsson, Per Persson, Akshay Viswanathan, Anders Holst
LREC1
2015 Navigating the Semantic Horizon using Relative Neighborhood Graphs
abstract
This paper introduces a novel way to navigate neighborhoods in distributional semantic models.The approach is based on relative neighborhood graphs, which uncover the topological structure of local neighborhoods in semantic space.This has the potential to overcome both the problem with selecting a proper k in k-NN search, and the problem that a ranked list of neighbors may conflate several different senses.We provide both qualitative and quantitative results that support the viability of the proposed method.
Amaru Cuba Gyllensten, Magnus Sahlgren
EMNLP2
2015 Factorization of Latent Variables in Distributional Semantic Models
abstract
This paper discusses the use of factorization techniques in distributional semantic models.We focus on a method for redistributing the weight of latent variables, which has previously been shown to improve the performance of distributional semantic models.However, this result has not been replicated and remains poorly understood.We refine the method, and provide additional theoretical justification, as well as empirical results that demonstrate the viability of the proposed approach.
Arvid Österlund, David Ödling, Magnus Sahlgren
EMNLP3
2012 Usefulness of Sentiment Analysis
Jussi Karlgren, Magnus Sahlgren, Fredrik Olsson, Fredrik Espinoza, Ola Hamfors
ECIR2
2010 Between Bags and Trees - Constructional Patterns in Text Used for Attitude Identification
Jussi Karlgren, Gunnar Eriksson, Magnus Sahlgren, Oscar Täckström
ECIR3
2009 Terminology mining in social media
abstract
The highly variable and dynamic word usage in social media presents serious challenges for both research and those commercial applications that are geared towards blogs or other user-generated non-editorial texts. This paper discusses and exemplifies a terminology mining approach for dealing with the productive character of the textual environment in social media. We explore the challenges of practically acquiring new terminology, and of modeling similarity and relatedness of terms from observing realistic amounts of data. We also discuss semantic evolution and density, and investigate novel measures for characterizing the preconditions for terminology mining.
Magnus Sahlgren, Jussi Karlgren
CIKM1
2008 Filaments of Meaning in Word Space
Jussi Karlgren, Anders Holst, Magnus Sahlgren
ECIR3
2006 Towards pertinent evaluation methodologies for word-space models
Magnus Sahlgren
LREC1
2005 Unsupervised Evaluation of Parser Robustness
Johnny Bigert, Jonas Sjöbergh, Ola Knutsson, Magnus Sahlgren
CICLing4
2005 Counting Lumps in Word Space: Density as a Measure of Corpus Homogeneity
Magnus Sahlgren, Jussi Karlgren
SPIRE1
2005 Automatic bilingual lexicon acquisition using random indexing of parallel corpora
abstract
This paper presents a very simple and effective approach to using parallel corpora for automatic bilingual lexicon acquisition. The approach, which uses the Random Indexing vector space methodology, is based on finding correlations between terms based on their distributional characteristics. The approach requires a minimum of preprocessing and linguistic knowledge, and is efficient, fast and scalable. In this paper, we explain how our approach differs from traditional cooccurrence-based word alignment algorithms, and we demonstrate how to extract bilingual lexica using the Random Indexing approach applied to aligned parallel data. The acquired lexica are evaluated by comparing them to manually compiled gold standards, and we report overlap of around 60%. We also discuss methodological problems with evaluating lexical resources of this kind.
Magnus Sahlgren, Jussi Karlgren
Nat. Lang. Eng.1
2004 Using Bag-of-Concepts to Improve the Performance of Support Vector Machines in Text Categorization
Magnus Sahlgren, Rickard Cöster
COLING1
2004 Automatic Bilingual Lexicon Acquisition Using Random Indexing of Aligned Bilingual Data
Magnus Sahlgren
LREC1