Hafsteinn Einarsson

dblp:155/6896 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 10 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
Hafsteinn Einarsson
LREC1
2026 Icelandic Math Eval: A Competitive Mathematics Benchmark for Large Language Models
Hafsteinn Einarsson, Jökull Ari Haraldsson, Ívar Armin Derayat, Sigrun Helga Lund, Benedikt Steinar Magnússon
LREC1
2026 Reformulate and Create, Don't Translate: Creating Natural Prompts for Underserved Languages
Annika Simonsen, Mathias Stenlund, Lars Bungum, Marc Daníel Skipstaþ Volhardt, Hafsteinn Einarsson
LREC5
2026 BRAGD: Constrained Multi-Label POS Tagging for Faroese
Annika Simonsen, Barbara Scalvini, Uni Johannesen, Iben Nyholm Debess, Hafsteinn Einarsson, Vésteinn Snæbjarnarson
LREC5
2026 Synthetic Instruction Generation for Low-Resource Nordic Languages: Viability and Limitations in LLM Instruction-Tuning
Mathias Stenlund, Annika Simonsen, Lars Bungum, Jan Ebert, Oleg Filatov, Hemanadhan Myneni, Morris Riedel, Hafsteinn Einarsson
LREC9
2024 Good or Bad News? Exploring GPT-4 for Sentiment Analysis for Faroese on a Public News Corpora
abstract
Sentiment analysis in low-resource languages presents unique challenges that Large Language Models may help address. This study explores the efficacy of GPT-4 for sentiment analysis on Faroese news texts, an uncharted task for this language. On the basis of guidelines presented, the sentiment analysis was performed with a multi-class approach at the sentence and document level with 225 sentences analysed in 170 articles. When comparing GPT-4 to human annotators, we observe that GPT-4 performs remarkably well. We explored two prompt configurations and observed a benefit from having clear instructions for the sentiment analysis task, but no benefit from translating the articles to English before the sentiment analysis task. Our results indicate that GPT-4 can be considered as a valuable tool for generating Faroese test data. Furthermore, our investigation reveals the intricacy of news sentiment. This motivates a more nuanced approach going forward, and we suggest a multi-label approach for future research in this domain. We further explored the efficacy of GPT-4 in topic classification on news texts and observed more negative sentiments expressed in international than national news. Overall, this work demonstrates GPT-4’s proficiency on a novel task and its utility for augmenting resources in low-data languages.
Iben Nyholm Debess, Annika Simonsen, Hafsteinn Einarsson
LREC/COLING3
2024 Gendered Grammar or Ingrained Bias? Exploring Gender Bias in Icelandic Language Models
abstract
Large language models, trained on vast datasets, exhibit increased output quality in proportion to the amount of data that is used to train them. This data-driven learning process has brought forth a pressing issue where these models may not only reflect but also amplify gender bias, racism, religious prejudice, and queerphobia present in their training data that may not always be recent. This study explores gender bias in language models trained on Icelandic, focusing on occupation-related terms. Icelandic is a highly grammatically gendered language that favors the masculine when referring to groups of people with indeterminable genders. Our aim is to explore whether language models merely mirror gender distributions within the corresponding professions or if they exhibit biases tied to their grammatical genders. Results indicate a significant overall predisposition towards the masculine but specific occupation terms consistently lean toward a particular gender, indicating complex interplays of societal and linguistic influences.
Steinunn Rut Friðriksdóttir, Hafsteinn Einarsson
LREC/COLING2
2024 A Human Perspective on GPT-4 Translations: Analysing Faroese to English News and Blog Text Translations
abstract
This study investigates the potential of Generative Pre-trained Transformer models, specifically GPT-4, to generate machine translation resources for the low-resource language, Faroese. Given the scarcity of high-quality, human-translated data for such languages, Large Language Models’ capabilities to produce native-sounding text offer a practical solution. This approach is particularly valuable for generating paired translation examples where one is in natural, authentic Faroese as opposed to traditional approaches that went from English to Faroese, addressing a common limitation in such approaches. By creating such a synthetic parallel dataset and evaluating it through the Multidimensional Quality Metrics framework, this research assesses the translation quality offered by GPT-4. The findings reveal GPT-4’s strengths in general translation tasks, while also highlighting its limitations in capturing cultural nuances.
Annika Simonsen, Hafsteinn Einarsson
EAMT (1)2
2022 Natural Questions in Icelandic
abstract
We present the first extractive question answering (QA) dataset for Icelandic, Natural Questions in Icelandic (NQiI). Developing such datasets is important for the development and evaluation of Icelandic QA systems. It also aids in the development of QA methods that need to work for a wide range of morphologically and grammatically different languages in a multilingual setting. The dataset was created by asking contributors to come up with questions they would like to know the answer to. Later, they were tasked with finding answers to each others questions following a previously published methodology. The questions are Natural in the sense that they are real questions posed out of interest in knowing the answer. The complete dataset contains 18 thousand labeled entries of which 5,568 are directly suitable for training an extractive QA system for Icelandic. The dataset is a valuable resource for Icelandic which we demonstrate by creating and evaluating a system capable of extractive QA in Icelandic.
Vésteinn Snæbjarnarson, Hafsteinn Einarsson
LREC2
2022 A Warm Start and a Clean Crawled Corpus - A Recipe for Good Language Models
abstract
We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error detection and constituency parsing. To train the models we introduce a new corpus of Icelandic text, the Icelandic Common Crawl Corpus (IC3), a collection of high quality texts found online by targeting the Icelandic top-level-domain .is. Several other public data sources are also collected for a total of 16GB of Icelandic text. To enhance the evaluation of model performance and to raise the bar in baselines for Icelandic, we manually translate and adapt the WinoGrande commonsense reasoning dataset. Through these efforts we demonstrate that a properly cleaned crawled corpus is sufficient to achieve state-of-the-art results in NLP applications for low to medium resource languages, by comparison with models trained on a curated corpus. We further show that initializing models using existing multilingual models can lead to state-of-the-art results for some downstream tasks.
Vésteinn Snæbjarnarson, Haukur Barri Símonarson, Pétur Orri Ragnarsson, Svanhvít Lilja Ingólfsdóttir, Haukur Páll Jónsson, Vilhjalmur Thorsteinsson, Hafsteinn Einarsson
LREC7
2019 The linear hidden subset problem for the (1 + 1) EA with scheduled and adaptive mutation rates
Hafsteinn Einarsson, Marcelo M. Gauy, Johannes Lengler, Florian Meier 0002, Asier Mujika, Angelika Steger, Felix Weissenberger
Theor. Comput. Sci.1
2018 The linear hidden subset problem for the (1 + 1) EA with scheduled and adaptive mutation rates
abstract
We study unbiased (1 + 1) evolutionary algorithms on linear functions with an unknown number n of bits with non-zero weight. Static algorithms achieve an optimal runtime of O(n(ln n)2+ε), however, it remained unclear whether more dynamic parameter policies could yield better runtime guarantees. We consider two setups: one where the mutation rate follows a fixed schedule, and one where it may be adapted depending on the history of the run. For the first setup, we give a schedule that achieves a runtime of (1±o(1))βn ln n, where β ≈ 3.552, which is an asymptotic improvement over the runtime of the static setup. Moreover, we show that no schedule admits a better runtime guarantee and that the optimal schedule is essentially unique. For the second setup, we show that the runtime can be further improved to (1 ± o(1))en ln n, which matches the performance of algorithms that know n in advance.
Hafsteinn Einarsson, Johannes Lengler, Marcelo M. Gauy, Florian Meier 0002, Asier Mujika, Angelika Steger, Felix Weissenberger
GECCO1
2017 Long Synfire Chains Emerge by Spike-Timing Dependent Plasticity Modulated by Population Activity
abstract
Sequences of precisely timed neuronal activity are observed in many brain areas in various species. Synfire chains are a well-established model that can explain such sequences. However, it is unknown under which conditions synfire chains can develop in initially unstructured networks by self-organization. This work shows that with spike-timing dependent plasticity (STDP), modulated by global population activity, long synfire chains emerge in sparse random networks. The learning rule fosters neurons to participate multiple times in the chain or in multiple chains. Such reuse of neurons has been experimentally observed and is necessary for high capacity. Sparse networks prevent the chains from being short and cyclic and show that the formation of specific synapses is not essential for chain formation. Analysis of the learning rule in a simple network of binary threshold neurons reveals the asymptotically optimal length of the emerging chains. The theoretical results generalize to simulated networks of conductance-based leaky integrate-and-fire (LIF) neurons. As an application of the emerged chain, we propose a one-shot memory for sequences of precisely timed neuronal activity.
Felix Weissenberger, Florian Meier 0002, Johannes Lengler, Hafsteinn Einarsson, Angelika Steger
Int. J. Neural Syst.4