VLDB 2026 Research / reviewers in the wild / expert
Aron Henriksson
dblp:19/10107
· DBLP profile ↗
23ranked-venue papers
8as first author
14since 2021 · last 2025
0000-0001-9731-1048ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Theory of computation · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Data-Constrained Synthesis of Training Data for De-IdentificationabstractMany sensitive domains — such as the clinical domain — lack widely available datasets due to privacy risks. The increasing generative capabilities of large language models (LLMs) have made synthetic datasets a viable path forward. In this study, we domain-adapt LLMs to the clinical domain and generate synthetic clinical texts that are machine-annotated with tags for personally identifiable information using capable encoder-based NER models. The synthetic corpora are then used to train synthetic NER models. The results show that training NER models using synthetic corpora incurs only a small drop in predictive performance. The limits of this process are investigated in a systematic ablation study — using both Swedish and Spanish data. Our analysis shows that smaller datasets can be sufficient for domain-adapting LLMs for data synthesis. Instead, the effectiveness of this process is almost entirely contingent on the performance of the machine-annotating NER models trained using the original data. Thomas Vakili, Aron Henriksson, Hercules Dalianis |
ACL (1) | 2 |
| 2025 | Detecting Suicidal Ideation on Social Media Using Large Language Models with Zero-Shot Prompting
Golnaz Nikmehr, Aritz Bilbao-Jayo, Aron Henriksson, Aitor Almeida |
ICT4AWE | 3 |
| 2025 | Mind the gap: from plausible to valid self-explanations in large language modelsabstractAbstract This paper investigates the reliability of explanations generated by large language models (LLMs) when prompted to explain their previous output. We evaluate two kinds of such self-explanations ( SE )—extractive and counterfactual—using state-of-the-art LLMs (1B to 70B parameters) on three different classification tasks (both objective and subjective). In line with Agarwal et al. (Faithfulness versus plausibility: On the (Un)reliability of explanations from large language models. 2024. https://doi.org/10.48550/arXiv.2402.04614 ), our findings indicate a gap between perceived and actual model reasoning: while SE largely correlate with human judgment (i.e. are plausible ), they do not fully and accurately follow the model’s decision process (i.e. are not faithful ). Additionally, we show that counterfactual SE are not even necessarily valid in the sense of actually changing the LLM’s prediction. Our results suggest that extractive SE provide the LLM’s “guess” at an explanation based on training data. Conversely, counterfactual SE can help understand the LLM’s reasoning: We show that the issue of validity can be resolved by sampling counterfactual candidates at high temperature—followed by a validity check—and introducing a formula to estimate the number of tries needed to generate valid explanations. This simple method produces plausible and valid explanations that offer a 16 times faster alternative to SHAP on average in our experiments. Korbinian Randl, John Pavlopoulos, Aron Henriksson, Tony Lindgren |
Mach. Learn. | 3 |
| 2024 | Supporting Teaching-to-the-Curriculum by Linking Diagnostic Tests to Curriculum Goals: Using Textbook Content as Context for Retrieval-Augmented Generation with Large Language Models
Xiu Li 0002, Aron Henriksson, Martin Duneld, Jalal Nouri, Yongchao Wu |
AIED (1) | 2 |
| 2024 | Evaluating the Reliability of Self-explanations in Large Language Models
Korbinian Randl, John Pavlopoulos, Aron Henriksson, Tony Lindgren |
DS (1) | 3 |
| 2024 | Selecting from Multiple Strategies Improves the Foreseeable Reasoning of Tool-Augmented Large Language Models
Yongchao Wu, Aron Henriksson |
ECML/PKDD (3) | 2 |
| 2023 | Towards Improving the Reliability and Transparency of ChatGPT for Educational Question Answering
Yongchao Wu, Aron Henriksson, Martin Duneld, Jalal Nouri |
EC-TEL | 2 |
| 2023 | Multimodal fine-tuning of clinical language models for predicting COVID-19 outcomesabstractClinical prediction models tend only to incorporate structured healthcare data, ignoring information recorded in other data modalities, including free-text clinical notes. Here, we demonstrate how multimodal models that effectively leverage both structured and unstructured data can be developed for predicting COVID-19 outcomes. The models are trained end-to-end using a technique we refer to as multimodal fine-tuning, whereby a pre-trained language model is updated based on both structured and unstructured data. The multimodal models are trained and evaluated using a multicenter cohort of COVID-19 patients encompassing all encounters at the emergency department of six hospitals. Experimental results show that multimodal models, leveraging the notion of multimodal fine-tuning and trained to predict (i) 30-day mortality, (ii) safe discharge and (iii) readmission, outperform unimodal models trained using only structured or unstructured healthcare data on all three outcomes. Sensitivity analyses are performed to better understand how well the multimodal models perform on different patient groups, while an ablation study is conducted to investigate the impact of different types of clinical notes on model performance. We argue that multimodal models that make effective use of routinely collected healthcare data to predict COVID-19 outcomes may facilitate patient management and contribute to the effective use of limited healthcare resources. Aron Henriksson, Yash Pawar, Pontus Hedberg, Pontus Nauclér |
Artif. Intell. Medicine | 1 |
| 2022 | Leveraging Clinical BERT in Multimodal Mortality Prediction Models for COVID-19abstractClinical prediction models are often based solely on the use of structured data in electronic health records, e.g. vital parameters and laboratory results, effectively ignoring potentially valuable information recorded in other modalities, such as free-text clinical notes. Here, we report on the development of a multimodal model that combines structured and unstructured data. In particular, we study how best to make use of a clinical language model in a multimodal setup for predicting 30-day all-cause mortality upon hospital admission in patients with COVID-19. We evaluate three strategies for incorporating a domain-specific clinical BERT model in multimodal prediction systems: (i) without fine-tuning, (ii) with unimodal fine-tuning, and (iii) with multimodal fine-tuning. The best-performing model leverages multimodal fine-tuning, in which the clinical BERT model is updated based also on the structured data. This multimodal mortality prediction model is shown to outperform unimodal models that are based on using either only structured data or only unstructured data. The experimental results indicate that clinical prediction models can be improved by including data in other modalities and that multimodal fine-tuning of a clinical language model is an effective strategy for incorporating information from clinical notes in multimodal prediction systems. Yash Pawar, Aron Henriksson, Pontus Hedberg, Pontus Nauclér |
CBMS | 2 |
| 2022 | Improving the Timeliness of Early Prediction Models for Sepsis through Utility OptimizationabstractEarly prediction of sepsis can facilitate early in-tervention and lead to improved clinical outcomes. However, for early prediction models to be clinically useful, and also to reduce alarm fatigue, detection of sepsis needs to be timely with respect to onset, being neither too late nor too early. In this paper, we propose a utility-based loss function for training early prediction models, where utility is defined by a function according to when the predictions are made in relation to onset as well as to specified early, optimal and late time points. Two versions of the utility- based loss function are evaluated and compared to a cross- entropy loss baseline. Experimental results, using real clinical data from electronic health records, show that incorporating the utility-based loss function leads to superior multimodal early prediction models, detecting sepsis both more accurately and more timely. We argue that improving the timeliness of early prediction models is important for increasing their utility and acceptance in a clinical setting. Anastasios Lamproudis, Aron Henriksson, John Karlsson Valik, Pontus Nauclér |
ICTAI | 2 |
| 2022 | Evaluating Pretraining Strategies for Clinical BERT ModelsabstractResearch suggests that using generic language models in specialized domains may be sub-optimal due to significant domain differences. As a result, various strategies for developing domain-specific language models have been proposed, including techniques for adapting an existing generic language model to the target domain, e.g. through various forms of vocabulary modifications and continued domain-adaptive pretraining with in-domain data. Here, an empirical investigation is carried out in which various strategies for adapting a generic language model to the clinical domain are compared to pretraining a pure clinical language model. Three clinical language models for Swedish, pretrained for up to ten epochs, are fine-tuned and evaluated on several downstream tasks in the clinical domain. A comparison of the language models’ downstream performance over the training epochs is conducted. The results show that the domain-specific language models outperform a general-domain language model; however, there is little difference in performance of the various clinical language models. However, compared to pretraining a pure clinical language model with only in-domain data, leveraging and adapting an existing general-domain language model requires fewer epochs of pretraining with in-domain data. Anastasios Lamproudis, Aron Henriksson, Hercules Dalianis |
LREC | 2 |
| 2022 | Downstream Task Performance of BERT Models Pre-Trained Using Automatically De-Identified Clinical DataabstractAutomatic de-identification is a cost-effective and straightforward way of removing large amounts of personally identifiable information from large and sensitive corpora. However, these systems also introduce errors into datasets due to their imperfect precision. These corruptions of the data may negatively impact the utility of the de-identified dataset. This paper de-identifies a very large clinical corpus in Swedish either by removing entire sentences containing sensitive data or by replacing sensitive words with realistic surrogates. These two datasets are used to perform domain adaptation of a general Swedish BERT model. The impact of the de-identification techniques is assessed by training and evaluating the models using six clinical downstream tasks. The results are then compared to a similar BERT model domain-adapted using an unaltered version of the clinical corpus. The results show that using an automatically de-identified corpus for domain adaptation does not negatively impact downstream performance. We argue that automatic de-identification is an efficient way of reducing the privacy risks of domain-adapted models and that the models created in this paper should be safe to distribute to other academic researchers. Thomas Vakili, Anastasios Lamproudis, Aron Henriksson, Hercules Dalianis |
LREC | 3 |
| 2022 | Holistic data-driven requirements elicitation in the big data eraabstractAbstract Digital transformation stimulates continuous generation of large amounts of digital data, both in organizations and in society at large. As a consequence, there have been growing efforts in the Requirements Engineering community to consider digital data as sources for requirements acquisition, in addition to human stakeholders. The volume, velocity and variety of the data make requirements discovery increasingly dynamic, but also unstructured and complex, which current elicitation methods are unable to consider and manage in a systematic and efficient manner. We propose a framework, in the form of a conceptual metamodel and a method, for continuous and automated acquisition, analysis and aggregation of heterogeneous digital sources that aims to support data-driven requirements elicitation and management. The usability of the framework is partially validated by an in-depth case study from the business sector of video game development. Aron Henriksson, Jelena Zdravkovic |
Softw. Syst. Model. | 1 |
| 2021 | Data-Driven Agile Requirements Elicitation through the Lenses of Situational Method EngineeringabstractUbiquitous digitalization has led to the continuous generation of large amounts of digital data, both in organizations and in society at large. In the requirements engineering community, there has been a growing interest in considering digital data as new sources for requirements elicitation, in addition to stake-holders. The volume, dynamics, and variety of data makes iterative requirements elicitation increasingly continuous, but also unstructured and complex, which current agile methods are unable to consider and manage in a systematic and efficient manner. There is also the need to support software evolution by enabling a synergy of stakeholder-driven requirements elicitation and management with data-driven approaches. In this study, we propose extension of agile requirements elicitation by applying situational method engineering. The research is grounded on two studies in the business domains of video games and online banking. Xavier Franch, Aron Henriksson, Jolita Ralyté, Jelena Zdravkovic |
RE | 2 |
| 2015 | Handling Temporality of Clinical Events for Drug Safety Surveillance
Jing Zhao 0017, Aron Henriksson, Maria Kvist, Lars Asker, Henrik Boström |
AMIA | 2 |
| 2015 | Modeling electronic health records in ensembles of semantic spaces for adverse drug event detectionabstractAdverse drug events (ADEs) are heavily under-reported in electronic health records (EHRs). Alerting systems that are able to detect potential ADEs on the basis of patient-specific EHR data would help to mitigate this problem. To that end, the use of machine learning has proven to be both efficient and effective; however, challenges remain in representing the heterogeneous EHR data, which moreover tends to be high-dimensional and exceedingly sparse, in a manner conducive to learning high-performing predictive models. Prior work has shown that distributional semantics - that is, natural language processing methods that, traditionally, model the meaning of words in semantic (vector) space on the basis of co-occurrence information - can be exploited to create effective representations of sequential EHR data of various kinds. When modeling data in semantic space, an important design decision concerns the size of the context window around an object of interest, which governs the scope of co-occurrence information that is taken into account and affects the composition of the resulting semantic space. Here, we report on experiments conducted on 27 clinical datasets, demonstrating that performance can be significantly improved by modeling EHR data in ensembles of semantic spaces, consisting of multiple semantic spaces built with different context window sizes. A follow-up investigation is conducted to study the impact on predictive performance as increasingly more semantic spaces are included in the ensemble, demonstrating that accuracy tends to improve with the number of semantic spaces, albeit not monotonically so. Finally, a number of different strategies for combining the semantic spaces are explored, demonstrating the advantage of early (feature) fusion over late (classifier) fusion. Semantic space ensembles allow multiple views of (sparse) data to be captured (densely) and thereby enable improved performance to be obtained on the task of detecting ADEs in EHRs. Aron Henriksson, Jing Zhao 0017, Henrik Boström, Hercules Dalianis |
BIBM | 1 |
| 2015 | Modeling heterogeneous clinical sequence data in semantic space for adverse drug event detectionabstractThe enormous amounts of data that are continuously recorded in electronic health record systems offer ample opportunities for data science applications to improve healthcare. There are, however, challenges involved in using such data for machine learning, such as high dimensionality and sparsity, as well as an inherent heterogeneity that does not allow the distinct types of clinical data to be treated in an identical manner. On the other hand, there are also similarities across data types that may be exploited, e.g., the possibility of representing some of them as sequences. Here, we apply the notions underlying distributional semantics, i.e., methods that model the meaning of words in semantic (vector) space on the basis of co-occurrence information, to four distinct types of clinical data: free-text notes, on the one hand, and clinical events, in the form of diagnosis codes, drug codes and measurements, on the other hand. Each semantic space contains continuous vector representations for every unique word and event, which can then be used to create representations of, e.g., care episodes that, in turn, can be exploited by the learning algorithm. This approach does not only reduce sparsity, but also takes into account, and explicitly models, similarities between various items, and it does so in an entirely data-driven fashion. Here, we report on a series of experiments using the random forest learning algorithm that demonstrate the effectiveness, in terms of accuracy and area under ROC curve, of the proposed representation form over the commonly used bag-of-items counterpart. The experiments are conducted on 27 real datasets that each involves the (binary) classification task of detecting a particular adverse drug event. It is also shown that combining structured and unstructured data leads to significant improvements over using only one of them. Aron Henriksson, Jing Zhao 0017, Henrik Boström, Hercules Dalianis |
DSAA | 1 |
| 2015 | Cascading adverse drug event detection in electronic health recordsabstractThe ability to detect adverse drug events (ADEs) in electronic health records (EHRs) is useful in many medical applications, such as alerting systems that indicate when an ADE-specific diagnosis code should be assigned. Automating the detection of ADEs can be attempted by applying machine learning to existing, labeled EHR data. How to do this in an effective manner is, however, an open question. The issues addressed in this study concern the granularity of the classification task: (1) If we wish to predict the occurrence of any ADE, is it advantageous to conflate the various ADE class labels prior to learning, or should they be merged post prediction? (2) If we wish to predict a family of ADEs or even a specific ADE, can the predictive performance be enhanced by dividing the classification task into a cascading scheme: predicting first, on a coarse level, whether there is an ADE or not, and, in the former case, followed by a more specific prediction on which family the ADE belongs to, and then finally a prediction on the specific ADE within that particular family? In this study, we conduct a series of experiments using a real, clinical dataset comprising healthcare episodes that have been assigned one of eight ADE-related diagnosis codes and a set of randomly extracted episodes that have not been assigned any ADE code. It is shown that, when distinguishing between ADEs and non-ADEs, merging the various ADE labels prior to learning leads to significantly higher predictive performance in terms of accuracy and area under ROC curve. A cascade of random forests is moreover constructed to determine either the family of ADEs or the specific class label; here, the performance is indeed enhanced compared to directly employing a one-step prediction. This study concludes that, if predictive performance is of primary importance, the cascading scheme should be the recommended approach over employing a one-step prediction for detecting ADEs in EHRs. Jing Zhao 0017, Aron Henriksson, Henrik Boström |
DSAA | 2 |
| 2015 | Identifying adverse drug event information in clinical notes with distributional semantic representations of contextabstract• A corpus of Swedish clinical notes was annotated for adverse drug event information. • Detecting adverse drug events in clinical notes can support pharmacovigilance. • Modeling context with distributional semantics yielded better predictive models. • Distributed word representations allowed more context information to be incorporated. • Inter-sentential relations between drugs and disorders/findings are hard to detect. For the purpose of post-marketing drug safety surveillance, which has traditionally relied on the voluntary reporting of individual cases of adverse drug events (ADEs), other sources of information are now being explored, including electronic health records (EHRs), which give us access to enormous amounts of longitudinal observations of the treatment of patients and their drug use. Adverse drug events, which can be encoded in EHRs with certain diagnosis codes, are, however, heavily underreported. It is therefore important to develop capabilities to process, by means of computational methods, the more unstructured EHR data in the form of clinical notes, where clinicians may describe and reason around suspected ADEs. In this study, we report on the creation of an annotated corpus of Swedish health records for the purpose of learning to identify information pertaining to ADEs present in clinical notes. To this end, three key tasks are tackled: recognizing relevant named entities (disorders, symptoms, drugs), labeling attributes of the recognized entities (negation, speculation, temporality), and relationships between them (indication, adverse drug event). For each of the three tasks, leveraging models of distributional semantics – i.e., unsupervised methods that exploit co-occurrence information to model, typically in vector space, the meaning of words – and, in particular, combinations of such models, is shown to improve the predictive performance. The ability to make use of such unsupervised methods is critical when faced with large amounts of sparse and high-dimensional data, especially in domains where annotated resources are scarce. Aron Henriksson, Maria Kvist, Hercules Dalianis, Martin Duneld |
J. Biomed. Informatics | 1 |
| 2014 | Generating features for named entity recognition by learning prototypes in semantic space: The case of de-identifying health recordsabstractCreating sufficiently large annotated resources for supervised machine learning, and doing so for every problem and every domain, is prohibitively expensive. Techniques that leverage large amounts of unlabeled data, which are often readily available, may decrease the amount of data that needs to be annotated to obtain a certain level of performance, as well as improve performance when large annotated resources are indeed available. Here, the development of one such method is presented, where semantic features are generated by exploiting the available annotations to learn prototypical (vector) representations of each named entity class in semantic space, constructed by employing a model of distributional semantics (random indexing) over a large, unannotated, in-domain corpus. Binary features that describe whether a given word belongs to a specific named entity class are provided to the learning algorithm; the feature values are determined by calculating the (cosine) distance in semantic space to each of the learned prototype vectors and ascertaining whether they are below or above a given threshold, set to optimize Fβ-score. The proposed method is evaluated empirically in a series of experiments, where the case is health-record deidentification, a task that involves identifying protected health information (PHI) in text. It is shown that a conditional random fields model with access to the generated semantic features, in addition to a set of orthographic and syntactic features, significantly outperforms, in terms of F1-score, a baseline model without access to the semantic features. Moreover, the quality of the features is further improved by employing a number of slightly different models of distributional semantics in an ensemble. Finally, the way in which the features are generated allows one to optimize them for various Fβ-scores, giving some degree of control to trade off precision and recall. Methods that are able to improve performance on named entity recognition tasks by exploiting large amounts of unlabeled data may substantially reduce costs involved in creating annotated resources for every domain and every problem. Aron Henriksson, Hercules Dalianis, Stewart Kowalski |
BIBM | 1 |
| 2014 | Detecting adverse drug events with multiple representations of clinical measurementsabstractAdverse drug events (ADEs) are grossly under-reported in electronic health records (EHRs). This could be mitigated by methods that are able to detect ADEs in EHRs, thereby allowing for missing ADE-specific diagnosis codes to be identified and added. A crucial aspect of constructing such systems is to find proper representations of the data in order to allow the predictive modeling to be as accurate as possible. One category of EHR data that can be used as indicators of ADEs are clinical measurements. However, using clinical measurements as features is not unproblematic due to the high rate of missing values and they can be repeated a variable number of times in each patient health record. In this study, five basic representations of clinical measurements are proposed and evaluated to handle these two problems. An empirical investigation using random forest on 27 datasets from a real EHR database with different ADE targets is presented, demonstrating that the predictive performance, in terms of accuracy and area under ROC curve, is higher when representing clinical measurements crudely as whether they were taken or how many times they were taken by a patient. Furthermore, a sixth alternative, combining all five basic representations, significantly outperforms using any of the basic representation except for one. A subsequent analysis of variable importance is also conducted with this fused feature set, showing that when clinical measurements have a high missing rate, the number of times they were taken by one patient is ranked as more informative than looking at their actual values. The observation from random forest is also confirmed empirically using other commonly employed classifiers. This study demonstrates that the way in which clinical measurements from EHRs are presented has a high impact for ADE detection, and that using multiple representations outperforms using a basic representation. Jing Zhao 0017, Aron Henriksson, Lars Asker, Henrik Boström |
BIBM | 2 |
| 2013 | Identifying Synonymy between SNOMED Clinical Terms of Varying Length Using Distributional Analysis of Electronic Health Records
Aron Henriksson, Mike Conway, Martin Duneld, Wendy W. Chapman |
AMIA | 1 |
| 2011 | Diagnosis Code Assignment Support Using Random Indexing of Patient Records - A Qualitative Feasibility Study
Aron Henriksson, Martin Hassel, Maria Kvist |
AIME | 1 |