Niklas Lavesson

dblp:55/1994 · DBLP profile ↗
← Back
35ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0002-0535-1761ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 11 · 4 first-authorSecurity and privacy · 5 · 1 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorTheory of computation · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Naming the Pain in machine learning-enabled systems engineering
abstract
Machine learning (ML)-enabled systems are being increasingly adopted by companies aiming to enhance their products and operational processes. This paper aims to deliver a comprehensive overview of the current status quo of engineering ML-enabled systems and lay the foundation to steer practically relevant and problem-driven academic research. We conducted an international survey to collect insights from practitioners on the current practices and problems in engineering ML-enabled systems. We received 188 complete responses from 25 countries. We conducted quantitative statistical analyses on contemporary practices using bootstrapping with confidence intervals and qualitative analyses on the reported problems using open and axial coding procedures. Our survey results reinforce and extend existing empirical evidence on engineering ML-enabled systems, providing additional insights into typical ML-enabled systems project contexts, the perceived relevance and complexity of ML life cycle phases, and current practices related to problem understanding, model deployment, and model monitoring. Furthermore, the qualitative analysis provides a detailed map of the problems practitioners face within each ML life cycle phase and the problems causing overall project failure. The results contribute to a better understanding of the status quo and problems in practical environments. We advocate for the further adaptation and dissemination of software engineering practices to enhance the engineering of ML-enabled systems. • International survey gathering insights from 188 practitioners across 25 countries. • Overview of current practices and challenges in engineering ML-enabled systems. • Inferential quantitative analysis reporting the status quo with confidence intervals. • Qualitative analysis mapping ML life cycle challenges and causes of project failure.
Marcos Kalinowski, Daniel Méndez 0001, Görkem Giray, Antonio Pedro Santos Alves, Kelly Azevedo, Tatiana Escovedo, Hugo Villamizar, Hélio Lopes 0001, Maria Teresa Baldassarre, Stefan Wagner 0001, Stefan Biffl, Jürgen Musil, Michael Felderer, Niklas Lavesson, Tony Gorschek
Inf. Softw. Technol.14
2025 Identifying key AI challenges in make-to-order manufacturing organisations: A multiple case study
abstract
Artificial Intelligence can make manufacturing organisations more effective and efficient, but it is not clear which AI tasks hold the greatest potential. Make-to-order manufacturers must constantly adapt to customers’ unique and rapidly changing needs, and therefore have different challenges than make-to-stock manufacturers. Our ambition is to develop an AI-enabled software system to support manufacturing organisations in improving their processes. To this end, we first seek to understand the data and technology requirements for key AI-enabled tasks in a make-to-order setting and determine the level of performance and explainability needed to address them. We perform a multiple case study of five make-to-order packaging manufacturers, interviewing personnel from sales, production, and supply chain to identify and prioritise operational challenges suitable for AI approaches. Demand forecasting emerges as the most important task, followed by predictive maintenance, quality inspection, complex decision risk estimation, and production planning. Participants emphasise the importance of explainable techniques to ensure trust in the systems. The results highlight a need for a greater control of the production process and a better understanding of customer needs. Although most of the tasks could be solved with current techniques, some, such as intermittent demand forecasting and complex decision risk estimation, would require further development. The study clarifies the potential of AI-enabled systems in make-to-order manufacturing and outlines the steps required to realise it.
Jonatan Flyckt, Tony Gorschek, Daniel Méndez 0001, Niklas Lavesson
J. Syst. Softw.4
2024 Adversarial Machine Learning in Industry: A Systematic Literature Review
abstract
Adversarial Machine Learning (AML) discusses the act of attacking and defending Machine Learning (ML) Models, an essential building block of Artificial Intelligence (AI). ML is applied in many software-intensive products and services and introduces new opportunities and security challenges. AI and ML will gain even more attention from the industry in the future, but threats caused by already-discovered attacks specifically targeting ML models are either overseen, ignored, or mishandled. Current AML research investigates attack and defense scenarios for ML in different industrial settings with a varying degree of maturity with regard to academic rigor and practical relevance. However, to the best of our knowledge, a synthesis of the state of academic rigor and practical relevance is missing. This literature study reviews studies in the area of AML in the context of industry, measuring and analyzing each study’s rigor and relevance scores. Overall, all studies scored a high rigor score and a low relevance score, indicating that the studies are thoroughly designed and documented but miss the opportunity to include touch points relatable for practitioners.
Felix Viktor Jedrzejewski, Lukas Thode, Jannik Fischbach, Tony Gorschek, Daniel Méndez 0001, Niklas Lavesson
Comput. Secur.6
2023 Status Quo and Problems of Requirements Engineering for Machine Learning: Results from an International Survey
Antonio Pedro Santos Alves, Marcos Kalinowski, Görkem Giray, Daniel Méndez 0001, Niklas Lavesson, Kelly Azevedo, Hugo Villamizar, Tatiana Escovedo, Hélio Lopes 0001, Stefan Biffl, Jürgen Musil, Michael Felderer, Stefan Wagner 0001, Maria Teresa Baldassarre, Tony Gorschek
PROFES (1)5
2023 Explaining rifle shooting factors through multi-sensor body tracking
abstract
There is a lack of data-driven training instructions for sports shooters, as instruction has commonly been based on subjective assessments. Many studies have correlated body posture and balance to shooting performance in rifle shooting tasks, but have mostly focused on single aspects of postural control. This study has focused on finding relevant rifle shooting factors by examining the entire body over sequences of time. A data collection was performed with 13 human participants carrying out live rifle shooting scenarios while being recorded with multiple body tracking sensors. A pre-processing pipeline produced a novel skeleton sequence representation, which was used to train a transformer model. The predictions from this model could be explained on a per sample basis using the attention mechanism, and visualised in an interactive format for humans to interpret. It was possible to separate the different phases of a shooting scenario from body posture with a high classification accuracy (80%). Shooting performance could be detected to an extent by separating participants using their strong and weak shooting hand. The dataset and pre-processing pipeline, as well as the techniques for generating explainable predictions presented in this study have laid the groundwork for future research in the sports shooting domain.
Jonatan Flyckt, Filip Andersson, Florian Westphal, Andreas Månsson, Niklas Lavesson
Intell. Data Anal.5
2022 Detecting ditches using supervised learning on high-resolution digital elevation models
abstract
Drained wetlands can constitute a large source of greenhouse gas emissions, but the drainage networks in these wetlands are largely unmapped, and better maps are needed to aid in forest production and to better understand the climate consequences. We develop a method for detecting ditches in high resolution digital elevation models derived from LiDAR scans. Thresholding methods using digital terrain indices can be used to detect ditches. However, a single threshold generally does not capture the variability in the landscape, and generates many false positives and negatives. We hypothesise that, by combining the digital terrain indices using supervised learning, we can improve ditch detection at a landscape-scale. In addition to digital terrain indices, additional features are generated by transforming the data to include neighbouring cells for better ditch predictions. A Random Forests classifier is used to locate the ditches, and its probability output is processed to remove noise, and binarised to produce the final ditch prediction. The confidence interval for the Cohen’s Kappa index ranges [0.655 , 0.781] between the evaluation plots with a confidence level of 95%. The study demonstrates that combining information from a suite of digital terrain indices using machine learning provides an effective technique for automatic ditch detection at a landscape-scale, aiding in both practical forest management and in combatting climate change.
Jonatan Flyckt, Filip Andersson, Niklas Lavesson, Liselott Nilsson, Anneli M. Ågren
Expert Syst. Appl.3
2021 Energy modeling of Hoeffding tree ensembles
abstract
Energy consumption reduction has been an increasing trend in machine learning over the past few years due to its socio-ecological importance. In new challenging areas such as edge computing, energy consumption and predictive accuracy are key variables during algorithm design and implementation. State-of-the-art ensemble stream mining algorithms are able to create highly accurate predictions at a substantial energy cost. This paper introduces the nmin adaptation method to ensembles of Hoeffding tree algorithms, to further reduce their energy consumption without sacrificing accuracy. We also present extensive theoretical energy models of such algorithms, detailing their energy patterns and how nmin adaptation affects their energy consumption. We have evaluated the energy efficiency and accuracy of the nmin adaptation method on five different ensembles of Hoeffding trees under 11 publicly available datasets. The results show that we are able to reduce the energy consumption significantly, by 21% on average, affecting accuracy by less than one percent on average.
Eva García Martín, Albert Bifet, Niklas Lavesson
Intell. Data Anal.3
2020 Representative Image Selection for Data Efficient Word Spotting
Florian Westphal, Håkan Grahn, Niklas Lavesson
DAS3
2019 A Case for Guided Machine Learning
Florian Westphal, Niklas Lavesson, Håkan Grahn
CD-MAKE2
2019 Higher Order Mining for Monitoring District Heating Substations
abstract
We propose a higher order mining (HOM) approach for modelling, monitoring and analyzing district heating (DH) substations' operational behaviour and performance. HOM is concerned with mining over patterns rather than primary or raw data. The proposed approach uses a combination of different data analysis techniques such as sequential pattern mining, clustering analysis, consensus clustering and minimum spanning tree (MST). Initially, a substation's operational behaviour is modeled by extracting weekly patterns and performing clustering analysis. The substation's performance is monitored by assessing its modeled behaviour for every two consecutive weeks. In case some significant difference is observed, further analysis is performed by integrating the built models into a consensus clustering and applying an MST for identifying deviating behaviours. The results of the study show that our method is robust for detecting deviating and sub-optimal behaviours of DH substations. In addition, the proposed method can facilitate domain experts in the interpretation and understanding of the substations' behaviour and performance by providing different data analysis and visualization techniques.
Shahrooz Abghari, Veselka Boeva, Jens Brage, Christian Johansson, Håkan Grahn, Niklas Lavesson
DSAA6
2019 Learning Character Recognition with Graph-Based Privileged Information
abstract
This paper proposes a pre-training method for neural network-based character recognizers to reduce the required amount of training data, and thus the human labeling effort. The proposed method transfers knowledge about the similarities between graph representations of characters to the recognizer by training to predict the graph edit distance. We show that convolutional neural networks trained with this method outperform traditional supervised learning if only ten or less labeled images per class are available. Furthermore, we show that our approach performs up to 33% better than a graph edit distance based recognition approach, even if only one labeled image per class is available.
Florian Westphal, Niklas Lavesson, Håkan Grahn
ICDAR2
2018 Document Image Binarization Using Recurrent Neural Networks
abstract
In the context of document image analysis, image binarization is an important preprocessing step for other document analysis algorithms, but also relevant on its own by improving the readability of images of historical documents. While historical document image binarization is challenging due to common image degradations, such as bleedthrough, faded ink or stains, achieving good binarization performance in a timely manner is a worthwhile goal to facilitate efficient information extraction from historical documents. In this paper, we propose a recurrent neural network based algorithm using Grid Long Short-Term Memory cells for image binarization, as well as a pseudo F-Measure based weighted loss function. We evaluate the binarization and execution performance of our algorithm for different choices of footprint size, scale factor and loss function. Our experiments show a significant trade-off between binarization time and quality for different footprint sizes. However, we see no statistically significant difference when using different scale factors and only limited differences for different loss functions. Lastly, we compare the binarization performance of our approach with the best performing algorithm in the 2016 handwritten document image binarization contest and show that both algorithms perform equally well.
Florian Westphal, Niklas Lavesson, Håkan Grahn
DAS2
2018 Hoeffding Trees with Nmin Adaptation
abstract
Machine learning software accounts for a significant amount of energy consumed in data centers. These algorithms are usually optimized towards predictive performance, i.e. accuracy, and scalability. This is the case of data stream mining algorithms. Although these algorithms are adaptive to the incoming data, they have fixed parameters from the beginning of the execution. We have observed that having fixed parameters lead to unnecessary computations, thus making the algorithm energy inefficient. In this paper we present the nmin adaptation method for Hoeffding trees. This method adapts the value of the nmin parameter, which significantly affects the energy consumption of the algorithm. The method reduces unnecessary computations and memory accesses, thus reducing the energy, while the accuracy is only marginally affected. We experimentally compared VFDT (Very Fast Decision Tree, the first Hoeffding tree algorithm) and CVFDT (Concept-adapting VFDT) with the VFDT-nmin (VFDT with nmin adaptation). The results show that VFDT-nmin consumes up to 27% less energy than the standard VFDT, and up to 92% less energy than CVFDT, trading off a few percent of accuracy in a few datasets.
Eva García Martín, Niklas Lavesson, Håkan Grahn, Emiliano Casalicchio, Veselka Boeva
DSAA2
2018 Evolutionary Clustering Techniques for Expertise Mining Scenarios
abstract
The problem addressed in this article concerns the development of evolutionary clustering techniques that can be applied to adapt the existing clustering solution to a clustering of newly collected ...
Veselka Boeva, Milena Angelova, Niklas Lavesson, Oliver Rosander, Elena Tsiporkova
ICAART (2)3
2018 A Minimum Spanning Tree Clustering Approach for Outlier Detection in Event Sequences
abstract
Outlier detection has been studied in many domains. Outliers arise due to different reasons such as mechanical issues, fraudulent behavior, and human error. In this paper, we propose an unsupervised approach for outlier detection in a sequence dataset. The proposed approach combines sequential pattern mining, cluster analysis, and a minimum spanning tree algorithm in order to identify clusters of outliers. Initially, the sequential pattern mining is used to extract frequent sequential patterns. Next, the extracted patterns are clustered into groups of similar patterns. Finally, the minimum spanning tree algorithm is used to find groups of outliers. The proposed approach has been evaluated on two different real datasets, i.e., smart meter data and video session data. The obtained results have shown that our approach can be applied to narrow down the space of events to a set of potential outliers and facilitate domain experts in further analysis and identification of system level issues.
Shahrooz Abghari, Veselka Boeva, Niklas Lavesson, Håkan Grahn, Selim Ickin, Jörgen Gustafsson
ICMLA3
2018 Efficient document image binarization using heterogeneous computing and parameter tuning
abstract
In the context of historical document analysis, image binarization is a first important step, which separates foreground from background, despite common image degradations, such as faded ink, stains, or bleed-through. Fast binarization has great significance when analyzing vast archives of document images, since even small inefficiencies can quickly accumulate to years of wasted execution time. Therefore, efficient binarization is especially relevant to companies and government institutions, who want to analyze their large collections of document images. The main challenge with this is to speed up the execution performance without affecting the binarization performance. We modify a state-of-the-art binarization algorithm and achieve on average a 3.5 times faster execution performance by correctly mapping this algorithm to a heterogeneous platform, consisting of a CPU and a GPU. Our proposed parameter tuning algorithm additionally improves the execution time for parameter tuning by a factor of 1.7, compared to previous parameter tuning algorithms. We see that for the chosen algorithm, machine learning-based parameter tuning improves the execution performance more than heterogeneous computing, when comparing absolute execution times.
Florian Westphal, Håkan Grahn, Niklas Lavesson
Int. J. Document Anal. Recognit.3
2017 Identification of Energy Hotspots: A Case Study of the Very Fast Decision Tree
Eva García Martín, Niklas Lavesson, Håkan Grahn
GPC2
2016 Large-scale information retrieval in software engineering - an experience report from industrial application
Michael Unterkalmsteiner, Tony Gorschek, Robert Feldt, Niklas Lavesson
Empir. Softw. Eng.4
2015 Energy Efficiency in Data Stream Mining
abstract
Data mining algorithms are usually designed to optimize a trade-off between predictive accuracy and computational efficiency. This paper introduces energy consumption and energy efficiency as important factors to consider during data mining algorithm analysis and evaluation. We extended the CRISP (Cross Industry Standard Process for Data Mining) framework to include energy consumption analysis. Based on this framework, we conducted an experiment to illustrate how energy consumption and accuracy are affected when varying the parameters of the Very Fast Decision Tree (VFDT) algorithm. The results indicate that energy consumption can be reduced by up to 92.5% (557 J) while maintaining accuracy.
Eva García Martín, Niklas Lavesson, Håkan Grahn
ASONAM2
2015 Improved concept drift handling in surgery prediction and other applications
Ayne A. Beyene, Tewelle Welemariam, Marie Persson Netz, Niklas Lavesson
Knowl. Inf. Syst.4
2014 A method for evaluation of learning components
Niklas Lavesson, Veselka Boeva, Elena Tsiporkova, Paul Davidsson
Autom. Softw. Eng.1
2014 Detecting serial residential burglaries using clustering
Anton Borg, Martin Boldt, Niklas Lavesson, Ulf Melander, Veselka Boeva
Expert Syst. Appl.3
2013 Open data for anomaly detection in maritime surveillance
Samira Kazemi, Shahrooz Abghari, Niklas Lavesson, Henric Johnson, Peter Ryman
Expert Syst. Appl.3
2012 E-mail Classification Using Social Network Information
abstract
A majority of E-mail is suspected to be spam. Traditional spam detection fails to differentiate between user needs and evolving social relationships. Online Social Networks (OSNs) contain more and more social information, contributed by users. OSN information may be used to improve spam detection. This paper presents a method that can use several social networks for detecting spam and a set of metrics for representing OSN data. The paper investigates the impact of using social network data extracted from an E-mail corpus to improve spam detection. The social data model is compared to traditional spam data models by generating and evaluating classifiers from both model types. The results show that accurate spam detectors can be generated from the low-dimensional social data model alone, however, spam detectors generated from combinations of the traditional and social models were more accurate than the detectors generated from either model in isolation.
Anton Borg, Niklas Lavesson
ARES2
2012 Veto-based Malware Detection
abstract
Malicious software (malware) represents a threat to the security and privacy of computer users. Traditional signature-based and heuristic-based methods are unsuccessful in detecting some forms of malware. This paper presents a malware detection approach based on supervised learning. The main contributions of the paper are an ensemble learning algorithm, two pre-processing techniques, and an empirical evaluation of the proposed algorithm. Sequences of operational codes are extracted as features from malware and benign files. These sequences are used to produce three different data sets with different configurations. A set of learning algorithms is evaluated on the data sets and the predictions are combined by the ensemble algorithm. The predicted output is decided on the basis of veto voting. The experimental results show that the approach can accurately detect both novel and known malware instances with higher recall in comparison to majority voting.
Raja Khurram Shahzad, Niklas Lavesson
ARES2
2012 Similarity assessment for removal of noisy end user license agreements
Niklas Lavesson, Stefan Axelsson
Knowl. Inf. Syst.1
2011 Accurate Adware Detection Using Opcode Sequence Extraction
abstract
Adware represents a possible threat to the security and privacy of computer users. Traditional signature-based and heuristic-based methods have not been proven to be successful at detecting this type of software. This paper presents an adware detection approach based on the application of data mining on disassembled code. The main contributions of the paper is a large publicly available adware data set, an accurate adware detection algorithm, and an extensive empirical evaluation of several candidate machine learning techniques that can be used in conjunction with the algorithm. We have extracted sequences of opcodes from adware and benign software and we have then applied feature selection, using different configurations, to obtain 63 data sets. Six data mining algorithms have been evaluated on these data sets in order to find an efficient and accurate detector. Our experimental results show that the proposed approach can be used to accurately detect both novel and known adware instances even though the binary difference between adware and legitimate software is usually small.
Raja Khurram Shahzad, Niklas Lavesson, Henric Johnson
ARES2
2011 CudaRF: A CUDA-based implementation of Random Forests
abstract
Machine learning algorithms are frequently applied in data mining applications. Many of the tasks in this domain concern high-dimensional data. Consequently, these tasks are often complex and computationally expensive. This paper presents a GPU-based parallel implementation of the Random Forests algorithm. In contrast to previous work, the proposed algorithm is based on the compute unified device architecture (CUDA). An experimental comparison between the CUDA-based algorithm (CudaRF), and state-of-the-art Random Forests algorithms (Fas-tRF and LibRF) shows that CudaRF outperforms both FastRF and LibRF for the studied classification task.
Håkan Grahn, Niklas Lavesson, Mikael Hellborg Lapajne, Daniel Slat
AICCSA2
2011 Your best might not be good enough: Ranking in collaborative social search engines
abstract
—A relevant feature of online social networks like Facebook is the scope for users to share external information from the web with their friends by sharing an URL. The phenomenon of sharing has bridged the web graph with the social network graph and the shared knowledge in ego networks
Prantik Bhattacharyya, Jeff Rowe, Shyhtsun Felix Wu, Karen Zita Haigh, Niklas Lavesson, Henric Johnson
CollaborateCom5
2011 Learning to detect spyware using end user license agreements
Niklas Lavesson, Martin Boldt, Paul Davidsson, Andreas Jacobsson
Knowl. Inf. Syst.1
2010 Detection of Spyware by Mining Executable Files
abstract
Spyware represents a serious threat to confidentiality since it may result in loss of control over private data for computer users. This type of software might collect the data and send it to a third party without informed user consent. Traditionally two approaches have been presented for the purpose of spyware detection: Signature-based Detection and Heuristic-based Detection. These approaches perform well against known Spyware but have not been proven to be successful at detecting new spyware. This paper presents a Spyware detection approach by using Data Mining (DM)technologies. Our approach is inspired by DM-based malicious code detectors, which are known to work well for detecting viruses and similar software. However, this type of detector has not been investigated in terms of how well it is able to detect spyware. We extract binary features, called n-grams, from both spyware and legitimate software and apply five different supervised learning algorithms to train classifiers that are able to classify unknown binaries by analyzing extracted n-grams. The experimental results suggest that our method is successful even when the training data is scarce.
Raja Khurram Shahzad, Syed Imran Haider, Niklas Lavesson
ARES3
2009 Analysis of Speed Sign Classification Algorithms Using Shape Based Segmentation of Binary Images
Azam Sheikh Muhammad, Niklas Lavesson, Paul Davidsson, Mikael G. Nilsson
CAIP2
2009 AMORI: A Metric-Based One Rule Inducer
abstract
The requirements of real-world data mining problems vary extensively. It is plausible to assume that some of these requirements can be expressed as application-specific performance metrics. An algorithm that is designed to maximize performance given a certain learning metric may not produce the best possible result according to these application-specific metrics. We have implemented A Metric-based One Rule Inducer (AMORI), for which it is possible to select the learning metric. We have compared the performance of this algorithm by embedding three different learning metrics (classification accuracy, the F-measure, and the area under the ROC curve), on 19 UCI data sets. In addition, we have compared the results of AMORI with those obtained using an existing rule learning algorithm of similar complexity (One Rule) and a state-of-the-art rule learner (Ripper). The experiments show that a performance gain is achieved, for all included metrics, when using identical metrics for learning and evaluation. We also show that each AMORI/metric combination outperforms One Rule when using identical learning and evaluation metrics. The performance of AMORI is acceptable when compared with Ripper. Overall, the results suggest that metric-based learning is a viable approach.
Niklas Lavesson, Paul Davidsson
SDM1
2008 Generic Methods for Multi-criteria Evaluation
abstract
When evaluating data mining algorithms that are applied to solve real-world problems there are often several, conflicting criteria that need to be considered. We investigate the concept of generic multi-criteria (MC) classifier and algorithm evaluation and perform a comparison of existing methods. This comparison makes explicit some of the important characteristics of MC analysis and focuses on finding out which method is most suitable for further development. Generic MC methods can be described as frameworks for combining evaluation metrics and are generic in the sense that the metrics used are not dictated by the method; the choice of metric is instead dependent on the problem at hand. We discuss some scenarios that benefit from the application of generic MC methods and synthesize what we believe are attractive properties from the reviewed methods into a new method called the candidate evaluation function (CEF). Finally, we present a case study in which we apply CEF to trade-off several criteria when solving a real-world problem.
Niklas Lavesson, Paul Davidsson
SDM1
2006 Quantifying the Impact of Learning Algorithm Parameter Tuning
Niklas Lavesson, Paul Davidsson
AAAI1