VLDB 2026 Research / reviewers in the wild / expert
Teodor Fredriksson
dblp:271/1353
· DBLP profile ↗
5ranked-venue papers
5as first author
3since 2021 · last 2023
0000-0001-8176-5846ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Classification of Complex-Valued Radar Data using Semi-Supervised Learning: a Case StudyabstractIn recent years, the interest in applying machine learning (ML) and deep learning (DL) has been increasing due to their ability to learn to predict and find structure in data. The most common approach of ML and DL is supervised learning. Supervised learning requires the input data to be labeled. However, as reported by many industries, such as the embedded systems domain, fully labeled datasets are difficult to obtain since data labeling is manually intensive. This paper uses a semi-supervised learning approach on real-world Pulse-Doppler data obtained from our industry collaborator Saab to address this challenge. We took inspiration from the FixMatch algorithm. To investigate whether unlabeled data can help improve classification accuracy, we compare FixMatch to a supervised baseline. We use five different settings for the number of available labels per class label to investigate how many labeled instances and how much manual effort is required for optimal accuracy. Bayesian Linear Regression is used to analyze the results. The results show that FixMatch can reach a higher accuracy than the supervised baseline. Furthermore, FixMatch requires more computation time but will help reduce manual effort. In addition, FixMatch will not underfit or overfit. Thanks to this study, practitioners know the benefits of utilizing FixMatch and when it is safe to use to improve a supervised baseline in the industry. Teodor Fredriksson, Jan Bosch, Helena Olsson |
SEAA | 1 |
| 2021 | An Empirical Evaluation of Algorithms for Data LabelingabstractThe lack of labeled data is a major problem in both research and industrial settings since obtaining labels is often an expensive and time-consuming activity. In the past years, several machine learning algorithms were developed to assist and perform automated labeling in partially labeled datasets. While many of these algorithms are available in open-source packages, there is a lack of research that investigates how these algorithms compare to each other for different types of datasets and with different percentages of available labels. To address this problem, this paper empirically evaluates and compares seven algorithms for automated labeling in terms of their accuracy. We investigate how these algorithms perform in twelve different and well-known datasets with three different types of data, images, texts, and numerical values. We evaluate these algorithms under two different experimental conditions, with 10% and 50% labels of available labels in the dataset. Each algorithm, in each dataset for each experimental condition, is evaluated independently ten times with different random seeds. The results are analyzed and the algorithms are compared utilizing a Bayesian Bradley-Terry model. The results indicate that the active learning algorithms using the query strategies uncertainty sampling, QBC and random sampling are always the best algorithms. However, this comes with the expense of increased manual labeling effort. These results help machine learning practitioners in choosing optimal machine learning algorithms to label their data. Teodor Fredriksson, David Issa Mattos, Jan Bosch, Helena Olsson |
COMPSAC | 1 |
| 2021 | Assessing the Suitability of Semi-Supervised Learning Datasets using Item Response TheoryabstractIn practice, supervised learning algorithms require fully labeled datasets to achieve the high accuracy demanded by current modern applications. However, in industrial settings supervised learning algorithms can perform poorly because of few labeled instances. Semi-supervised learning (SSL) is an automatic labeling approach that utilizes complete labels to infer missing labels in partially complete datasets. The high number of available SSL algorithms and the lack of systematic comparison between them leaves practitioners without guidelines to select the appropriate one for their application. Moreover, each SSL algorithm is often validated and evaluated in a small number of common datasets. However, there is no research that examines what datasets are suitable for comparing different SSL algorihtms. The purpose of this paper is to empirically evaluate the suitability of the datasets commonly used to evaluate and compare different SSL algorithms. We performed a simulation study using twelve datasets of three different datatypes (numerical, text, image) on thirteen different SSL algorithms. The contributions of this paper are two-fold. First, we propose the use of Bayesian congeneric item response theory model to assess the suitability of commonly used datasets. Second, we compare the different SSL algorithms using these datasets. The results show that with except of three datasets, the others have very low discrimination factors and are easily solved by the current algorithms. Additionally, the SSL algorithms have overlapping 90% credible intervals, indicating uncertainty in the difference between the accuracy of these SSL models. The paper concludes suggesting that researchers and practitioners should better consider the choice of datasets used for comparing SSL algorithms. Teodor Fredriksson, David Issa Mattos, Jan Bosch, Helena Olsson |
SEAA | 1 |
| 2020 | Machine Learning Models for Automatic Labeling: A Systematic Literature ReviewabstractAutomatic labeling is a type of classification problem. Classification has been studied with the help of statistical methods for a long time. With the explosion of new better computer processing units (CPUs) and graphical processing units (GPUs) the interest in machine learning has grown exponentially and we can use both statistical learning algorithms as well as deep neural networks (DNNs) to solve the classification tasks. Classification is a supervised machine learning problem and there exists a large amount of methodology for performing such task. However, it is very rare in industrial applications that data is fully labeled which is why we need good methodology to obtain error-free labels. The purpose of this paper is to examine the current literature on how to perform labeling using ML, we will compare these models in terms of popularity and on what datatypes they are used on. We performed a systematic literature review of empirical studies for machine learning for labeling. We identified 43 primary studies relevant to our search. From this we were able to determine the most common machine learning models for labeling. Lack of unlabeled instances is a major problem for industry as supervised learning is the most widely used. Obtaining labels is costly in terms of labor and financial costs. Based on our findings in this review we present alternate ways for labeling data for use in supervised learning tasks. Teodor Fredriksson, Jan Bosch, Helena Olsson |
ICSOFT | 1 |
| 2020 | Data Labeling: An Empirical Investigation into Industrial Challenges and Mitigation Strategies
Teodor Fredriksson, David Issa Mattos, Jan Bosch, Helena Olsson |
PROFES | 1 |