VLDB 2026 Research / reviewers in the wild / expert
Anna Wingkvist
dblp:55/7247
· DBLP profile ↗
19ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0002-0835-823XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 6 since 2021Human-computer interaction and ubiquitous computing · 7Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph Representation Learning for Software Architecture RecoveryabstractSoftware architecture recovery aims to infer a system's modular organization from source code, bridging the gap between design intent and implementation structure.Traditional techniques rely on handcrafted heuristics and often fail to capture deeper architectural relationships.We investigate whether GNNs can recover these relationships by framing the task as unsupervised representation learning over a multi-relational software graph.Our approach learns node embeddings that reflect architectural boundaries, offering a promising alternative to existing recovery methods. Rakhshanda Jabeen, Morgan Ericsson, Jonas Nordqvist, Anna Wingkvist |
ESANN | 4 |
| 2022 | Contextual Operationalization of Metrics as Scores: Is My Metric Value Good?abstractSoftware quality models aggregate metrics to indicate quality. Most metrics reflect counts derived from events or attributes that cannot directly be associated with quality. Worse, what constitutes a desirable value for a metric may vary across contexts. We demonstrate an approach to transforming arbitrary metrics into absolute quality scores by leveraging metrics captured from similar contexts. In contrast to metrics, scores represent freestanding quality properties that are also comparable. We provide a web-based tool for obtaining contextualized scores for metrics as obtained from one’s software. Our results indicate that significant differences among various metrics and contexts exist. The suggested approach works with arbitrary contexts. Given sufficient contextual information, it allows for answering the question of whether a metric value is good/bad or common/extreme. Sebastian Hönel, Morgan Ericsson, Welf Löwe, Anna Wingkvist |
QRS | 4 |
| 2022 | To automatically map source code entities to architectural modules with Naive BayesabstractThe process of mapping a source code entity onto an architectural module is to a large degree a manual task. Automating this process could increase the use of static architecture conformance checking methods, such as reflexion modeling, in industry. Current techniques rely on user parameterization and a highly cohesive design. A machine learning approach would potentially require less parameters and better use of the available information to aid in automatic mapping. We investigate how a classifier can be trained to map from source code to architecture modules automatically. This classifier is trained with semantic and syntactic dependency information extracted from the source code and from architecture descriptions. The classifier is implemented using multinomial naive Bayes and evaluated. We perform experiments and compare the classifier with three state-of-the-art mapping functions in eight open-source Java systems with known ground-truth-mappings. We find that the classifier outperforms the state-of-the-art in all cases and that it provides a useful baseline for further research in the area of semi-automatic incremental clustering. We conclude that machine learning is a useful approach that performs better and with less need for parameterization compared to other approaches. Future work includes investigating problematic mappings and a more diverse set of subject systems. Tobias Olsson, Morgan Ericsson, Anna Wingkvist |
J. Syst. Softw. | 3 |
| 2022 | Assessing the linguistic quality of REST APIs for IoT applicationsabstractInternet of Things (IoT) is a growing technology that relies on connected ‘things’ that gather data from peer devices and send data to servers via APIs (Application Programming Interfaces). The design quality of those APIs has a direct impact on their understandability and reusability. This study focuses on the linguistic design quality of REST APIs for IoT applications and assesses their linguistic quality by performing the detection of linguistic patterns and antipatterns in REST APIs for IoT applications. Linguistic antipatterns are considered poor practices in the naming, documentation, and choice of identifiers. In contrast, linguistic patterns represent best practices to APIs design. The linguistic patterns and their corresponding antipatterns are hence contrasting pairs. We propose the SARAv2 (Semantic Analysis of REST APIs version two) approach to perform syntactic and semantic analyses of REST APIs for IoT applications. Based on the SARAv2 approach, we develop the REST-Ling tool and empirically validate the detection results of nine linguistic antipatterns. We analyse 19 REST APIs for IoT applications. Our detection results show that the linguistic antipatterns are prevalent and the REST-Ling tool can detect linguistic patterns and antipatterns in REST APIs for IoT applications with an average accuracy of over 80%. Moreover, the tool performs the detection of linguistic antipatterns on average in the order of seconds, i.e., 8.396 s. We found that APIs generally follow good linguistic practices, although the prevalence of poor practices exists. Francis Palma, Tobias Olsson, Anna Wingkvist, Javier Gonzalez-Huerta |
J. Syst. Softw. | 3 |
| 2021 | Optimized Dependency Weights in Source Code Clustering
Tobias Olsson, Morgan Ericsson, Anna Wingkvist |
ECSA | 3 |
| 2021 | Weighted software metrics aggregation and its application to defect predictionabstractAbstract It is a well-known practice in software engineering to aggregate software metrics to assess software artifacts for various purposes, such as their maintainability or their proneness to contain bugs. For different purposes, different metrics might be relevant. However, weighting these software metrics according to their contribution to the respective purpose is a challenging task. Manual approaches based on experts do not scale with the number of metrics. Also, experts get confused if the metrics are not independent, which is rarely the case. Automated approaches based on supervised learning require reliable and generalizable training data, a ground truth, which is rarely available. We propose an automated approach to weighted metrics aggregation that is based on unsupervised learning. It sets metrics scores and their weights based on probability theory and aggregates them. To evaluate the effectiveness, we conducted two empirical studies on defect prediction, one on ca. 200 000 code changes, and another ca. 5 000 software classes. The results show that our approach can be used as an agnostic unsupervised predictor in the absence of a ground truth. Maria Ulan, Welf Löwe, Morgan Ericsson, Anna Wingkvist |
Empir. Softw. Eng. | 4 |
| 2021 | Copula-based software metrics aggregationabstractAbstract A quality model is a conceptual decomposition of an abstract notion of quality into relevant, possibly conflicting characteristics and further into measurable metrics. For quality assessment and decision making, metrics values are aggregated to characteristics and ultimately to quality scores. Aggregation has often been problematic as quality models do not provide the semantics of aggregation. This makes it hard to formally reason about metrics, characteristics, and quality. We argue that aggregation needs to be interpretable and mathematically well defined in order to assess, to compare, and to improve quality. To address this challenge, we propose a probabilistic approach to aggregation and define quality scores based on joint distributions of absolute metrics values. To evaluate the proposed approach and its implementation under realistic conditions, we conduct empirical studies on bug prediction of ca. 5000 software classes, maintainability of ca. 15000 open-source software systems, and on the information quality of ca. 100000 real-world technical documents. We found that our approach is feasible, accurate, and scalable in performance. Maria Ulan, Welf Löwe, Morgan Ericsson, Anna Wingkvist |
Softw. Qual. J. | 4 |
| 2020 | Current State and Next Steps on Automated Hints for Students Learning to CodeabstractThe core of this work-in-progress is that the best way to learn how to code is to practice by solving problems. However, if students have trouble with this, they can get frustrated and give up. Automated Tutoring Systems (ATS) aim to provide hints to help them solve the problems they encounter. Many of the existing systems offer general hints, e.g., "check the conditional statement" or help the student interpret the compiler or test-case errors. While this can be useful, we think that an ATS should provide interactive and specialized feedback for each program. We snowballed through publications on promising ATS and found that there are several such systems (in 27 publications), but we could also identify many challenges and that our requirements were not met by any existing system. For example, few of them work on general-purpose programming languages, e.g., Java, or scale to realistic problems consisting of multiple methods and classes. From the search, we find ATS based on Automated Program Repair (APR) shows the most promise. However, while program repair has the potential to generate specialized hints to help guide the student to a working state, studies that looked into these have identified further challenges. For example, many APR ATS tools only show the repaired program to the students, who then have to compare and modify their program accordingly. Another issue is that APR generally only modifies a few lines, so if the student solution is far from correct, the repair might fail. This can be solved by partial repair, i.e., the program is repaired so at least one additional test-case passes. While this increases the repair rate, it might make hints more difficult or point the students in a non-obvious or even "wrong" direction. The APR can take several minutes, which also makes it unsuitable for interactive ATS. We take a design science approach to define an ATS based on APR that attempts to address the identified challenges. We give a review of the state-of-the-art for the required components, e.g., APR, how to generate hints from differences between two programs. From this, we suggest a three-step roadmap; 1. identify suitable APR-tools, 2. construct an oversized test-suite, and 3. adopt APR to the tutoring context. Daniel Toll, Anna Wingkvist, Morgan Ericsson |
FIE | 2 |
| 2020 | Using source code density to improve the accuracy of automatic commit classification into maintenance activities
Sebastian Hönel, Morgan Ericsson, Welf Löwe, Anna Wingkvist |
J. Syst. Softw. | 4 |
| 2019 | TDMentions: a dataset of technical debt mentions in online postsabstractThe term technical debt is easy to understand as a metaphor, but can quickly grow complex in practice. We contribute with a dataset, TDMentions, that enables researchers to study how developers and end users use the term technical debt in online posts and discussions. The dataset consists of posts from news aggregators and Q&A-sites, blog posts, and issues and commits on GitHub. Morgan Ericsson, Anna Wingkvist |
TechDebt@ICSE | 2 |
| 2019 | Optimization of Software Estimation ModelsabstractIn software engineering, estimations are frequently used to determine expected but yet unknown properties of software development processes or the developed systems, such as costs, time, number of developers, efforts, sizes, and complexities. Plenty of estimation models exist, but it is hard to compare and improve them as software technologies evolve quickly. We suggest an approach to estimation model design and automated optimization allowing for model comparison and improvement based on commonly collected data points. This way, the approach simplifies model optimization and selection. It contributes to a convergence of existing estimation models to meet contemporary software technology practices and provide a possibility for selecting the most appropriate ones. Chris Kopetschny, Morgan Ericsson, Welf Löwe, Anna Wingkvist |
ICSOFT | 4 |
| 2019 | Importance and Aptitude of Source Code Density for Commit Classification into Maintenance ActivitiesabstractCommit classification, the automatic classification of the purpose of changes to software, can support the understanding and quality improvement of software and its development process. We introduce code density of a commit, a measure of the net size of a commit, as a novel feature and study how well it is suited to determine the purpose of a change. We also compare the accuracy of code-density-based classifications with existing size-based classifications. By applying standard classification models, we demonstrate the significance of code density for the accuracy of commit classification. We achieve up to 89% accuracy and a Kappa of 0.82 for the cross-project commit classification where the model is trained on one project and applied to other projects. Such highly accurate classification of the purpose of software changes helps to improve the confidence in software (process) quality analyses exploiting this classification information. Sebastian Hönel, Morgan Ericsson, Welf Löwe, Anna Wingkvist |
QRS | 4 |
| 2018 | Visualizing Programming Session TimelinesabstractLearning programming with tutor tools has grown in popularity. These tools present programming assignments and provide feedback in the form of test-cases and compilation errors. Our timeline visualization of data from one such tool allows us to tell a story about what files were accessed and for how long, in what order files were edited, grown or shrunk, what errors the student ran into, and how those errors were addressed. This can be done without a need to read and replay the entire programming session. In sum, the tool has been used to visualize logs from students that tried to solve programming assignments and we find interesting stories that can help us improve how we address new assignments. Daniel Toll, Anna Wingkvist |
VINCI | 2 |
| 2018 | Quality Models Inside Out: Interactive Visualization of Software Metrics by Means of Joint ProbabilitiesabstractAssessing software quality, in general, is hard; each metric has a different interpretation, scale, range of values, or measurement method. Combining these metrics automatically is especially difficult, because they measure different aspects of software quality, and creating a single global final quality score limits the evaluation of the specific quality aspects and trade-offs that exist when looking at different metrics. We present a way to visualize multiple aspects of software quality. In general, software quality can be decomposed hierarchically into characteristics, which can be assessed by various direct and indirect metrics. These characteristics are then combined and aggregated to assess the quality of the software system as a whole. We introduce an approach for quality assessment based on joint distributions of metrics values. Visualizations of these distributions allow users to explore and compare the quality metrics of software systems and their artifacts, and to detect patterns, correlations, and anomalies. Furthermore, it is possible to identify common properties and flaws, as our visualization approach provides rich interactions for visual queries to the quality models' multivariate data. We evaluate our approach in two use cases based on: 30 real-world technical documentation projects with 20,000 XML documents, and an open source project written in Java with 1000 classes. Our results show that the proposed approach allows an analyst to detect possible causes of bad or good quality. Maria Ulan, Sebastian Hönel, Rafael Messias Martins, Morgan Ericsson, Welf Löwe, Anna Wingkvist, Andreas Kerren |
VISSOFT | 6 |
| 2017 | How Tool Support and Peer Scoring Improved Our Students' Attitudes Toward Peer ReviewsabstractWe wanted to introduce peer reviews for the final report in a course on Software Testing. The students had however experienced issues with peer reviews in a previous course which made this a challenge. To get a better understanding of the situation, we distributed a pre-questionnaire to the students. 48 of the 83 students provided their expectations on peer reviews. To deal with some of the perceived issues, we developed a peer review tool where we introduce anonymity, grading of reviews, teacher interventions, as well as let students score and comment on the reviews they receive. In total, 67 reports were submitted by 83 students and 325 reviews were completed. The post-questionnaire was answered by 48 students (not necessarily the same respondents as for the pre-questionnaire as both were collected anonymously). While 27 of the students expected incorrect feedback only 13 students agreed to have got incorrect feedback in the post-questionnaire. The students reported that they found the feedback from their peers more valuable (+15%) than expected, and 88% of the students reported that they learned from doing peer reviews. Overall, we find that the students' attitudes towards peer reviews have improved. Daniel Toll, Anna Wingkvist |
ITiCSE | 2 |
| 2015 | Detailed Recordings of Student Programming SessionsabstractObservation is important when we teach programming. It can help identify students that struggle, concepts that are not clearly presented during lectures, poor assignments, etc. However, as development tools become more widely available or courses move off-campus and online, we lose our ability to naturally observe students. Online programming environments provide an opportunity to record how students solve assignments and the data recorded allows for in-depth analysis. For example, file activities, mouse movements, text-selections, and text caret movements provide a lot of information on when a programmer collects information and what task is currently worked on. We developed CSQUIZ to allow us to observe students on our online courses through data analysis. Based on our experience with the tool in a course, we find recorded sessions a sufficient replacement for natural observations. Daniel Toll, Tobias Olsson, Morgan Ericsson, Anna Wingkvist |
ITiCSE | 4 |
| 2014 | Mining job ads to find what skills are sought after from an employers' perspective on IT graduatesabstractWe mine job ads to discover what skills are required from an employers' perspective. Some obvious trends appear, such as skills related to web and mobile technology. We aim to uncover more detailed information as the study continues to allow course content to better match the expressed needs. Morgan Ericsson, Anna Wingkvist |
ITiCSE | 2 |
| 2014 | The challenge of teaching students the value of programming best practicesabstractWe investigate the benefits of our programming assignments in correlation to what the students learn and show in their programming solutions. The assignments are supposed to teach the students to use best practices related to program comprehension, but do the programming assignments clearly show the benefits of best practices? We performed an experiment that showed no significant result which suggests that the assignments did not emphasise the value of best practices. As lecturers, we understand that constructing assignments that match the sought after outcome in students learning is a complex task. The experiment provided valuable insights that we will use to improve the assignments to better mirror best practices. Daniel Toll, Tobias Olsson, Anna Wingkvist, Morgan Ericsson |
ITiCSE | 3 |
| 2013 | A Study of the Effect of Data Normalization on Software and Information Quality AssessmentabstractIndirect metrics in quality models define weighted integrations of direct metrics to provide higher-level quality indicators. This paper presents a case study that investigates to what degree quality models depend on statistical assumptions about the distribution of direct metrics values when these are integrated and aggregated. We vary the normalization used by the quality assessment efforts of three companies, while keeping quality models, metrics, metrics implementation and, hence, metrics values constant. We find that normalization has a considerable impact on the ranking of an artifact (such as a class). We also investigate how normalization affects the quality trend and find that normalizations have a considerable effect on quality trends. Based on these findings, we find it questionable to continue to aggregate different metrics in a quality model as we do today. Morgan Ericsson, Welf Löwe, Tobias Olsson, Daniel Toll, Anna Wingkvist |
APSEC (2) | 5 |