VLDB 2026 Research / reviewers in the wild / expert
Mika Mäntylä
dblp:m/MikaMantyla · also Mika V. Mäntylä, Mika Viking Mäntylä
· DBLP profile ↗
97ranked-venue papers
20as first author
29since 2021 · last 2026
0000-0002-2841-5879ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 95 · 20 first-author · 27 since 2021Databases, data management, data science and information retrieval · 12 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AnoMod: A Dataset for Anomaly Detection and Root Cause Analysis in Microservice SystemabstractMicroservice systems (MSS) have become a predominant architectural style for cloud services. Yet the community still lacks high-quality, publicly available datasets for anomaly detection (AD) and root cause analysis (RCA) in MSS. Most benchmarks emphasize performance-related faults and provide only one or two monitoring modalities, limiting research on broader failure modes and cross-modal methods. To address these gaps, we introduce a new multimodal anomaly dataset built on two open-source microservice systems: SocialNetwork and TrainTicket. We design and inject four categories of anomalies (Ano): performance-level, service-level, database-level, and code-level, to emulate realistic anomaly modes. For each scenario, we collect five modalities (Mod): logs, metrics, distributed traces, API responses, and code coverage reports, offering a richer, end-to-end view of system state and inter-service interactions. We name our dataset, reflecting its unique properties, as AnoMod. This dataset enables (1) evaluation of cross-modal anomaly detection and fusion/ablation strategies, and (2) fine-grained RCA studies across service and code regions, supporting end-to-end troubleshooting pipelines that jointly consider detection and localization. Ke Ping, Hamza Bin Mazhar, Yuqing Wang 0002, Mika Mäntylä |
MSR | 5 |
| 2026 | VisualLogAnalyzer: An Interactive Web Application for Multi-Level Log Analysis
Jesse Nyyssölä, Simo Sipilä, Mika Mäntylä |
SANER | 3 |
| 2026 | Token interdependency parsing (Tipping) - A map-reduced based statistical log parserabstractContext: Over the last decade, impressive growth in software adaptations has led to a surge in log data production, making manual log analysis impractical and underscoring the need for automated methods. On the other hand, most automated analysis tools benefit from a component that separates log templates from their parameters, commonly referred to as a “log parser”. Effectiveness and efficiency are inherently competing attributes, such that enhancing one typically requires compromising the other, making the simultaneous maximization of both challenging. Objective: This paper aims to introduce a new log parser capable of processing logs faster than the statistics-based methods while maintaining a respectable effectiveness, named “Tipping”. Method: Tipping combines rule-based tokenizers, interdependency token graphs, strongly connected components, and various techniques to ensure rapid, scalable, and precise log parsing. Furthermore, Tipping was designed with map-reduce principles to be parallelizable, enabling it to run on multiple processing cores to speed up the process. We evaluated Tipping against other statistics-based log parsers in terms of effectiveness, efficiency, and the downstream task of anomaly detection. Results: Accordingly, our evaluations demonstrate that Tipping outperforms most statistics-based parsing methods in both effectiveness and efficiency. More in-depth, Tipping can parse 11 million lines of logs in less than 20 s on a development machine. Furthermore, Tipping ranks among the top performers in terms of effectiveness and efficiency across the Loghub-2k, Loghub-2.0, LogPM, and LogLead benchmarks. Conclusion: Tipping’s robustness, versatility, efficiency, and scalability make it a viable tool for modern automated log analysis. However, Tipping currently cannot operate in online or streaming mode, which limits its applicability in environments where logs arrive continuously and must be processed incrementally in real time. Shayan Hashemi, Mika Mäntylä |
Inf. Softw. Technol. | 2 |
| 2026 | Studying SATD in drone systems with Human-AI collaboration
Leevi Rantala, Lwin Khin Shar, Mika Mäntylä, Wei Minn, Yan Naing Tun |
J. Syst. Softw. | 3 |
| 2025 | SESR-Eval: Dataset for Evaluating LLMs in the Title-Abstract Screening of Systematic ReviewsabstractBackground: The use of large language models (LLMs) in the title-abstract screening process of systematic reviews (SRs) has shown promising results, but suffers from limited performance evaluation. Aims: Create a benchmark dataset to evaluate the performance of LLMs in the title-abstract screening process of SRs. Provide evidence whether using LLMs in title-abstract screening in software engineering is advisable. Method: We start with 169 SR research artifacts and find 24 of those to be suitable for inclusion in the dataset. Using the dataset we benchmark title-abstract screening using 9 LLMs. Results: We present the SESR-Eval (Software Engineering Systematic Review Evaluation) dataset containing 34,528 labeled primary studies, sourced from 24 secondary studies published in software engineering (SE) journals. Most LLMs performed similarly and the differences in screening accuracy between secondary studies are greater than differences between LLMs. The cost of using an LLM is relatively low - less than 40 per secondary study even for the most expensive model. Conclusions: Our benchmark enables monitoring AI performance in the screening task of SRs in software engineering. At present, LLMs are not yet recommended for automating the title-abstract screening process, since accuracy varies widely across secondary studies, and no LLM managed a high recall with reasonable precision. In future, we plan to investigate factors that influence LLM screening performance between studies. Aleksi Huotala, Miikka Kuutila, Mika Mäntylä |
ESEM | 3 |
| 2025 | Cross-System Software Log-based Anomaly Detection Using Meta-LearningabstractModern software systems produce vast amounts of logs, serving as an essential resource for anomaly detection. Artificial Intelligence for IT Operations (AIOps) tools have been developed to automate the process of log-based anomaly detection for software systems. Three practical challenges are widely recognized in this field: data labeling costs, evolving logs in dynamic systems, and adaptability across different systems. In this paper, we propose CroSysLog, an AIOps tool for log-event level anomaly detection, considering these challenges. Following prior approaches, CroSysLog uses a neural representation approach to gain a nuanced understanding of logs and generate representations for individual log events accordingly. CroSysLog can be trained on source systems with sufficient labeled logs from open datasets to achieve robustness, and then efficiently adapt to target systems with a few labeled log events for effective anomaly detection. We evaluate CroSysLog using open datasets of four large-scale distributed supercomputing systems: BGL, Thunderbird, Liberty, and Spirit. We used random log splits, maintaining the chronological order of consecutive log events, from these systems to train and evaluate CroSysLog. These splits were widely distributed across a one/two-year span of each system's log collection duration, capturing the evolving nature of the logs in each system. Our results show that, after training CroSysLog on Liberty and BGL as source systems, CroSysLog can efficiently adapt to target systems Thunderbird and Spirit using a few labeled log events from each target system, effectively performing anomaly detection for these target systems. The results demonstrate that CroSysLog is a practical, scalable, and adaptable tool for log-event level anomaly detection in operational and maintenance contexts of software systems. Yuqing Wang 0002, Mika Mäntylä, Jesse Nyyssölä, Ke Ping |
SANER | 2 |
| 2025 | Detection, classification and prevalence of self-admitted aging debtabstractAbstract Context Previous research on software aging is limited, with a focus on dynamic runtime indicators like memory and performance, often neglecting evolutionary indicators like source code comments and narrowly examining legacy issues within the Technical Debt (TD) context. Objective We introduce the concept of Aging Debt (AD), representing the increased maintenance efforts and costs needed to keep software updated. We study AD through Self-Admitted Aging Debt (SAAD) observed in source code comments left by software developers. Method We employ a mixed-methods approach, combining qualitative and quantitative analyses to detect and measure AD in software. This includes framing SAAD patterns from the source code comments after analysing the source code context, then utilizing the SAAD patterns to detect SAAD comments. In the process, we develop a taxonomy for SAAD that reflects the temporal aging of software and its associated debt. Then we utilize the taxonomy to quantify the different types of AD prevalent in Open Source Software (OSS) repositories. Results Our proposed taxonomy categorizes evolutionary software aging into Active and Dormant types. Our extensive analysis of over 9,000+ OSS repositories reveals that more than 21% repositories exhibit signs of SAAD as observed from our gold standard SAAD dataset. Notably, Dormant AD emerges as the predominant category, highlighting a critical but often overlooked aspect of software maintenance. Conclusion As software volume grows annually, so do evolutionary aging and maintenance challenges; our proposed taxonomy can aid researchers in detailed software aging studies and help practitioners develop improved and proactive maintenance strategies. Murali Sridharan, Mika Mäntylä, Leevi Rantala |
Empir. Softw. Eng. | 2 |
| 2025 | Credtwi: Investigating Social Media Credibility with a Browser PluginabstractPeople now look for information online and on social media for everyday problems. Organizations and malevolent actors have taken the opportunity to spread misinformation/disinformation. It is increasingly important to understand the credibility of online information. We designed and implemented a research browser plugin, Credtwi. It injects credibility questionnaires directly into the user’s Twitter feed, enabling crowdsourced data collection. We carried out a week-long field study where participants assessed the credibility of tweets on various topics. We provide insights into information credibility in the Twitter ecosystem by analyzing the assessments and study questionnaires. The participants’ perception of Twitter as a credible information source decreased after using Credtwi. Our results suggest that the author’s verification status and bio are the most important factors for their perceived credibility. Finally, we discovered significant differences between the assessments of the different genders. Our results contribute to the research on online social media content credibility. Eetu Huusko, Nazanin Nakhaie Ahooie, Miikka Kuutila, Aleksi Huotala, Mika Mäntylä, Simo Hosio |
Int. J. Hum. Comput. Interact. | 6 |
| 2025 | Research artifacts in secondary studies: A systematic mapping in software engineeringabstractContext: Systematic reviews (SRs) summarize state-of-the-art evidence in science, including software engineering (SE). Objective: Our objective is to evaluate how SRs report research artifacts and to provide a comprehensive list of these artifacts. Method: We examined 537 secondary studies published between 2013 and 2023 to analyze the availability and reporting of research artifacts. Results: Our findings indicate that only 31.5% of the reviewed studies include research artifacts. Encouragingly, the situation is gradually improving, as our regression analysis shows a significant increase in the availability of research artifacts over time. However, in 2023, just 62.0% of secondary studies provide a research artifact while an even lower percentage, 30.4% use a permanent repository with a digital object identifier (DOI) for storage. Conclusion: To enhance transparency and reproducibility in SE research, we advocate for the mandatory publication of research artifacts in secondary studies. Aleksi Huotala, Miikka Kuutila, Mika Mäntylä |
Inf. Softw. Technol. | 3 |
| 2024 | The Promise and Challenges of Using LLMs to Accelerate the Screening Process of Systematic ReviewsabstractContext: Systematic review (SR) is a popular research method in software engineering (SE). However, conducting an SR takes an average of 67 weeks. Thus, automating any step of the SR process could reduce the effort associated with SRs. Objective: Our objective is to investigate the extent to which Large Language Models (LLMs) can accelerate title-abstract screening by (1) simplifying abstracts for human screeners, and (2) automating title-abstract screening entirely. Method: We performed an experiment where human screeners performed title-abstract screening for 20 papers with both original and simplified abstracts from a prior SR. The experiment with human screeners was reproduced by instructing GPT-3.5 and GPT-4 LLMs to perform the same screening tasks. We also studied whether different prompting techniques (Zero-shot (ZS), One-shot (OS), Few-shot (FS), and Few-shot with Chain-of-Thought (FS-CoT) prompting) improve the screening performance of LLMs. Lastly, we studied if redesigning the prompt used in the LLM reproduction of title-abstract screening leads to improved screening performance. Results: Text simplification did not increase the screeners’ screening performance, but reduced the time used in screening. Screeners’ scientific literacy skills and researcher status predict screening performance. Some LLM and prompt combinations perform as well as human screeners in the screening tasks. Our results indicate that a more recent LLM (GPT-4) is better than its predecessor LLM (GPT-3.5). Additionally, Few-shot and One-shot prompting outperforms Zero-shot prompting. Conclusion: Using LLMs for text simplification in the screening process does not significantly improve human performance. Using LLMs to automate title-abstract screening seems promising, but current LLMs are not significantly more accurate than human screeners. To recommend the use of LLMs in the screening process of SRs, more research is needed. We recommend future SR studies to publish replication packages with screening data to enable more conclusive experimenting with LLM screening. Aleksi Huotala, Miikka Kuutila, Paul Ralph, Mika Mäntylä |
EASE | 4 |
| 2024 | What Makes Programmers Laugh? Exploring the Submissions of the Subreddit r/ProgrammerHumorabstractBackground: Humor is a fundamental part of human communication, with prior work linking positive humor in the workplace to positive outcomes, such as improved performance and job satisfaction. Aims: This study aims to investigate programming-related humor in a large social media community. Methodology: We collected 139,718 submissions from Reddit subreddit r/ProgrammerHumor. Both textual and image-based (memes) submissions were considered. The image data was processed with OCR to extract text from images for NLP analysis. Multiple regression models were built to investigate what makes submissions humorous. Additionally, a random sample of 800 submissions was labeled by human annotators regarding their relation to theories of humor, suitability for the workplace, the need for programming knowledge to understand the submission, and whether images in image-based submissions added context to the submission. Results: Our results indicate that predicting the humor of software developers is difficult. Our best regression model was able to explain only 10% of the variance. However, statistically significant differences were observed between topics, submission times, and associated humor theories. Our analysis reveals that the highest submission scores are achieved by image-based submissions that are created during the winter months in the northern hemisphere, between 2-3pm UTC on weekends, which are distinctly related to superiority and incongruity theories of humor, and are about the topic of "Learning". Conclusions: Predicting humor with natural language processing methods is challenging. We discuss the benefits and inherent difficulties in assessing perceived humor of submissions, as well as possible avenues for future work. Additionally, our replication package should help future studies and can act as a joke repository for the software industry and education. Miikka Kuutila, Leevi Rantala, Simo Hosio, Mika Mäntylä |
ESEM | 5 |
| 2024 | A Dataset of Microservices-based Open-Source ProjectsabstractResearchers in the microservices community often resort to demonstrating the impact of their proposed advancements on custom-made microservices projects. This is a possible source of bias that can reduce the trustworthiness of the results. Moreover, it is hard to compare advances in small projects, often developed due to lack of time. It is common across disciplines to recognize benchmarks that mitigate bias and unify the advancements' impact. To facilitate the identification of available open-source microservice projects (OSS-MS), we performed a comprehensive study to identify, curate, and catalog OSS-MS. We started with 389559 projects and filtered them down to 3804 projects that we manually labeled. After manual labeling, our dataset contains 378 projects with three or more microservices and with over 100 commits. We document the projects from many perspectives, including project size, platform, number of contributors, project purpose, and foundation support. This dataset can serve researchers as a roadmap to identify benchmarks, as our dataset can be used to answer questions such as whether the number of services impacts the issue count. Dario Amoroso d'Aragona, Alexander Bakhtin, Xiaozhou Li 0002, Ruoyu Su, Lauren Adams, Ernesto Aponte, Francis Boyle, Patrick Boyle, Rachel Koerner, Joseph Lee, Fangchao Tian, Yuqing Wang 0002, Jesse Nyyssölä, Ernesto Quevedo Caballero, Md Shahidur Rahaman, Amr S. Abdelfattah, Mika Mäntylä, Tomás Cerný, Davide Taibi 0001 |
MSR | 17 |
| 2024 | From Reinvention to Reuse: An Empirical Example Study on Technical Debt Dataset
Leevi Rantala, Mika Mäntylä, Murali Sridharan |
PROFES | 2 |
| 2024 | Event-level Anomaly Detection on Software logs: Role of Algorithm, Threshold, and Window SizeabstractAnomaly detection on software logs has proven to be an efficient way to identify root causes of software related issues. Rather than identifying anomalies in sequences, this study aims to utilize an event-based anomaly detection approach to perform a ground-truth evaluation on the BlueGene/L (BGL) dataset. In the study we determined two approaches for adjusting the threshold of event-based anomaly detection: one that accounts for the absolute number of predictions (top-k) and another that accounts for cumulative probability of those predictions (top-p). These approaches were assessed on precision, recall and F1-score, and we found that top-k generally yields better results. We found that on deep learning (DL) models adjusting the threshold has the largest effect on F1-score among all our test configurations. If possible, we recommend using the top-k approach with a kvalue tailored to the dataset. Adjusting the threshold increased the F1-score from 0.846 of the baseline to 0.957 with optimal k-value. Regarding window size, we found that the optimal size is dependent on the model. N-Gram model benefits from as short sequences as possible while the DL models experience improvement up to the window size of 5. Adjusting the mask position of the prediction within the window can also improve the F1-score. Regarding the algorithm selection, with the best configuration the results show similar F1-scores on LSTM (0.9727) and CNN (0.9731) while the N-Gram (0.970) and Transformer (0.9675) model performed somewhat worse. As the best configurations allow us to consistently reach F1-scores of over 0.970 on BGL, we reached state of the art results on event-based anomaly detection. Jesse Nyyssölä, Mika Mäntylä |
QRS | 2 |
| 2024 | Speed and Performance of Parserless and Unsupervised Anomaly Detection Methods on Software LogsabstractSoftware log analysis can be laborious and time consuming. Time and labeled data are usually lacking in industrial settings. This paper studies unsupervised and time efficient methods for anomaly detection. We study two custom and two established models. The custom models are: an OOV (Out-Of-Vocabulary) detector, which counts the terms in the test data that are not present in the training data, and the Rarity Model (RM), which calculates a rarity score for terms based on their infrequency. The established models are KMeans and Isolation Forest. The models are evaluated on four public datasets (BGL, Thunderbird, Hadoop, HDFS) with three different representation techniques for the log messages (Words, character Trigrams, Parsed events). For training, we used both normal-only data, which is free of all anomalies, and unfiltered data, which contains both normal and anomalous instances. We used primarily the AUC-ROC metric for evaluation due to challenges in setting a threshold but we also include F1-scores for further insight. Different configurations are advised based on specific requirements. When training data is unfiltered, includes both normal and anomalous instances, the most effective combination is the Isolation Forest with event representation, achieving an AUC-ROC of 0.829. If it’s possible to create a normal-only training dataset, combining the Out-Of-Vocabulary (OOV) detector with trigram representation yields the highest AUC-ROC of 0.846. For speed considerations, the OOV detector is optimal for filtered data, while the Rarity Model is the best choice for unfiltered data. Jesse Nyyssölä, Mika Mäntylä |
QRS | 2 |
| 2024 | LogPM: Character-Based Log Parser BenchmarkabstractLog parsers transform free-form textual log messages into categorical data and are important tools in automated log analysis pipelines. However, selecting a suitable log parsing algorithm poses a formidable obstacle, thereby underscoring the importance of having a comprehensive benchmark to facilitate decision-making. This paper introduces a novel log parsing benchmark, focusing on predicted template precision at the character level rather than accurate grouping used in the past. We present a new metric called Parameter Mask Agreement that measures template accuracy at the character level alongside a dataset tailored for the task. We identified several challenges that parsers encounter for each dataset, which can aid in developing new log parsers. Moreover, a small empirical study was conducted using the proposed benchmark, evaluating the performance of three renowned parsers: Drain, Spell, and Lenma. The findings revealed that Lenma demonstrated the highest parsing accuracy, whereas Drain exhibited superior parsing speed. Finally, we propose that our benchmark is more appropriate than previous approaches in scenarios where accurate template detection is essential and computational efficiency needs to be assessed. Shayan Hashemi, Jesse Nyyssölä, Mika Mäntylä |
SANER | 3 |
| 2024 | LogLead - Fast and Integrated Log Loader, Enhancer, and Anomaly DetectorabstractThis paper introduces LogLead, a tool designed for efficient log analysis benchmarking. LogLead combines three essential steps in log processing: loading, enhancing, and anomaly detection. The tool leverages Polars, a high-speed DataFrame library. We currently have Loaders for eight systems that are publicly available (HDFS, Hadoop, BGL, Thunderbird, Spirit, Liberty, TrainTicket, and GC Webshop). We have multiple enhancers with three parsers (Drain, Spell, LenMa), Bert embedding creation and other log representation techniques like bag-of-words. LogLead integrates to five supervised and four unsupervised machine learning algorithms for anomaly detection from SKLearn. By integrating diverse datasets, log representation methods and anomaly detectors, LogLead facilitates comprehensive benchmarking in log analysis research. We show that log loading from raw file to dataframe is over 10x faster with LogLead compared to past solutions. We demonstrate roughly 2x improvement in Drain parsing speed by off-loading log message normalization to LogLead. Our brief benchmarking on HDFS indicates that log representations extending beyond the bag-of-words approach offer limited additional benefits. Tool URL: https://github.com/EvoTestOps/LogLead. Mika Mäntylä, Yuqing Wang 0002, Jesse Nyyssölä |
SANER | 1 |
| 2024 | OneLog: towards end-to-end software log anomaly detectionabstractAbstract With the growth of online services, IoT devices, and DevOps-oriented software development, software log anomaly detection is becoming increasingly important. Prior works mainly follow a traditional four-staged architecture (Preprocessor, Parser, Vectorizer, and Classifier). This paper proposes OneLog, which utilizes a single deep neural network instead of multiple separate components. OneLog harnesses convolutional neural network (CNN) at the character level to take digits, numbers, and punctuations, which were removed in prior works, into account alongside the main natural language text. We evaluate our approach in six message- and sequence-based data sets: HDFS, Hadoop, BGL, Thunderbird, Spirit, and Liberty. We experiment with Onelog with single-, multi-, and cross-project setups. Onelog offers state-of-the-art performance in our datasets. Onelog can utilize multi-project datasets simultaneously during training, which suggests our model can generalize between datasets. Multi-project training also improves Onelog performance making it ideal when limited training data is available for an individual project. We also found that cross-project anomaly detection is possible with a single project pair (Liberty and Spirit). Analysis of model internals shows that one log has multiple modes of detecting anomalies and that the model learns manually validated parsing rules for the log messages. We conclude that character-based CNNs are a promising approach toward end-to-end learning in log anomaly detection. They offer good performance and generalization over multiple datasets. We will make our scripts publicly available upon the acceptance of this paper. Shayan Hashemi, Mika Mäntylä |
Autom. Softw. Eng. | 2 |
| 2024 | Keyword-labeled self-admitted technical debt and static code analysis have significant relationship but limited overlapabstractAbstract Technical debt presents sub-optimal choices made in development, which are beneficial in the short term but not in the long run. Consciously admitted debt, which is marked with a keyword, e.g., TODO, is called keyword-labeled self-admitted technical debt (KL-SATD). KL-SATD can lead to adverse effects in software development, e.g., to a rise in complexity within the developed software. We investigated the relationship between KL-SATD from source code comments and reports from the highly popular industrial program analysis tool SonarQube. The goal was to find which SonarQube metrics and issues are related to KL-SATD introduction and removal and how many KL-SATD in the context of an issue addresses that issue. We performed a study with 33 software repositories. We analyzed the changes in SonarQube reports (sqale index, reliability and security remediation metrics, and SonarQube issues) and the relationship to KL-SATD addition and removal with mixed model analysis. We manually annotated a sample to investigate how many KL-SATD comments are in the context of SonarQube issues and how many address them directly. KL-SATD is associated with a reduction in code maintainability measured with SonarQube’s sqale index. KL-SATD removal is associated with an increase in code maintainability (sqale index) and reliability measured with SonarQube’s reliability remediation effort. The introduction and removal of KL-SATD have a predominantly relationship with code smells, and not with vulnerabilities and bugs. Manual annotation revealed that 36% of KL-SATD comments are in the context of a SonarQube issue, but only 15% of the comment address an issue. This means that despite of statistical relationship between KL-SATD comments and SonarQube reports there is a large set of KL-SATD comments that are in areas that Sonarqube reports as clean or free of maintainability issues. KL-SATD introduction and removal are connected mainly to code smells, connecting them to maintainability rather than reliability or security. This is reinforced by the relationship with the sqale index, as well as the dominance of code smells in SonarQube issues. Many KL-SATD issues have characteristics going beyond static analysis tools and require future studies extending the capabilities of the current tools. As KL-SATD comments and SonarQube reports appear to have limited overlap, it suggests that they are complementary and both are needed for getting a comprehensive view coverage of code maintainability. The study also presents rules violations developers should be aware of regarding KL-SATD introduction and removal. Leevi Rantala, Mika Mäntylä, Valentina Lenarduzzi |
Softw. Qual. J. | 2 |
| 2023 | It is an online platform and not the real world, I don't care much: Investigating Twitter Profile Credibility With an Online Machine Learning-Based ToolabstractSocial media is now an important source of everyday information. Given the plethora of scandals concerning the rapid spread of misinformation and disinformation on social media, the credibility of the content on these platforms is now a pivotal research area. Much of the existing work on social media credibility focuses on content credibility. In this study, however, we focus on the credibility of the profile as the virtual representation of the content author. We developed a real-time machine-learning-based online tool that assesses the credibility of profiles on Twitter, one of the most common and versatile social media platforms. To investigate user perceptions on credibility-related issues, we used our tool as a stimulus for people to reflect on their profile’s credibility and collected 100 responses. The combination of our quantitative and qualitative analysis reveals that the latest tweets and retweet behavior are two of the most critical factors for profile credibility. It is also observed that people demonstrate a limited interest in their profile credibility but agree that the author’s credibility is of paramount importance. With an open-source tool to assess user credibility on Twitter and a user study to establish its utility, we contribute a timely piece of research on the topic of online credibility. Ville Paananen, Sharadhi Alape Suryanarayana, Eetu Huusko, Miikka Kuutila, Mika Mäntylä, Simo Hosio |
CHIIR | 6 |
| 2023 | PENTACET data - 23 Million Contextual Code Comments and 250,000 SATD commentsabstractMost Self-Admitted Technical Debt (SATD) research utilizes explicit SATD features such as ‘TODO’ and ‘FIXME’ for SATD detection. A closer look reveals several SATD research uses simple SATD (‘Easy to Find’) code comments without contextual data (preceding and succeeding source code context). This work addresses this gap through PENTACET (or 5C dataset) data. PENTACET is a large Curated Contextual Code Comments per Contributor and the most extensive SATD data. We mine 9,096 Open Source Software Java projects totaling over 400 million LOC. The outcome is a dataset with 23 million code comments, preceding and succeeding source code context for each comment, and more than 250,000 SATD comments, including both ‘Easy to Find’ and ‘Hard to Find’ SATD. We believe PENTACET data will further SATD research using Artificial Intelligence techniques. Murali Sridharan, Leevi Rantala, Mika Mäntylä |
MSR | 3 |
| 2022 | How to Configure Masked Event Anomaly Detection on Software Logs?abstractSoftware Log anomaly event detection with masked event prediction has various technical approaches with countless configurations and parameters. Our objective is to provide a baseline of settings for similar studies in the future. The models we use are the N-Gram model, which is a classic approach in the field of natural language processing (NLP), and two deep learning (DL) models long short-term memory (LSTM) and convolutional neural network (CNN). For datasets we used four datasets Profilence, BlueGene/L (BGL), Hadoop Distributed File System (HDFS) and Hadoop. Other settings are the size of the sliding window which determines how many surrounding events we are using to predict a given event, mask position (the position within the window we are predicting), the usage of only unique sequences, and the portion of data that is used for training. The results show clear indications of settings that can be generalized across datasets. The performance of the DL models does not deteriorate as the window size increases while the N-Gram model shows worse performance with large window sizes on the BGL and Profilence datasets. Despite the popularity of Next Event Prediction, the results show that in this context it is better not to predict events at the edges of the subsequence, i.e., first or last event, with the best result coming from predicting the fourth event when the window size is five. Regarding the amount of data used for training, the results show differences across datasets and models. For example, the N-Gram model appears to be more sensitive toward the lack of data than the DL models. Overall, for similar experimental setups we suggest the following general baseline: Window size 10, mask position second to last, do not filter out non-unique sequences, and use a half of the total data for training. Jesse Nyyssölä, Mika Mäntylä, Martín Varela 0001 |
ICSME | 2 |
| 2022 | SoCCMiner: A Source Code-Comments and Comment-Context MinerabstractNumerous tools exist for mining source code and software development process metrics. However, very few publicly available tools focus on source code comments, a crucial software artifact. This paper presents SoCCMiner (Source Code-Comments and Comment-Context Miner), a tool that offers multiple mining pipelines. It is the first readily available (plug-and-play) and customizable open-source tool for mining source code contextual information of comments at different granularities (Class comments, Method comments, Interface comments, and other granular comments). Mining comments at different source code granularities can aid researchers and practitioners working in a host of applications that focus on source code comments, such as Self-Admitted Technical Debt, Program Comprehension, and other applications. Furthermore, SoCCMiner is highly adaptable and extendable to include additional attributes and support other programming languages. This prototype supports the Java programming language. Murali Sridharan, Mika Mäntylä, Maëlick Claes, Leevi Rantala |
MSR | 2 |
| 2022 | SiaLog: detecting anomalies in software execution logs using the siamese networkabstractAbstract Detecting anomalies in software logs has become a notable concern for software engineers and maintainers as they represent anomalies in software execution paths and states. This paper propose a novel anomaly detection approach based on the Siamese network on top of Recurrent Neural Networks(RNN). Accordingly, we introduce a novel training pair generation algorithm to train the Siamese network which reduces generated training significantly while maintaining the $$F_1$$ F 1 score. Additionally, we propose a hybrid model by combining the Siamese network with a traditional feedforward neural network to make end-to-end training possible, reducing engineering effort in setting up a deep-learning-based log anomaly detector. Furthermore, we provides validations of the approach on the Hadoop Distributed File System (HDFS), Blue Gene/L (BGL), and Hadoop map-reduce task log datasets. To the best of our knowledge, the proposed approach outperforms other methods on the same dataset at the $$F_1$$ F 1 scores of respectively 0.99, 0.99, and 0.94 on HDFS, BGL, and Hadoop datasets, resulting in a new state-of-the-art performance.To further evaluate the proposed method, we examine our method’s robustness to log evolutions by evaluating the model on synthetically evolved log sequences; we got the $$F_1$$ F 1 score of 0.95 on the HDFS dataset at the noise ratio of $$20\%$$ 20 % . Finally, we dive deep into some of the side benefits of the Siamese network. Accordingly, we introduce an unsupervised log evolution monitoring method alongside a visualization technique that facilitates model interpretability. Shayan Hashemi, Mika Mäntylä |
Autom. Softw. Eng. | 2 |
| 2022 | Introduction to the Special Issue on: Grey Literature and Multivocal Literature Reviews (MLRs) in software engineering
Vahid Garousi, Austen Rainer, Michael Felderer, Mika Mäntylä |
Inf. Softw. Technol. | 4 |
| 2022 | Test automation maturity improves product quality - Quantitative study of open source projects using continuous integrationabstractThe popularity of continuous integration (CI) is increasing as a result of market pressure to release product features or updates frequently. The ability of CI to deliver quality at speed depends on reliable test automation. In this paper, we present an empirical study to observe the effect of test automation maturity (assessed by standard best practices in the literature) on product quality, test automation effort, and release cycle in the CI context of open source projects. We run our test automation maturity survey and got responses from 37 open source java projects. We also mined software repositories of the same projects. The main results of regression analysis reveal that, higher levels of test automation maturity are positively associated with higher product quality (p-value=0.000624) and shorter release cycle (p-value=0.01891); There is no statistically significant evidence of increased test automation effort due to higher levels of test automation maturity and product quality. Thus, we conclude that, a potential benefit of improving test automation maturity (using standard best practices) is product quality improvement and release cycle acceleration in the CI context of open source projects. We encourage future research to extend our findings by adding more datasets with different programming languages and CI tools, closed source projects, and large-scale industrial projects. Our recommendation to practitioners (in the similar CI context) is to utilize standard best practices to improve test automation maturity. Yuqing Wang 0002, Mika Mäntylä, Jouni Markkula |
J. Syst. Softw. | 2 |
| 2022 | Improving test automation maturity: A multivocal literature reviewabstractAbstract Mature test automation is key for achieving software quality at speed. In this paper, we present a multivocal literature review with the objective to survey and synthesize the guidelines given in the literature for improving test automation maturity. We selected and reviewed 81 primary studies, consisting of 26 academic literature and 55 grey literature sources. From primary studies, we extracted 26 test automation best practices (e.g., Define an effective test automation strategy, Set up good test environments, and Develop high‐quality test scripts) and collected many pieces of advice (e.g., in forms of implementation/improvement approaches, technical techniques, concepts, and experience‐based heuristics) on how to conduct these best practices. We made main observations: (1) There are only six best practices whose positive effect on maturity improvement have been evaluated by academic studies using formal empirical methods; (2) several technical related best practices in this MLR were not presented in test maturity models; (3) some best practices can be linked to success factors and maturity impediments proposed by other scholars; (4) most pieces of advice on how to conduct proposed best practices were identified from experience studies and their effectiveness need to be further evaluated with cross‐site empirical evidence using formal empirical methods; (5) in the literature, some advice on how to conduct certain best practices are conflicting, and some advice on how to conduct certain best practices still need further qualitative analysis. Yuqing Wang 0002, Mika Mäntylä, Jouni Markkula, Päivi Raulamo-Jurvanen |
Softw. Test. Verification Reliab. | 2 |
| 2021 | Data Balancing Improves Self-Admitted Technical Debt DetectionabstractA high imbalance exists between technical debt and non-technical debt source code comments. Such imbalance affects Self-Admitted Technical Debt (SATD) detection performance, and existing literature lacks empirical evidence on the choice of balancing technique. In this work, we evaluate the impact of multiple balancing techniques, including Data level, Classifier level, and Hybrid, for SATD detection in Within-Project and Cross-Project setup. Our results show that the Data level balancing technique SMOTE or Classifier level Ensemble approaches Random Forest or XGBoost are reasonable choices depending on whether the goal is to maximize Precision, Recall, F1, or AUC-ROC. We compared our best-performing model with the previous SATD detection benchmark (cost-sensitive Convolution Neural Network). Interestingly the top-performing XGBoost with SMOTE sampling improved the Within-project F1 score by 10% but fell short in Cross-Project set up by 9%. This supports the higher generalization capability of deep learning in Cross-Project SATD detection, yet while working within individual projects, classical machine learning algorithms can deliver better performance. We also evaluate and quantify the impact of duplicate source code comments in SATD detection performance. Finally, we employ SHAP and discuss the interpreted SATD features. We have included the replication package1and shared a web-based SATD prediction tool2with the balancing techniques in this study. Murali Sridharan, Mika Mäntylä, Leevi Rantala, Maëlick Claes |
MSR | 2 |
| 2021 | Individual differences limit predicting well-being and productivity using software repositories: a longitudinal industrial studyabstractReports of poor work well-being and fluctuating productivity in software engineering have been reported in both academic and popular sources. Understanding and predicting these issues through repository analysis might help manage software developers' well-being. Our objective is to link data from software repositories, that is commit activity, communication, expressed sentiments, and job events, with measures of well-being obtained with a daily experience sampling questionnaire. To achieve our objective, we studied a single software project team for eight months in the software industry. Additionally, we performed semi-structured interviews to explain our results. The acquired quantitative data are analyzed with generalized linear mixed-effects models with autocorrelation structure. We find that individual variance accounts for most of the $R^2$ values in models predicting developers' experienced well-being and productivity. In other words, using software repository variables to predict developers' well-being or productivity is challenging due to individual differences. Prediction models developed for each developer individually work better, with fixed effects $R^2$ value of up to 0.24. The semi-structured interviews give insights into the well-being of software developers and the benefits of chat interaction. Our study suggests that individualized prediction models are needed for well-being and productivity prediction in software development. Miikka Kuutila, Mika Mäntylä, Maëlick Claes, Marko Elovainio, Bram Adams |
Empir. Softw. Eng. | 2 |
| 2020 | Prevalence, Contents and Automatic Detection of KL-SATD
Leevi Rantala, Mika Mäntylä, David Lo 0001 |
SEAA | 2 |
| 2020 | Software Test Automation Maturity: A Survey of the State of the PracticeabstractThe software industry has seen an increasing interest in test automation. In this paper, we present a test automation maturity survey serving as a self-assessment for practitioners. Based on responses of 151 practitioners coming from above 101 organizations in 25 countries, we make observations regarding the state of the practice of test automation maturity: a) The level of test automation maturity in different organizations is differentiated by the practices they adopt; b) Practitioner reported the quite diverse situation with respect to different practices, e.g., 85\% practitioners agreed that their test teams have enough test automation expertise and skills, while 47\% of practitioners admitted that there is lack of guidelines on designing and executing automated tests; c) Some practices are strongly correlated and/or closely clustered; d) The percentage of automated test cases and the use of Agile and/or DevOps development models are good indicators for a higher test automation maturity level; (e) The roles of practitioners may affect response variation, e.g., QA engineers give the most optimistic answers, consultants give the most pessimistic answers. Our results give an insight into present test automation processes and practices and indicate chances for further improvement in the present industry. Yuqing Wang 0002, Mika Mäntylä, Serge Demeyer, Kristian Wiklund, Sigrid Eldh, Tatu Kairi |
ICSOFT | 2 |
| 2020 | 20-MAD: 20 Years of Issues and Commits of Mozilla and Apache DevelopmentabstractData of long-lived and high profile projects is valuable for research on successful software engineering in the wild. Having a dataset with different linked software repositories of such projects, enables deeper diving investigations. This paper presents 20-MAD, a dataset linking the commit and issue data of Mozilla and Apache projects. It includes over 20 years of information about 765 projects, 3.4M commits, 2.3M issues, and 17.3M issue comments, and its compressed size is over 6 GB. The data contains all the typical information about source code commits (e.g., lines added and removed, message and commit time) and issues (status, severity, votes, and summary). The issue comments have been pre-processed for natural language processing and sentiment analysis. This includes emoticons and valence and arousal scores. Linking code repository and issue tracker information, allows studying individuals in two types of repositories and provide more accurate time zone information for issue trackers as well. To our knowledge, this the largest linked dataset in size and in project lifetime that is not based on GitHub. Maëlick Claes, Mika Mäntylä |
MSR | 2 |
| 2020 | Time pressure in software engineering: A systematic review
Miikka Kuutila, Mika Mäntylä, Umar Farooq 0006, Maëlick Claes |
Inf. Softw. Technol. | 2 |
| 2020 | Predicting technical debt from commit contents: reproduction and extension with automated feature selectionabstractAbstract Self-admitted technical debt refers to sub-optimal development solutions that are expressed in written code comments or commits. We reproduce and improve on a prior work by Yan et al. (2018) on detecting commits that introduce self-admitted technical debt. We use multiple natural language processing methods: Bag-of-Words, topic modeling, and word embedding vectors. We study 5 open-source projects. Our NLP approach uses logistic Lasso regression from Glmnet to automatically select best predictor words. A manually labeled dataset from prior work that identified self-admitted technical debt from code level commits serves as ground truth. Our approach achieves + 0.15 better area under the ROC curve performance than a prior work, when comparing only commit message features, and + 0.03 better result overall when replacing manually selected features with automatically selected words. In both cases, the improvement was statistically significant (p< 0.0001). Our work has four main contributions, which are comparing different NLP techniques for SATD detection, improved results over previous work, showing how to generate generalizable predictor words when using multiple repositories, and producing a list of words correlating with SATD. As a concrete result, we release a list of the predictor words that correlate positively with SATD, as well as our used datasets and scripts to enable replication studies and to aid in the creation of future classifiers. Leevi Rantala, Mika Mäntylä |
Softw. Qual. J. | 2 |
| 2019 | Practitioner Evaluations on Software Testing ToolsabstractIn software engineering practice, evaluating and selecting the software testing tools that best fit the project at hand is an important and challenging task. In scientific studies of software engineering, practitioner evaluations and beliefs have recently gained interest, and some studies suggest that practitioners find beliefs of peers more credible than empirical evidence. To study how software practitioners evaluate testing tools, we applied online opinion surveys (n=89). We analyzed the reliability of the opinions utilizing Krippendorff's alpha, intra-class correlation coefficient (ICC), and coefficients of variation (CV). Negative binomial regression was used to evaluate the effect of demographics. We find that opinions towards a specific tool can be conflicting. We show how increasing the number of respondents improves the reliability of the estimates measured with ICC. Our results indicate that on average, opinions from seven experts provide a moderate level of reliability. From demographics, we find that technical seniority leads to more negative evaluations. To improve the understanding, robustness, and impact of the findings, we need to conduct further studies by utilizing diverse sources and complementary methods. Päivi Raulamo-Jurvanen, Simo Hosio, Mika Mäntylä |
EASE | 3 |
| 2019 | A Self-assessment Instrument for Assessing Test Automation MaturityabstractTest automation is important in the software industry but self-assessment instruments for assessing its maturity are not sufficient. The two objectives of this study are to synthesize what an organization should focus to assess its test automation; develop a self-assessment instrument (a survey) for assessing test automation maturity and scientifically evaluate it. We carried out the study in four stages. First, a literature review of 25 sources was conducted. Second, the initial instrument was developed. Third, seven experts from five companies evaluated the initial instrument. Content Validity Index and Cognitive Interview methods were used. Fourth, we revised the developed instrument. Our contributions are as follows: (a) we collected practices mapped into 15 key areas that indicate where an organization should focus to assess its test automation; (b) we developed and evaluated a self-assessment instrument for assessing test automation maturity; (c) we discuss important topics such as response bias that threatens self-assessment instruments. Our results help companies and researchers to understand and improve test automation practices and processes. Yuqing Wang 0002, Mika Mäntylä, Sigrid Eldh, Jouni Markkula, Kristian Wiklund, Tatu Kairi, Päivi Raulamo-Jurvanen, Antti Haukinen |
EASE | 2 |
| 2019 | Applying Surveys and Interviews in Software Test Tool Evaluation
Päivi Raulamo-Jurvanen, Simo Hosio, Mika Mäntylä |
PROFES | 3 |
| 2019 | Characterizing industry-academia collaborations in software engineering: evidence from 101 projectsabstractResearch collaboration between industry and academia supports improvement and innovation in industry and helps ensure the industrial relevance of academic research. However, many researchers and practitioners in the community believe that the level of joint industry-academia collaboration (IAC) projects in Software Engineering (SE) research is relatively low, creating a barrier between research and practice. The goal of the empirical study reported in this paper is to explore and characterize the state of IAC with respect to industrial needs, developed solutions, impacts of the projects and also a set of challenges, patterns and anti-patterns identified by a recent Systematic Literature Review (SLR) study. To address the above goal, we conducted an opinion survey among researchers and practitioners with respect to their experience in IAC. Our dataset includes 101 data points from IAC projects conducted in 21 different countries. Our findings include: (1) the most popular topics of the IAC projects, in the dataset, are: software testing, quality, process, and project managements; (2) over 90% of IAC projects result in at least one publication; (3) almost 50% of IACs are initiated by industry, busting the myth that industry tends to avoid IACs; and (4) 61% of the IAC projects report having a positive impact on their industrial context, while 31% report no noticeable impacts or were “not sure”. To improve this situation, we present evidence-based recommendations to increase the success of IAC projects, such as the importance of testing pilot solutions before using them in industry. This study aims to contribute to the body of evidence in the area of IAC, and benefit researchers and practitioners. Using the data and evidence presented in this paper, they can conduct more successful IAC projects in SE by being aware of the challenges and how to overcome them, by applying best practices (patterns), and by preventing anti-patterns. Vahid Garousi, Dietmar Pfahl, João M. Fernandes 0001, Michael Felderer, Mika Mäntylä, David C. Shepherd, Andrea Arcuri, Ahmet Coskunçay, Bedir Tekinerdogan |
Empir. Softw. Eng. | 5 |
| 2019 | Guidelines for including grey literature and conducting multivocal literature reviews in software engineering
Vahid Garousi, Michael Felderer, Mika Mäntylä |
Inf. Softw. Technol. | 3 |
| 2018 | On the use of emoticons in open source software developmentabstractBackground: Using sentiment analysis to study software developers' behavior comes with challenges such as the presence of a large amount of technical discussion unlikely to express any positive or negative sentiment. However, emoticons provide information about developer sentiments that can easily be extracted from software repositories. Aim: We investigate how software developers use emoticons differently in issue trackers in order to better understand the differences between developers and determine to which extent emoticons can be used as in place of sentiment analysis. Method: We extract emoticons from 1.3M comments from Apache's issue tracker and 4.5M from Mozilla's issue tracker using regular expressions built from a list of emoticons used by SentiStrength and Wikipedia. We check for statistical differences using Mann-Whitney U tests and determine the effect size with Cliff's δ. Results: Overall Mozilla developers rely more on emoticons than Apache developers. While the overall rate of comments with emoticons is of 1% and 3% for Apache and Mozilla, some individual developers can have a rate up to 21%. Looking specifically at Mozilla developers, we find that western developers use significantly more emoticons (with medium size effect) than eastern developers. While the majority of emoticons are used to express joy, we find that Mozilla developers use emoticons more frequently to express sadness and surprise than Apache developers. Finally, we find that Apache developers use overall more emoticons during weekends than during weekdays, with the share of sad and surprised emoticons increasing during weekends. Conclusions: While emoticons are primarily used to express joy, the more occasional use of sad and surprised emoticons can potentially be utilized to detect frustration in place of sentiment analysis among developers using emoticons frequently enough. Maëlick Claes, Mika Mäntylä, Umar Farooq 0006 |
ESEM | 2 |
| 2018 | Using experience sampling to link software repositories with emotions and work well-beingabstractBackground: The experience sampling method studies everyday experiences of humans in natural environments. In psychology it has been used to study the relationships between work well-being and productivity. To our best knowledge, daily experience sampling has not been previously used in software engineering. Aims: Our aim is to identify links between software developers self-reported affective states and work well-being and measures obtained from software repositories. Method: We perform an experience sampling study in a software company for a period of eight months, we use logistic regression to link the well-being measures with development activities, i.e. number of commits and chat messages. Results: We find several significant relationships between questionnaire variables and software repository variables. To our surprise relationship between hurry and number of commits is negative, meaning more perceived hurry is linked with a smaller number of commits. We also find a negative relationship between social interaction and hindered work well-being. Conclusions: The negative link between commits and hurry is counter-intuitive and goes against previous lab-experiments in software engineering that show increased efficiency under time pressure. Overall, our is an initial step in using experience sampling in software engineering and validating theories on work well-being from other fields in the domain of software engineering. Miikka Kuutila, Mika Mäntylä, Maëlick Claes, Marko Elovainio, Bram Adams |
ESEM | 2 |
| 2018 | Measuring LDA topic stability from clusters of replicated runsabstractBackground: Unstructured and textual data is increasing rapidly and Latent Dirichlet Allocation (LDA) topic modeling is a popular data analysis methods for it. Past work suggests that instability of LDA topics may lead to systematic errors. Aim: We propose a method that relies on replicated LDA runs, clustering, and providing a stability metric for the topics. Method: We generate k LDA topics and replicate this process n times resulting in n*k topics. Then we use K-medioids to cluster the n*k topics to k clusters. The k clusters now represent the original LDA topics and we present them like normal LDA topics showing the ten most probable words. For the clusters, we try multiple stability metrics, out of which we recommend Rank-Biased Overlap, showing the stability of the topics inside the clusters. Results: We provide an initial validation where our method is used for 270,000 Mozilla Firefox commit messages with k=20 and n=20. We show how our topic stability metrics are related to the contents of the topics. Conclusions: Advances in text mining enable us to analyze large masses of text in software engineering but non-deterministic algorithms, such as LDA, may lead to unreplicable conclusions. Our approach makes LDA stability transparent and is also complementary rather than alternative to many prior works that focus on LDA parameter tuning. Mika Mäntylä, Maëlick Claes, Umar Farooq 0006 |
ESEM | 1 |
| 2018 | Do programmers work at night or during the weekend?abstractAbnormal working hours can reduce work health, general well-being, and productivity, independent from a profession. To inform future approaches for automatic stress and overload detection, this paper establishes empirically collected measures of the work patterns of software engineers. To this aim, we perform the first large-scale study of software engineers' working hours by investigating the time stamps of commit activities of 86 large open source software projects, both containing hired and volunteer developers. We find that two thirds of software engineers mainly follow typical office hours, empirically established to be from 10h to 18h, and do not usually work during nights and weekends. Large variations between projects and individuals exist. Surprisingly, we found no support that project maturation would decrease abnormal working hours. In the Firefox case study, we found that hired developers work more during office hours while seniority, either in terms of number of commits or job status, did not impact working hours. We conclude that the use of working hours or timestamps of work products for stress detection requires establishing baselines at the level of individuals. Maëlick Claes, Mika Mäntylä, Miikka Kuutila, Bram Adams |
ICSE | 2 |
| 2018 | Towards automatically identifying paid open source developersabstractOpen source development contains contributions from both hired and volunteer software developers. Identification of this status is important when we consider the transferability of research results to the closed source software industry, as they include no volunteer developers. While many studies have taken the employment status of developers into account, this information is often gathered manually due to the lack of accurate automatic methods. In this paper, we present an initial step towards predicting paid and unpaid open source development using machine learning and compare our results with automatic techniques used in prior work. By relying on code source repository meta-data from Mozilla, and manually collected employment status, we built a dataset of the most active developers, both volunteer and hired by Mozilla. We define a set of metrics based on developers' usual commit time pattern and use different classification methods (logistic regression, classification tree, and random forest). The results show that our proposed method identify paid and unpaid commits with an AUC of 0.75 using random forest, which is higher than the AUC of 0.64 obtained with the best of the previously used automatic methods. Maëlick Claes, Mika Mäntylä, Miikka Kuutila, Umar Farooq 0006 |
MSR | 2 |
| 2018 | Natural language or not (NLON): a package for software engineering text analysis pipelineabstractThe use of natural language processing (NLP) is gaining popularity in software engineering. In order to correctly perform NLP, we must pre-process the textual information to separate natural language from other information, such as log messages, that are often part of the communication in software engineering. We present a simple approach for classifying whether some textual input is natural language or not. Although our NLoN package relies on only 11 language features and character tri-grams, we are able to achieve an area under the ROC curve performances between 0.976-0.987 on three different data sources, with Lasso regression from Glmnet as our learner and two human raters for providing ground truth. Cross-source prediction performance is lower and has more fluctuation with top ROC performances from 0.913 to 0.980. Compared with prior work, our approach offers similar performance but is considerably more lightweight, making it easier to apply in software engineering text mining pipelines. Our source code and data are provided as an R-package for further improvements. Mika Mäntylä, Fabio Calefato, Maëlick Claes |
MSR | 1 |
| 2018 | Test Case Prioritization Using Test Similarities
Alireza Haghighatkhah, Mika Mäntylä, Markku Oivo, Pasi Kuvaja |
PROFES | 2 |
| 2018 | A benchmark study on the effectiveness of search-based data selection and feature selection for cross project defect prediction
Seyedrebvar Hosseini, Burak Turhan, Mika Mäntylä |
Inf. Softw. Technol. | 3 |
| 2018 | Test prioritization in continuous integration environments
Alireza Haghighatkhah, Mika Mäntylä, Markku Oivo, Pasi Kuvaja |
J. Syst. Softw. | 2 |
| 2017 | Industry-academia collaborations in software engineering: An empirical analysis of challenges, patterns and anti-patterns in research projectsabstractResearch collaboration between industry and academia supports improvement and innovation in industry and helps to ensure industrial relevance in academic research. However, many researchers and practitioners believe that the level of joint industry-academia collaboration (IAC) in software engineering (SE) research is still relatively low, compared to the amount of activity in each of the two communities. The goal of the empirical study reported in this paper is to exploratory characterize the state of IAC with respect to a set of challenges, patterns and anti-patterns identified by a recent Systematic Literature Review study. To address the above goal, we gathered the opinions of researchers and practitioners w.r.t. their experiences in IAC projects. Our dataset includes 47 opinion data points related to a large set of projects conducted in 10 different countries. We aim to contribute to the body of evidence in the area of IAC, for the benefit of researchers and practitioners in conducting future successful IAC projects in SE. As an output, the study presents a set of empirical findings and evidence-based recommendations to increase the success of IAC projects. Vahid Garousi, Michael Felderer, João M. Fernandes 0001, Dietmar Pfahl, Mika Mäntylä |
EASE | 5 |
| 2017 | Choosing the Right Test Automation Tool: a Grey Literature Review of Practitioner SourcesabstractBackground: Choosing the right software test automation tool is not trivial, and recent industrial surveys indicate lack of right tools as the main obstacle to test automation. Aim: In this paper, we study how practitioners tackle the problem of choosing the right test automation tool. Method: We synthesize the "voice" of the practitioners with a grey literature review originating from 53 different companies. The industry experts behind the sources had roles such as "Software Test Automation Architect", and "Principal Software Engineer". Results: Common consensus about the important criteria exists but those are not applied systematically. We summarize the scattered steps from individual sources by presenting a comprehensive process for tool evaluation with 12 steps and a total of 14 different criteria for choosing the right tool. Conclusions: The practitioners tend to have general interest in and be influenced by related grey literature as about 78% of our sources had at least 20 backlinks (a reference comparable to a citation) while the variation was between 3 and 759 backlinks. There is a plethora of different software testing tools available, yet the practitioners seem to prefer and adopt the widely known and used tools. The study helps to identify the potential pitfalls of existing processes and opportunities for comprehensive tool evaluation. Päivi Raulamo-Jurvanen, Mika Mäntylä, Vahid Garousi |
EASE | 2 |
| 2017 | Abnormal working hours: effect of rapid releases and implications to work contentabstractDuring the past years, overload at work leading to psychological diseases, such as burnouts, have drawn more public attention. This paper is a preliminary step toward an analysis of the work patterns and possible indicators of overload and time pressure on software developers with mining software repositories approach. We explore the working pattern of developers in the context of Mozilla Firefox, a large and long-lived open source project. To that end we investigate the impact of the move from traditional to rapid release cycle on work pattern. Moreover we compare Mozilla Firefox work pattern with another Mozilla product, Firefox OS, which has a different release cycle than Firefox. We find that both projects exhibit healthy working patterns, i.e. lower activity during the weekends and outside of office hours. Firefox experiences proportionally more activity on weekends than Firefox OS (Cohen's d = 0.94). We find that switching to rapid releases has reduced weekend work (Cohen's d = 1.43) and working during the night (Cohen's d = 0.45). This result holds even when we limit the analyzes on the hired resources, i.e. considering only individuals with Mozilla foundation email address, although, the effect sizes are smaller for weekends (Cohen's d = 0.64) and nights (Cohen's d = 0.23). Moreover, we use dissimilarity word clouds and find that work during the weekend is more technical while work during the week expresses more positive sentiment with words like "good" and "nice". Our results suggest that moving to rapid releases have positive impact on the work health and work-life-balance of software engineers. However, caution is needed as our results are based on a limited set of quantitative data from a single organization. Maëlick Claes, Mika Mäntylä, Miikka Kuutila, Bram Adams |
MSR | 2 |
| 2017 | Bootstrapping a lexicon for emotional arousal in software engineeringabstractEmotional arousal increases activation and performance but may also lead to burnout in software development. We present the first version of a Software Engineering Arousal lexicon (SEA) that is specifically designed to address the problem of emotional arousal in the software developer ecosystem. SEA is built using a bootstrapping approach that combines word embedding model trained on issue-tracking data and manual scoring of items in the lexicon. We show that our lexicon is able to differentiate between issue priorities, which are a source of emotional activation and then act as a proxy for arousal. The best performance is obtained by combining SEA (428 words) with a previously created general purpose lexicon by Warriner et al. (13,915 words) and it achieves Cohen's d effect sizes up to 0.5. Mika Mäntylä, Nicole Novielli, Filippo Lanubile, Maëlick Claes, Miikka Kuutila |
MSR | 1 |
| 2017 | Guest editorial for special section on success and failure in software engineering
Mika Mäntylä, Magne Jørgensen, Paul Ralph, Hakan Erdogmus |
Empir. Softw. Eng. | 1 |
| 2017 | Prioritizing manual test cases in rapid release environmentsabstractSummary Test case prioritization is an important testing activity, in practice, specially for large scale systems. The goal is to rank the existing test cases in a way that they detect faults as soon as possible, so that any partial execution of the test suite detects the maximum number of defects for the given budget. Test prioritization becomes even more important when the test execution is time consuming, for example, manual system tests versus automated unit tests. Most existing test case prioritization techniques are based on code coverage, which requires access to source code. However, manual testing is mainly performed in a black‐box manner (manual testers do not have access to the source code). Therefore, in this paper, the existing test case prioritization techniques (e.g. diversity‐based and history‐based techniques) are examined and modified to be applicable on manual black‐box system testing. An empirical study on four older releases of desktop Firefox showed that none of the techniques were strongly dominating the others in all releases. However, when nine more recent releases of desktop Firefox, where the development has been moved from a traditional to a more agile and rapid release environment, were studied, a very significant difference between the history‐based approach and its alternatives was observed. The higher effectiveness of the history‐based approach compared with alternatives also held on 28 additional rapid releases of other Firefox projects – mobile Firefox and tablet Firefox. The conclusion of the paper is that test cases in rapid release environments can be very effectively prioritized for execution, based on their historical failure knowledge. In particular, it is the recency of historical knowledge that explains its effectiveness in rapid release environments rather than other changes in the process. Copyright © 2016 John Wiley & Sons, Ltd. Hadi Hemmati, Zhihan Fang, Mika Mäntylä, Bram Adams |
Softw. Test. Verification Reliab. | 3 |
| 2016 | The need for multivocal literature reviews in software engineering: complementing systematic literature reviews with grey literatureabstractSystematic Literature Reviews (SLR) may not provide insight into the "state of the practice" in SE, as they do not typically include the "grey" (non-published) literature. A Multivocal Literature Review (MLR) is a form of a SLR which includes grey literature in addition to the published (formal) literature. Only a few MLRs have been published in SE so far. We aim at raising the awareness for MLRs in SE by addressing two research questions (RQs): (1) What types of knowledge are missed when a SLR does not include the multivocal literature in a SE field? and (2) What do we, as a community, gain when we include the multivocal literature and conduct MLRs? To answer these RQs, we sample a few example SLRs and MLRs and identify the missing and the gained knowledge due to excluding or including the grey literature. We find that (1) grey literature can give substantial benefits in certain areas of SE, and that (2) the inclusion of grey literature brings forward certain challenges as evidence in them is often experience and opinion based. Given these conflicting viewpoints, the authors are planning to prepare systematic guidelines for performing MLRs in SE. Vahid Garousi, Michael Felderer, Mika Mäntylä |
EASE | 3 |
| 2016 | Mining valence, arousal, and dominance: possibilities for detecting burnout and productivity?abstractSimilar to other industries, the software engineering domain is plagued by psychological diseases such as burnout, which lead developers to lose interest, exhibit lower activity and/or feel powerless. Prevention is essential for such diseases, which in turn requires early identification of symptoms. The emotional dimensions of Valence, Arousal and Dominance (VAD) are able to derive a person's interest (attraction), level of activation and perceived level of control for a particular situation from textual communication, such as emails. As an initial step towards identifying symptoms of productivity loss in software engineering, this paper explores the VAD metrics and their properties on 700,000 Jira issue reports containing over 2,000,000 comments, since issue reports keep track of a developer's progress on addressing bugs or new features. Using a general-purpose lexicon of 14,000 English words with known VAD scores, our results show that issue reports of different type (e.g., Feature Request vs. Bug) have a fair variation of Valence, while increase in issue priority (e.g., from Minor to Critical) typically increases Arousal. Furthermore, we show that as an issue's resolution time increases, so does the arousal of the individual the issue is assigned to. Finally, the resolution of an issue increases valence, especially for the issue Reporter and for quickly addressed issues. The existence of such relations between VAD and issue report activities shows promise that text mining in the future could offer an alternative way for work health assessment surveys. Mika Mäntylä, Bram Adams, Giuseppe Destefanis, Daniel Graziotin, Marco Ortu |
MSR | 1 |
| 2016 | Gamification of Software Testing - An MLR
Mika Mäntylä, Kari Smolander |
PROFES | 1 |
| 2016 | Using Surveys and Web-Scraping to Select Tools for Software Testing Consultancy
Päivi Raulamo-Jurvanen, Kari Kakkonen, Mika Mäntylä |
PROFES | 3 |
| 2016 | Comparing and experimenting machine learning techniques for code smell detection
Francesca Arcelli Fontana, Mika Mäntylä, Marco Zanoni, Alessandro Marino |
Empir. Softw. Eng. | 2 |
| 2016 | When and what to automate in software testing? A multi-vocal literature review
Vahid Garousi, Mika Mäntylä |
Inf. Softw. Technol. | 2 |
| 2016 | A systematic literature review of literature reviews in software testing
Vahid Garousi, Mika Mäntylä |
Inf. Softw. Technol. | 2 |
| 2015 | Citation and Topic Analysis of the ESEM PapersabstractContext: The pool of papers published in ESEM. Objective: To utilize citation analysis and automated topic analysis to characterize the SE research literature over the years focusing on those papers published in ESEM. Method: We collected data from Scopus database consisting of 513 ESEM papers. For thematic analysis, we used topic modeling to automatically generate the most probable topic distributions given the data. Results: Nearly 42% of the papers have not been cited at all but the effect seems to wear off as time passes. Using text mining of article titles and abstracts, we found that currently the most popular research topics in the ESEM community are: systematic reviews, testing, defects, cost estimation, and team work. Conclusions: While this study analyzes the paper pool of the ESEM symposium, the approach can easily be applied to any other sub-set of SE papers to conduct large scale studies. Due to large volumes of research in SE, we suggest using the automated analysis of bibliometrics as we have done in this paper. Päivi Raulamo-Jurvanen, Mika Mäntylä, Vahid Garousi |
ESEM | 2 |
| 2015 | Prioritizing Manual Test Cases in Traditional and Rapid Release EnvironmentsabstractTest case prioritization is one of the most practically useful activities in testing, specially for large scale systems. The goal is ranking the existing test cases in a way that they detect faults as soon as possible, so that any partial execution of the test suite detects maximum number of defects for the given budget. Test prioritization becomes even more important when the test execution is time consuming, e.g., manual system tests vs. automated unit tests. Most existing test case prioritization techniques are based on code coverage, which requires access to source code. However, manual testing is mainly done in a black- box manner (manual testers do not have access to the source code). Therefore, in this paper, we first examine the existing test case prioritization techniques and modify them to be applicable on manual black-box system testing. We specifically study a coverage- based, a diversity-based, and a risk driven approach for test case prioritization. Our empirical study on four older releases of Mozilla Firefox shows that none of the techniques are strongly dominating the others in all releases. However, when we study nine more recent releases of Firefox, where the development has been moved from a traditional to a more agile and rapid release environment, we see a very signifiant difference (on average 65% effectiveness improvement) between the risk-driven approach and its alternatives. Our conclusion, based on one case study of 13 releases of an industrial system, is that test suites in rapid release environments, potentially, can be very effectively prioritized for execution, based on their historical riskiness; whereas the same conclusions do not hold in the traditional software development environments. Hadi Hemmati, Zhihan Fang, Mika Mäntylä |
ICST | 3 |
| 2015 | On rapid releases and software testing: a case study and a semi-systematic literature review
Mika Mäntylä, Bram Adams, Foutse Khomh, Emelie Engström, Kai Petersen |
Empir. Softw. Eng. | 1 |
| 2015 | Using metrics in Agile and Lean Software Development - A systematic literature review of industrial studies
Eetu Kupiainen, Mika Mäntylä, Juha Itkonen |
Inf. Softw. Technol. | 2 |
| 2015 | Diagrams or structural lists in software project retrospectives - An experimental comparisonabstractRoot cause analysis (RCA) is a recommended practice in retrospectives and cause–effect diagram (CED) is a commonly recommended technique for RCA. Our objective is to evaluate whether CED improves the outcome and perceived utility of RCA. We conducted a controlled experiment with 11 student software project teams by using a single factor paired design resulting in a total of 22 experimental units. Two visualization techniques of underlying causes were compared: CED and a structural list of causes. We used the output of RCA, questionnaires, and group interviews to compare the two techniques. In our results, CED increased the total number of detected causes. CED also increased the links between causes, thus, suggesting more structured analysis of problems. Furthermore, the participants perceived that CED improved organizing and outlining the detected causes. The implication of our results is that using CED in the RCA of retrospectives is recommended, yet, not mandatory as the groups also performed well with the structural list. In addition to increased number of detected causes, CED is visually more attractive and preferred by retrospective participants, even though it is somewhat harder to read and requires specific software tools. Timo O. A. Lehtinen, Mika Mäntylä, Juha Itkonen, Jari Vanhanen |
J. Syst. Softw. | 2 |
| 2014 | A replicated study on duplicate detection: using apache lucene to search among Android defectsabstractContext: Duplicate detection is a fundamental part of issue management. Systems able to predict whether a new defect report will be closed as a duplicate, may decrease costs by limiting rework and collecting related pieces of information. Goal: Our work explores using Apache Lucene for large-scale duplicate detection based on textual content. Also, we evaluate the previous claim that results are improved if the title is weighted as more important than the description. Method: We conduct a conceptual replication of a well-cited study conducted at Sony Ericsson, using Lucene for searching in the public Android defect repository. In line with the original study, we explore how varying the weighting of the title and the description affects the accuracy. Results: We show that Lucene obtains the best results when the defect report title is weighted three times higher than the description, a bigger difference than has been previously acknowledged. Conclusions: Our work shows the potential of using Lucene as a scalable solution for duplicate detection. Markus Borg, Per Runeson, Jens Johansson, Mika Mäntylä |
ESEM | 4 |
| 2014 | How is exploratory testing used? A state-of-the-practice surveyabstractContext: Exploratory Testing has experienced a rise in popularity in the industry with the emergence of agile development practices, yet it remains unclear, in which domains and how it is used in practice. Dietmar Pfahl, Huishi Yin, Mika Mäntylä, Jürgen Münch |
ESEM | 3 |
| 2014 | Time pressure: a controlled experiment of test case development and requirements reviewabstractTime pressure is prevalent in the software industry in which shorter and shorter deadlines and high customer demands lead to increasingly tight deadlines. However, the effects of time pressure have received little attention in software engineering research. We performed a controlled experiment on time pressure with 97 observations from 54 subjects. Using a two-by-two crossover design, our subjects performed requirements review and test case development tasks. We found statistically significant evidence that time pressure increases efficiency in test case development (high effect size Cohen’s d=1.279) and in requirements review (medium effect size Cohen’s d=0.650). However, we found no statistically significant evidence that time pressure would decrease effectiveness or cause adverse effects on motivation, frustration or perceived performance. We also investigated the role of knowledge but found no evidence of the mediating role of knowledge in time pressure as suggested by prior work, possibly due to our subjects. We conclude that applying moderate time pressure for limited periods could be used to increase efficiency in software engineering tasks that are well structured and straight forward. Mika Mäntylä, Kai Petersen, Timo O. A. Lehtinen, Casper Lassenius |
ICSE | 1 |
| 2014 | Supporting Regression Test Scoping with Visual AnalyticsabstractBackground: Test managers have to repeatedly select test cases for test activities during evolution of large software systems. Researchers have widely studied automated test scoping, but have not fully investigated decision support with human interaction. We previously proposed the introduction of visual analytics for this purpose. Aim: In this empirical study we investigate how to design such decision support. Method: We explored the use of visual analytics using heat maps of historical test data for test scoping support by letting test managers evaluate prototype visualizations in three focus groups with in total nine industrial test experts. Results: All test managers in the study found the visual analytics useful for supporting test planning. However, our results show that different tasks and contexts require different types of visualizations. Conclusion: Important properties for test planning support are: ability to overview testing from different perspectives, ability to filter and zoom to compare subsets of the testing with respect to various attributes and the ability to manipulate the subset under analysis by selecting and deselecting test cases. Our results may be used to support the introduction of visual test analytics in practice. Emelie Engström, Mika Mäntylä, Per Runeson, Markus Borg |
ICST | 2 |
| 2014 | Are test cases needed? Replicated comparison between exploratory and test-case-based software testing
Juha Itkonen, Mika Mäntylä |
Empir. Softw. Eng. | 2 |
| 2014 | Perceived causes of software project failures - An analysis of their relationships
Timo O. A. Lehtinen, Mika Mäntylä, Jari Vanhanen, Juha Itkonen, Casper Lassenius |
Inf. Softw. Technol. | 2 |
| 2014 | A tool supporting root cause analysis for synchronous retrospectives in distributed software teams
Timo O. A. Lehtinen, Risto Virtanen, Juha O. Viljanen, Mika Mäntylä, Casper Lassenius |
Inf. Softw. Technol. | 4 |
| 2014 | How are software defects found? The role of implicit defect detection, individual responsibility, documents, and knowledge
Mika Mäntylä, Juha Itkonen |
Inf. Softw. Technol. | 1 |
| 2013 | Defect Bash - Literature ReviewabstractDefect bash is a co-located testing session performed by a group of people. We performed a systematic review of the academic and grey literature, i.e. informally published writings, of the defect bash. Altogether, we found 44 items (17 academic and 27 grey literature sources) that were identified useful for the review. Based on the review the definition of defect bash is presented, benefits and limitations of using defect bash are given. Finally, the process of doing defect bash is outlined. This review provides initial understanding on how defect bash could be useful in achieving the software quality and lays foundation for further academic studies of this topic. 1 Xuejiao Zhou, Mika Mäntylä |
ENASE | 2 |
| 2013 | Code Smell Detection: Towards a Machine Learning-Based ApproachabstractSeveral code smells detection tools have been developed providing different results, because smells can be subjectively interpreted and hence detected in different ways. Usually the detection techniques are based on the computation of different kinds of metrics, and other aspects related to the domain of the system under analysis, its size and other design features are not taken into account. In this paper we propose an approach we are studying based on machine learning techniques. We outline some common problems faced for smells detection and we describe the different steps of our approach and the algorithms we use for the classification. Francesca Arcelli Fontana, Marco Zanoni, Alessandro Marino, Mika Mäntylä |
ICSM | 4 |
| 2013 | On Rapid Releases and Software TestingabstractLarge open and closed source organizations like Google, Facebook and Mozilla are migrating their products towards rapid releases. While this allows faster time-to-market and user feedback, it also implies less time for testing and bug fixing. Since initial research results indeed show that rapid releases fix proportionally less reported bugs than traditional releases, this paper investigates the changes in software testing effort after moving to rapid releases. We analyze the results of 312,502 execution runs of the 1,547 mostly manual system level test cases of Mozilla Fire fox from 2006 to 2012 (5 major traditional and 9 major rapid releases), and triangulated our findings with a Mozilla QA engineer. In rapid releases, testing has a narrower scope that enables deeper investigation of the features and regressions with the highest risk, while traditional releases run the whole test suite. Furthermore, rapid releases make it more difficult to build a large testing community, forcing Mozilla to increase contractor resources in order to sustain testing for rapid releases. Mika Mäntylä, Foutse Khomh, Bram Adams, Emelie Engström, Kai Petersen |
ICSM | 1 |
| 2013 | A Systematic Mapping Study of Empirical studies on the Use of Pair Programming in IndustryabstractPrevious systematic literature reviews on pair programming (PP) lack in their coverage of industrial PP data as well as certain factors of PP such as infrastructure. Therefore, we conducted a systematic mapping study on empirical, industrial PP research. Based on 154 research papers, we built a new PP framework containing 18 factors. We analyzed the previous research on each factor through several research properties. The most thoroughly studied factors in industry are communication, knowledge of work, productivity and quality. Many other factors largely lack comparative data, let alone data from reliable data collection methods such as measurement. Based on these gaps in research further studies would be most valuable for development process, targets of PP, developers’ characteristics, and feelings of work. We propose how they could be studied better. If the gaps had been commonly known, they could have been covered rather easily in the previous empirical studies. Our results help to focus further studies on the most relevant gaps in research and design them based on the previous studies. The results also help to identify the factors for which systematic reviews that synthesize the findings of the primary studies would already be feasible. Jari Vanhanen, Mika Mäntylä |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2013 | Analyzing an automotive testing process with evidence-based software engineering
Abhinaya Kasoju, Kai Petersen, Mika Mäntylä |
Inf. Softw. Technol. | 3 |
| 2013 | More testers - The effect of crowd size and time restriction in software testing
Mika Mäntylä, Juha Itkonen |
Inf. Softw. Technol. | 1 |
| 2013 | The Role of the Tester's Knowledge in Exploratory Software TestingabstractWe present a field study on how testers use knowledge while performing exploratory software testing (ET) in industrial settings. We video recorded 12 testing sessions in four industrial organizations, having our subjects think aloud while performing their usual functional testing work. Using applied grounded theory, we analyzed how the subjects performed tests and what type of knowledge they utilized. We discuss how testers recognize failures based on their personal knowledge without detailed test case descriptions. The knowledge is classified under the categories of domain knowledge, system knowledge, and general software engineering knowledge. We found that testers applied their knowledge either as a test oracle to determine whether a result was correct or not, or for test design, to guide them in selecting objects for test and designing tests. Interestingly, a large number of failures, windfall failures, were found outside the actual focus areas of testing as a result of exploratory investigation. We conclude that the way exploratory testers apply their knowledge for test design and failure recognition differs clearly from the test-case-based paradigm and is one of the explanatory factors of the effectiveness of the exploratory testing approach. Juha Itkonen, Mika Mäntylä, Casper Lassenius |
IEEE Trans. Software Eng. | 2 |
| 2012 | Testing highly complex system of systems: an industrial case studyabstractContext: Systems of systems (SoS) are highly complex and are integrated on multiple levels (unit, component, system, system of systems). Many of the characteristics of SoS (such as operational and managerial independence, integration of system into system of systems, SoS comprised of complex systems) make their development and testing challenging. Nauman Bin Ali, Kai Petersen, Mika Mäntylä |
ESEM | 3 |
| 2012 | How many individuals to use in a QA task with fixed total effort?abstractIncreasing the number of persons working on quality assurance (QA) tasks, e.g., reviews and testing, increases the number of defects detected -- but it also increases the total effort unless effort is controlled with fixed effort budgets. Our research investigates how QA tasks should be configured regarding two parameters, i.e., time and number of people. We define an optimization problem to answer this question. As a core element of the optimization problem we discuss and describe how defect detection probability should be modeled as a function of time. We apply the formulas used in the definition of the optimization problem to empirical defect data of an experiment previously conducted with university students. The results show that the optimal choice of the number of persons depends on the actual defect detection probabilities of the individual defects over time, but also on the size of the effort budget. Future work will focus on generalizing the optimization problem to a larger set of parameters, including not only task time and number of persons but also experience and knowledge of the personnel involved, and methods and tools applied when performing a QA task. Mika Mäntylä, Kai Petersen, Dietmar Pfahl |
ESEM | 1 |
| 2012 | Who tested my software? Testing as an organizationally cross-cutting activityabstractThere is a recognized disconnect between testing research and industry practice, and more studies are needed on understanding how testing is conducted in real-world circumstances instead of demonstrating the superiority of specific methods. Recent literature indicates that testing is a cross-cutting activity that involves various organizational roles rather than the sole involvement of specialized testers. This research empirically investigates how testing involves employees in varying organizational roles in software product companies. We studied the organization and values of testing using an exploratory case study methodology through interviews, defect database analysis, workshops, analyses of documentation, and informal communications at three software product companies. We analyzed which employee groups test software in the case companies, and how many defects they find. Two companies organized testing as a team effort, and one company had a specialized testing group because of its different development model. We found evidence that testing was not an action conducted only by testing specialists. Testing by individuals with customer contact and domain expertise was an important validation method. We discovered that defects found by developers had the highest fix rates while those revealed by specialized testers had the lowest. The defect importance was susceptible to organizational competition of resources (i.e., overvaluing defects of reporter’s own products or projects). We conclude that it is important to understand the diversity of individuals participating in software testing and the relevance of validation from the end users’ viewpoint. Future research is required to evaluate testing approaches for diverse organizational roles. Finally, to improve defect information, we suggest increasing automation in defect data collection. Mika Mäntylä, Juha Itkonen, Joonas Iivonen |
Softw. Qual. J. | 1 |
| 2011 | Survey Reproduction of Defect Reporting in Industrial Software DevelopmentabstractContext: Defect reporting is an important part of software development in-vivo, but previous work from open source context suggests that defect reports often have insufficient information for defect fixing. Objective: Our goal was to reproduce and partially replicate one of those open source studies in industrial context to see how well the results could be generalized. Method: We surveyed developers from six industrial software development organizations about the defect report information, from three viewpoints: concerning quality, usefulness and automation possibilities of the information. Seventy-four developers out of 142 completed our survey. Results: Our reproduction confirms the results of the prior study in that "steps to reproduce" and "observed behaviour" are highly important defect information. Our results extend the results of the prior study as we found that "part of the application", "configuration of the application", and "operating data" are also highly important, but they were not surveyed in the prior study. Finally, we classified defect information as "critical problems", "solutions", "boosters", and "essentials" based on the survey answers. Conclusion: The quality of defect reports is a problem in the software industry as well as in the open source community. Thus, we suggest that a part of the defect reporting should be automated since many of the defect reporters lack technical knowledge or interest to produce high-quality defect reports. Eero I. Laukkanen, Mika Mäntylä |
ESEM | 2 |
| 2011 | What are Problem Causes of Software Projects? Data of Root Cause Analysis at Four Software CompaniesabstractRoot cause analysis (RCA) is a structured investigation of a problem to detect the causes that need to be prevented. We applied ARCA, an RCA method, to target problems of four medium-sized software companies and collected 648 causes of software engineering problems. Thereafter, we applied grounded theory to the causes to study their types and related process areas. We detected 14 types of causes in 6 process areas. Our results indicate that development work and software testing are the most common process areas, whereas lack of instructions and experiences, insufficient work practices, low quality task output, task difficulty, and challenging existing product are the most common types of the causes. As the types of causes are evenly distributed between the cases, we hypothesize that the distributions could be generalizable. Finally, we found that only 2.5% of the causes are related to software development tools that are widely investigated in software engineering research. Timo O. A. Lehtinen, Mika Mäntylä |
ESEM | 2 |
| 2011 | Development and evaluation of a lightweight root cause analysis method (ARCA method) - Field studies at four software companies
Timo O. A. Lehtinen, Mika Mäntylä, Jari Vanhanen |
Inf. Softw. Technol. | 2 |
| 2010 | Characteristics of high performing testers: a case studyabstractObjective: We studied what are the characteristics of high performing software testers in the industry. Method: We conducted an exploratory case study, collecting data through recorded interviews of one development manager and three testers in each of the three companies, analysis of the defect database, and informal communication within our research partnership with the companies. Results: We found that experience, reflection, motivation and personal characteristics were the top level themes. Experience related to the domain, e.g. processes of the customer, and on the other hand, specialized technical skills, e.g. performance testing, were seen more important than skills of test case design and test planning. Joonas Iivonen, Mika Mäntylä, Juha Itkonen |
ESEM | 2 |
| 2010 | Empirical software evolvability - code smells and human evaluationsabstractLow software evolvability may increase costs of software development for over 30%. In practice, human evaluations and discoveries of software evolvability dictate the actions taken to improve the software evolvability, but the human side has often been ignored in prior research. This dissertation synopsis proposes a new group of code smells called the solution approach, which is based on a study of 563 evolvability issues found in industrial and student code reviews. Solution approach issues require re-thinking of the existing implementation rather than just reorganizing the code through refactoring. This work also contributes to the body of knowledge about software quality assurance practices by confirming that 75% of defects found in code reviews affect software evolvability rather than functionality. We also found evidence indicating that context-specific demographics, i.e., role in organization and code ownership, affect evolvability evaluations, but general demographics, i.e., work experience and education, do not. Mika Mäntylä |
ICSM | 1 |
| 2009 | How do testers do it? An exploratory study on manual testing practicesabstractWe present the results of a qualitative observation study on the manual testing practices in four software development companies. Manual testing practices are seldom studied, and based on the literature we conjecture that they have a strong effect on the effectiveness of manual testing. We observed testing sessions of 11 software professionals performing system level functional testing. As a result we identified 22 manual testing practices that we classified into 9 test session strategies and 13 detailed test execution techniques. Many of the identified techniques were based on similar ideas as traditional test case design techniques. However, the subjects applied these techniques during manual testing without separate test design phase. The results indicate that software professionals use a wide set of strategies and techniques when performing manual testing. Testers seem to need and use techniques even if applying exploratory testing. Juha Itkonen, Mika Mäntylä, Casper Lassenius |
ESEM | 2 |
| 2009 | What Types of Defects Are Really Discovered in Code Reviews?abstractResearch on code reviews has often focused on defect counts instead of defect types, which offers an imperfect view of code review benefits. In this paper, we classified the defects of nine industrial (C/C++) and 23 student (Java) code reviews, detecting 388 and 371 defects, respectively. First, we discovered that 75 percent of defects found during the review do not affect the visible functionality of the software. Instead, these defects improved software evolvability by making it easier to understand and modify. Second, we created a defect classification consisting of functional and evolvability defects. The evolvability defect classification is based on the defect types found in this study, but, for the functional defects, we studied and compared existing functional defect classifications. The classification can be useful for assigning code review roles, creating checklists, assessing software evolvability, and building software engineering tools. We conclude that, in addition to functional defects, code reviews find many evolvability defects and, thus, offer additional benefits over execution-based quality assurance methods that cannot detect evolvability defects. We suggest that code reviews may be most valuable for software products with long life cycles as the value of discovering evolvability defects in them is greater than for short life cycle systems. Mika Mäntylä, Casper Lassenius |
IEEE Trans. Software Eng. | 1 |
| 2007 | Defect Detection Efficiency: Test Case Based vs. Exploratory TestingabstractThis paper presents a controlled experiment comparing the defect detection efficiency of exploratory testing (ET) and test case based testing (TCT). While traditional testing literature emphasizes test cases, ET stresses the individual tester's skills during test execution and does not rely upon predesigned test cases. In the experiment, 79 advanced software engineering students performed manual functional testing on an open-source application with actual and seeded defects. Each student participated in two 90-minute controlled sessions, using ET in one and TCT in the other. We found no significant differences in defect detection efficiency between TCT and ET. The distributions of detected defects did not differ significantly regarding technical type, detection difficulty, or severity. However, TCT produced significantly more false defect reports than ET. Surprisingly, our results show no benefit of using predesigned test cases in terms of defect detection efficiency, emphasizing the need for further studies of manual testing. Juha Itkonen, Mika Mäntylä, Casper Lassenius |
ESEM | 2 |
| 2007 | Issues and Tactics when Adopting Pair Programming: A Longitudinal Case StudyabstractWe present experiences from a two-year study of adopting pair programming (PP) in a Finnish software product company. When adopting PP, the company used five tactics: the creation of simple PP guidelines, the use of a PP champion, making the use of PP voluntary, creating a positive atmosphere for PP, and instituting a separate PP room. By the end of the study the feelings of PP considerably surpassed developers' preconceptions of PP, and even the feelings of solo programming. Issues identified in the infrastructure for PP were solved through the adoption of the PP room. In the end of the study, a majority of the developers thought that PP should be utilized more than the reached ca. 10% of development effort. Unresolved issues in resourcing PP probably hindered reaching the desired level for the use of PP. Jari Vanhanen, Casper Lassenius, Mika Mäntylä |
ICSEA | 3 |
| 2006 | Subjective evaluation of software evolvability using code smells: An empirical study
Mika Mäntylä, Casper Lassenius |
Empir. Softw. Eng. | 1 |
| 2004 | Developing New Approaches for Software Design Quality Improvement Based on Subjective EvaluationsabstractThis research abstract presents two approaches for utilizing the developers' subjective design quality evaluations during the software lifecycle. In process-based approach developers study and improve their system's structure at fixed intervals. Tool-based approach uses subjective evaluations as input to tool analysis. These approaches or their combination are expected to improve software design and promote organizational learning about software design. Mika Mäntylä |
ICSE | 1 |
| 2004 | Bad Smells - Humans as Code CriticsabstractThis work presents the results of an initial empirical study on the subjective evaluation of bad code smells, which identify poor structures in software. Based on a case study in a Finnish software product company, we make two contributions. First, we studied the evaluator effect when subjectively evaluating the existence of smells in code modules. We found that the use of smells for code evaluation purposes is hard due to conflicting perceptions of different evaluators. Second, we applied source code metrics for identifying three smells and compared these results to the subjective evaluations. Surprisingly, the metrics and smell evaluations did not correlate. Mika Mäntylä, Jari Vanhanen, Casper Lassenius |
ICSM | 1 |
| 2003 | A Taxonomy and an Initial Empirical Study of Bad Smells in CodeabstractThis paper presents research in progress, as well as tentative findings related to the empirical study of so called bad code smells. We present a taxonomy that categorizes similar bad smells. We believe that taxonomy makes the smells more understandable and recognizes the relationships between smells. Additionally, we present our initial findings from an empirical study of the use of the smells for evaluating code quality in a small Finnish software product company. Our findings indicate that the taxonomy for the smells could help explain the identified correlations between the subjective evaluations of the existence of the smells. Mika Mäntylä, Jari Vanhanen, Casper Lassenius |
ICSM | 1 |