VLDB 2026 Research / reviewers in the wild / expert
Miikka Kuutila
dblp:185/5232
· DBLP profile ↗
14ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-3695-7280ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SESR-Eval: Dataset for Evaluating LLMs in the Title-Abstract Screening of Systematic ReviewsabstractBackground: The use of large language models (LLMs) in the title-abstract screening process of systematic reviews (SRs) has shown promising results, but suffers from limited performance evaluation. Aims: Create a benchmark dataset to evaluate the performance of LLMs in the title-abstract screening process of SRs. Provide evidence whether using LLMs in title-abstract screening in software engineering is advisable. Method: We start with 169 SR research artifacts and find 24 of those to be suitable for inclusion in the dataset. Using the dataset we benchmark title-abstract screening using 9 LLMs. Results: We present the SESR-Eval (Software Engineering Systematic Review Evaluation) dataset containing 34,528 labeled primary studies, sourced from 24 secondary studies published in software engineering (SE) journals. Most LLMs performed similarly and the differences in screening accuracy between secondary studies are greater than differences between LLMs. The cost of using an LLM is relatively low - less than 40 per secondary study even for the most expensive model. Conclusions: Our benchmark enables monitoring AI performance in the screening task of SRs in software engineering. At present, LLMs are not yet recommended for automating the title-abstract screening process, since accuracy varies widely across secondary studies, and no LLM managed a high recall with reasonable precision. In future, we plan to investigate factors that influence LLM screening performance between studies. Aleksi Huotala, Miikka Kuutila, Mika Mäntylä |
ESEM | 2 |
| 2025 | User Personas Improve Social Sustainability by Encouraging Software Developers to Deprioritize Antisocial FeaturesabstractBackground: Sustainable software development involves creating software in a manner that meets present goals without undermining our ability to meet future goals. In a software engineering context, sustainability has at least four dimensions: ecological, economic, social, and technical. No interventions for improving social sustainability in software engineering have been tested in rigorous lab-based experiments, and little evidence-based guidance is available. Objective: The purpose of this study is to evaluate the effectiveness of two interventions—stakeholder maps and persona models—for improving social sustainability through software feature prioritization. Method: We conducted a randomized controlled factorial experiment with 79 undergraduate computer science students. Participants were randomly assigned to one of four groups and asked to prioritize a backlog of prosocial, neutral, and antisocial user stories for a shopping mall's digital screen display and facial recognition software. Participants received either persona models, a stakeholder map, both, or neither. We compared the differences in prioritization levels assigned to prosocial and antisocial user stories using Cumulative Link Mixed Model regression. Results: Participants who received persona models gave significantly lower priorities to antisocial user stories but no significant difference was evident for prosocial user stories. The effects of the stake-holder map were not significant. The interaction effects were not significant. Conclusion: Providing aspiring software professionals with well-crafted persona models causes them to de-prioritize antisocial software features. The impact of persona modelling on sustainable software development therefore warrants further study with more experience professionals. Moreover, the novel methodological strategy of assessing social sustainability behavior through backlog prioritization appears feasible in lab-based settings. Bimpe Ayoola, Miikka Kuutila, Rina R. Wehbe, Paul Ralph |
ICSE | 2 |
| 2025 | Credtwi: Investigating Social Media Credibility with a Browser PluginabstractPeople now look for information online and on social media for everyday problems. Organizations and malevolent actors have taken the opportunity to spread misinformation/disinformation. It is increasingly important to understand the credibility of online information. We designed and implemented a research browser plugin, Credtwi. It injects credibility questionnaires directly into the user’s Twitter feed, enabling crowdsourced data collection. We carried out a week-long field study where participants assessed the credibility of tweets on various topics. We provide insights into information credibility in the Twitter ecosystem by analyzing the assessments and study questionnaires. The participants’ perception of Twitter as a credible information source decreased after using Credtwi. Our results suggest that the author’s verification status and bio are the most important factors for their perceived credibility. Finally, we discovered significant differences between the assessments of the different genders. Our results contribute to the research on online social media content credibility. Eetu Huusko, Nazanin Nakhaie Ahooie, Miikka Kuutila, Aleksi Huotala, Mika Mäntylä, Simo Hosio |
Int. J. Hum. Comput. Interact. | 4 |
| 2025 | Research artifacts in secondary studies: A systematic mapping in software engineeringabstractContext: Systematic reviews (SRs) summarize state-of-the-art evidence in science, including software engineering (SE). Objective: Our objective is to evaluate how SRs report research artifacts and to provide a comprehensive list of these artifacts. Method: We examined 537 secondary studies published between 2013 and 2023 to analyze the availability and reporting of research artifacts. Results: Our findings indicate that only 31.5% of the reviewed studies include research artifacts. Encouragingly, the situation is gradually improving, as our regression analysis shows a significant increase in the availability of research artifacts over time. However, in 2023, just 62.0% of secondary studies provide a research artifact while an even lower percentage, 30.4% use a permanent repository with a digital object identifier (DOI) for storage. Conclusion: To enhance transparency and reproducibility in SE research, we advocate for the mandatory publication of research artifacts in secondary studies. Aleksi Huotala, Miikka Kuutila, Mika Mäntylä |
Inf. Softw. Technol. | 2 |
| 2024 | The Promise and Challenges of Using LLMs to Accelerate the Screening Process of Systematic ReviewsabstractContext: Systematic review (SR) is a popular research method in software engineering (SE). However, conducting an SR takes an average of 67 weeks. Thus, automating any step of the SR process could reduce the effort associated with SRs. Objective: Our objective is to investigate the extent to which Large Language Models (LLMs) can accelerate title-abstract screening by (1) simplifying abstracts for human screeners, and (2) automating title-abstract screening entirely. Method: We performed an experiment where human screeners performed title-abstract screening for 20 papers with both original and simplified abstracts from a prior SR. The experiment with human screeners was reproduced by instructing GPT-3.5 and GPT-4 LLMs to perform the same screening tasks. We also studied whether different prompting techniques (Zero-shot (ZS), One-shot (OS), Few-shot (FS), and Few-shot with Chain-of-Thought (FS-CoT) prompting) improve the screening performance of LLMs. Lastly, we studied if redesigning the prompt used in the LLM reproduction of title-abstract screening leads to improved screening performance. Results: Text simplification did not increase the screeners’ screening performance, but reduced the time used in screening. Screeners’ scientific literacy skills and researcher status predict screening performance. Some LLM and prompt combinations perform as well as human screeners in the screening tasks. Our results indicate that a more recent LLM (GPT-4) is better than its predecessor LLM (GPT-3.5). Additionally, Few-shot and One-shot prompting outperforms Zero-shot prompting. Conclusion: Using LLMs for text simplification in the screening process does not significantly improve human performance. Using LLMs to automate title-abstract screening seems promising, but current LLMs are not significantly more accurate than human screeners. To recommend the use of LLMs in the screening process of SRs, more research is needed. We recommend future SR studies to publish replication packages with screening data to enable more conclusive experimenting with LLM screening. Aleksi Huotala, Miikka Kuutila, Paul Ralph, Mika Mäntylä |
EASE | 2 |
| 2024 | What Makes Programmers Laugh? Exploring the Submissions of the Subreddit r/ProgrammerHumorabstractBackground: Humor is a fundamental part of human communication, with prior work linking positive humor in the workplace to positive outcomes, such as improved performance and job satisfaction. Aims: This study aims to investigate programming-related humor in a large social media community. Methodology: We collected 139,718 submissions from Reddit subreddit r/ProgrammerHumor. Both textual and image-based (memes) submissions were considered. The image data was processed with OCR to extract text from images for NLP analysis. Multiple regression models were built to investigate what makes submissions humorous. Additionally, a random sample of 800 submissions was labeled by human annotators regarding their relation to theories of humor, suitability for the workplace, the need for programming knowledge to understand the submission, and whether images in image-based submissions added context to the submission. Results: Our results indicate that predicting the humor of software developers is difficult. Our best regression model was able to explain only 10% of the variance. However, statistically significant differences were observed between topics, submission times, and associated humor theories. Our analysis reveals that the highest submission scores are achieved by image-based submissions that are created during the winter months in the northern hemisphere, between 2-3pm UTC on weekends, which are distinctly related to superiority and incongruity theories of humor, and are about the topic of "Learning". Conclusions: Predicting humor with natural language processing methods is challenging. We discuss the benefits and inherent difficulties in assessing perceived humor of submissions, as well as possible avenues for future work. Additionally, our replication package should help future studies and can act as a joke repository for the software industry and education. Miikka Kuutila, Leevi Rantala, Simo Hosio, Mika Mäntylä |
ESEM | 1 |
| 2023 | It is an online platform and not the real world, I don't care much: Investigating Twitter Profile Credibility With an Online Machine Learning-Based ToolabstractSocial media is now an important source of everyday information. Given the plethora of scandals concerning the rapid spread of misinformation and disinformation on social media, the credibility of the content on these platforms is now a pivotal research area. Much of the existing work on social media credibility focuses on content credibility. In this study, however, we focus on the credibility of the profile as the virtual representation of the content author. We developed a real-time machine-learning-based online tool that assesses the credibility of profiles on Twitter, one of the most common and versatile social media platforms. To investigate user perceptions on credibility-related issues, we used our tool as a stimulus for people to reflect on their profile’s credibility and collected 100 responses. The combination of our quantitative and qualitative analysis reveals that the latest tweets and retweet behavior are two of the most critical factors for profile credibility. It is also observed that people demonstrate a limited interest in their profile credibility but agree that the author’s credibility is of paramount importance. With an open-source tool to assess user credibility on Twitter and a user study to establish its utility, we contribute a timely piece of research on the topic of online credibility. Ville Paananen, Sharadhi Alape Suryanarayana, Eetu Huusko, Miikka Kuutila, Mika Mäntylä, Simo Hosio |
CHIIR | 5 |
| 2021 | Individual differences limit predicting well-being and productivity using software repositories: a longitudinal industrial studyabstractReports of poor work well-being and fluctuating productivity in software engineering have been reported in both academic and popular sources. Understanding and predicting these issues through repository analysis might help manage software developers' well-being. Our objective is to link data from software repositories, that is commit activity, communication, expressed sentiments, and job events, with measures of well-being obtained with a daily experience sampling questionnaire. To achieve our objective, we studied a single software project team for eight months in the software industry. Additionally, we performed semi-structured interviews to explain our results. The acquired quantitative data are analyzed with generalized linear mixed-effects models with autocorrelation structure. We find that individual variance accounts for most of the $R^2$ values in models predicting developers' experienced well-being and productivity. In other words, using software repository variables to predict developers' well-being or productivity is challenging due to individual differences. Prediction models developed for each developer individually work better, with fixed effects $R^2$ value of up to 0.24. The semi-structured interviews give insights into the well-being of software developers and the benefits of chat interaction. Our study suggests that individualized prediction models are needed for well-being and productivity prediction in software development. Miikka Kuutila, Mika Mäntylä, Maëlick Claes, Marko Elovainio, Bram Adams |
Empir. Softw. Eng. | 1 |
| 2020 | Time pressure in software engineering: A systematic review
Miikka Kuutila, Mika Mäntylä, Umar Farooq 0006, Maëlick Claes |
Inf. Softw. Technol. | 1 |
| 2018 | Using experience sampling to link software repositories with emotions and work well-beingabstractBackground: The experience sampling method studies everyday experiences of humans in natural environments. In psychology it has been used to study the relationships between work well-being and productivity. To our best knowledge, daily experience sampling has not been previously used in software engineering. Aims: Our aim is to identify links between software developers self-reported affective states and work well-being and measures obtained from software repositories. Method: We perform an experience sampling study in a software company for a period of eight months, we use logistic regression to link the well-being measures with development activities, i.e. number of commits and chat messages. Results: We find several significant relationships between questionnaire variables and software repository variables. To our surprise relationship between hurry and number of commits is negative, meaning more perceived hurry is linked with a smaller number of commits. We also find a negative relationship between social interaction and hindered work well-being. Conclusions: The negative link between commits and hurry is counter-intuitive and goes against previous lab-experiments in software engineering that show increased efficiency under time pressure. Overall, our is an initial step in using experience sampling in software engineering and validating theories on work well-being from other fields in the domain of software engineering. Miikka Kuutila, Mika Mäntylä, Maëlick Claes, Marko Elovainio, Bram Adams |
ESEM | 1 |
| 2018 | Do programmers work at night or during the weekend?abstractAbnormal working hours can reduce work health, general well-being, and productivity, independent from a profession. To inform future approaches for automatic stress and overload detection, this paper establishes empirically collected measures of the work patterns of software engineers. To this aim, we perform the first large-scale study of software engineers' working hours by investigating the time stamps of commit activities of 86 large open source software projects, both containing hired and volunteer developers. We find that two thirds of software engineers mainly follow typical office hours, empirically established to be from 10h to 18h, and do not usually work during nights and weekends. Large variations between projects and individuals exist. Surprisingly, we found no support that project maturation would decrease abnormal working hours. In the Firefox case study, we found that hired developers work more during office hours while seniority, either in terms of number of commits or job status, did not impact working hours. We conclude that the use of working hours or timestamps of work products for stress detection requires establishing baselines at the level of individuals. Maëlick Claes, Mika Mäntylä, Miikka Kuutila, Bram Adams |
ICSE | 3 |
| 2018 | Towards automatically identifying paid open source developersabstractOpen source development contains contributions from both hired and volunteer software developers. Identification of this status is important when we consider the transferability of research results to the closed source software industry, as they include no volunteer developers. While many studies have taken the employment status of developers into account, this information is often gathered manually due to the lack of accurate automatic methods. In this paper, we present an initial step towards predicting paid and unpaid open source development using machine learning and compare our results with automatic techniques used in prior work. By relying on code source repository meta-data from Mozilla, and manually collected employment status, we built a dataset of the most active developers, both volunteer and hired by Mozilla. We define a set of metrics based on developers' usual commit time pattern and use different classification methods (logistic regression, classification tree, and random forest). The results show that our proposed method identify paid and unpaid commits with an AUC of 0.75 using random forest, which is higher than the AUC of 0.64 obtained with the best of the previously used automatic methods. Maëlick Claes, Mika Mäntylä, Miikka Kuutila, Umar Farooq 0006 |
MSR | 3 |
| 2017 | Abnormal working hours: effect of rapid releases and implications to work contentabstractDuring the past years, overload at work leading to psychological diseases, such as burnouts, have drawn more public attention. This paper is a preliminary step toward an analysis of the work patterns and possible indicators of overload and time pressure on software developers with mining software repositories approach. We explore the working pattern of developers in the context of Mozilla Firefox, a large and long-lived open source project. To that end we investigate the impact of the move from traditional to rapid release cycle on work pattern. Moreover we compare Mozilla Firefox work pattern with another Mozilla product, Firefox OS, which has a different release cycle than Firefox. We find that both projects exhibit healthy working patterns, i.e. lower activity during the weekends and outside of office hours. Firefox experiences proportionally more activity on weekends than Firefox OS (Cohen's d = 0.94). We find that switching to rapid releases has reduced weekend work (Cohen's d = 1.43) and working during the night (Cohen's d = 0.45). This result holds even when we limit the analyzes on the hired resources, i.e. considering only individuals with Mozilla foundation email address, although, the effect sizes are smaller for weekends (Cohen's d = 0.64) and nights (Cohen's d = 0.23). Moreover, we use dissimilarity word clouds and find that work during the weekend is more technical while work during the week expresses more positive sentiment with words like "good" and "nice". Our results suggest that moving to rapid releases have positive impact on the work health and work-life-balance of software engineers. However, caution is needed as our results are based on a limited set of quantitative data from a single organization. Maëlick Claes, Mika Mäntylä, Miikka Kuutila, Bram Adams |
MSR | 3 |
| 2017 | Bootstrapping a lexicon for emotional arousal in software engineeringabstractEmotional arousal increases activation and performance but may also lead to burnout in software development. We present the first version of a Software Engineering Arousal lexicon (SEA) that is specifically designed to address the problem of emotional arousal in the software developer ecosystem. SEA is built using a bootstrapping approach that combines word embedding model trained on issue-tracking data and manual scoring of items in the lexicon. We show that our lexicon is able to differentiate between issue priorities, which are a source of emotional activation and then act as a proxy for arousal. The best performance is obtained by combining SEA (428 words) with a previously created general purpose lexicon by Warriner et al. (13,915 words) and it achieves Cohen's d effect sizes up to 0.5. Mika Mäntylä, Nicole Novielli, Filippo Lanubile, Maëlick Claes, Miikka Kuutila |
MSR | 5 |