Aleksi Huotala

dblp:343/3184 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-5220-8730ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SESR-Eval: Dataset for Evaluating LLMs in the Title-Abstract Screening of Systematic Reviews
abstract
Background: The use of large language models (LLMs) in the title-abstract screening process of systematic reviews (SRs) has shown promising results, but suffers from limited performance evaluation. Aims: Create a benchmark dataset to evaluate the performance of LLMs in the title-abstract screening process of SRs. Provide evidence whether using LLMs in title-abstract screening in software engineering is advisable. Method: We start with 169 SR research artifacts and find 24 of those to be suitable for inclusion in the dataset. Using the dataset we benchmark title-abstract screening using 9 LLMs. Results: We present the SESR-Eval (Software Engineering Systematic Review Evaluation) dataset containing 34,528 labeled primary studies, sourced from 24 secondary studies published in software engineering (SE) journals. Most LLMs performed similarly and the differences in screening accuracy between secondary studies are greater than differences between LLMs. The cost of using an LLM is relatively low - less than 40 per secondary study even for the most expensive model. Conclusions: Our benchmark enables monitoring AI performance in the screening task of SRs in software engineering. At present, LLMs are not yet recommended for automating the title-abstract screening process, since accuracy varies widely across secondary studies, and no LLM managed a high recall with reasonable precision. In future, we plan to investigate factors that influence LLM screening performance between studies.
Aleksi Huotala, Miikka Kuutila, Mika Mäntylä
ESEM1
2025 Credtwi: Investigating Social Media Credibility with a Browser Plugin
abstract
People now look for information online and on social media for everyday problems. Organizations and malevolent actors have taken the opportunity to spread misinformation/disinformation. It is increasingly important to understand the credibility of online information. We designed and implemented a research browser plugin, Credtwi. It injects credibility questionnaires directly into the user’s Twitter feed, enabling crowdsourced data collection. We carried out a week-long field study where participants assessed the credibility of tweets on various topics. We provide insights into information credibility in the Twitter ecosystem by analyzing the assessments and study questionnaires. The participants’ perception of Twitter as a credible information source decreased after using Credtwi. Our results suggest that the author’s verification status and bio are the most important factors for their perceived credibility. Finally, we discovered significant differences between the assessments of the different genders. Our results contribute to the research on online social media content credibility.
Eetu Huusko, Nazanin Nakhaie Ahooie, Miikka Kuutila, Aleksi Huotala, Mika Mäntylä, Simo Hosio
Int. J. Hum. Comput. Interact.5
2025 Research artifacts in secondary studies: A systematic mapping in software engineering
abstract
Context: Systematic reviews (SRs) summarize state-of-the-art evidence in science, including software engineering (SE). Objective: Our objective is to evaluate how SRs report research artifacts and to provide a comprehensive list of these artifacts. Method: We examined 537 secondary studies published between 2013 and 2023 to analyze the availability and reporting of research artifacts. Results: Our findings indicate that only 31.5% of the reviewed studies include research artifacts. Encouragingly, the situation is gradually improving, as our regression analysis shows a significant increase in the availability of research artifacts over time. However, in 2023, just 62.0% of secondary studies provide a research artifact while an even lower percentage, 30.4% use a permanent repository with a digital object identifier (DOI) for storage. Conclusion: To enhance transparency and reproducibility in SE research, we advocate for the mandatory publication of research artifacts in secondary studies.
Aleksi Huotala, Miikka Kuutila, Mika Mäntylä
Inf. Softw. Technol.1
2024 The Promise and Challenges of Using LLMs to Accelerate the Screening Process of Systematic Reviews
abstract
Context: Systematic review (SR) is a popular research method in software engineering (SE). However, conducting an SR takes an average of 67 weeks. Thus, automating any step of the SR process could reduce the effort associated with SRs. Objective: Our objective is to investigate the extent to which Large Language Models (LLMs) can accelerate title-abstract screening by (1) simplifying abstracts for human screeners, and (2) automating title-abstract screening entirely. Method: We performed an experiment where human screeners performed title-abstract screening for 20 papers with both original and simplified abstracts from a prior SR. The experiment with human screeners was reproduced by instructing GPT-3.5 and GPT-4 LLMs to perform the same screening tasks. We also studied whether different prompting techniques (Zero-shot (ZS), One-shot (OS), Few-shot (FS), and Few-shot with Chain-of-Thought (FS-CoT) prompting) improve the screening performance of LLMs. Lastly, we studied if redesigning the prompt used in the LLM reproduction of title-abstract screening leads to improved screening performance. Results: Text simplification did not increase the screeners’ screening performance, but reduced the time used in screening. Screeners’ scientific literacy skills and researcher status predict screening performance. Some LLM and prompt combinations perform as well as human screeners in the screening tasks. Our results indicate that a more recent LLM (GPT-4) is better than its predecessor LLM (GPT-3.5). Additionally, Few-shot and One-shot prompting outperforms Zero-shot prompting. Conclusion: Using LLMs for text simplification in the screening process does not significantly improve human performance. Using LLMs to automate title-abstract screening seems promising, but current LLMs are not significantly more accurate than human screeners. To recommend the use of LLMs in the screening process of SRs, more research is needed. We recommend future SR studies to publish replication packages with screening data to enable more conclusive experimenting with LLM screening.
Aleksi Huotala, Miikka Kuutila, Paul Ralph, Mika Mäntylä
EASE1
2022 Benefits and Challenges of Isomorphism in Single-page Applications: Case Study and Review of Gray Literature
abstract
An isomorphic web application shares code between the server and the client by cleverly combining suitable parts of server-rendered applications and single-page applications. In this article, we study the benefits and challenges of isomorphism in single-page applications in terms of a gray literature review and a case study. The case study was conducted as a developer interview, where developers familiar with isomorphic web applications were interviewed. The results of both studies are then compared and the key findings are compared together. The results show that isomorphism in single-page applications brings benefits to both the developers and the end-users. Isomorphism in single-page applications is challenging to implement and has some downsides, but they mostly affect developers. Implementing isomorphism enables sharing code between the server and the client, but it increases the complexity of the application. Framework and library compatibility are issues that must be addressed by the developers.
Aleksi Huotala, Matti Luukkainen, Tommi Mikkonen
J. Web Eng.1