VLDB 2026 Research / reviewers in the wild / expert
Mika Saari
dblp:165/6843
· DBLP profile ↗
11ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0001-7677-2355ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Collaboration Between Databases and Large Language Models: Creating Skills by Free Online Training CoursesabstractAs they have been for decades, databases are a key component in various types of business applications. Nowadays, modern databases also include features that support the construction of applications based on artificial intelligence. The starting point of this paper is the questions of what are the typical use cases in which a database and generative artificial intelligence work together, what database features are used in these applications, and what other technologies the applications rely on. The results contribute to one of our practice-oriented research projects, in which we build an environment that supports AI experiments for participating companies. As input, we use open online learning materials provided by three database vendors. The databases we selected for review are all non-relational, representing three different genres, and are the most popular representatives of their genre. For each training program of the vendors, we will examine the structure and scope of the program, the courses on generative AI included in the program, and the use cases and technologies related to generative AI presented in them. Finally, we will prepare a summary of the use cases and technologies found. Timo Mäkinen, Hannu Jaakkola, Jari Soini, Mika Saari |
EJC | 4 |
| 2025 | Specializing LLMs with Custom Datasets: A Four-Round Open vs Closed ComparisonabstractAs general-purpose LLMs become widely accessible, tailoring their outputs to domain-specific tasks remains challenging. This paper investigates how custom, pattern-guided datasets and fine-tuning can specialize LLM behavior. We run a four-round open-vs-closed comparison: (1) evaluate commercially hosted models with an open-source dataset; (2) repeat with a synthetic, pattern-aligned variant; (3) replicate both rounds with open-source LLMs; and (4) demonstrate a real-world case by fine-tuning an open model to answer WordPress technical support queries using a finalized, pattern-constrained dataset. Across rounds, fine-tuned open-source models achieved performance comparable to closed-source baselines while offering greater controllability and deployment flexibility. We detail dataset design, the role of response patterns in producing consistent outputs, and practical trade-offs observed during fine-tuning and real-world use. The findings indicate that with well-crafted custom datasets, open-source LLMs can be reliably specialized for targeted applications without reliance on proprietary APIs. Mubashir Saeed, Mika Saari, Kari Systä, Pekka Abrahamsson |
EJC | 2 |
| 2025 | Autonomous Legacy Web Application Upgrades Using a Multi-Agent SystemabstractThe use of Large Language Models (LLMs) for autonomous code generation is gaining attention in emerging technologies. As LLM capabilities expand, they offer new possibilities such as code refactoring, security enhancements, and legacy application upgrades. Many outdated web applications pose security and reliability challenges, yet companies continue using them due to the complexity and cost of upgrades. To address this, we propose an LLM-based multi-agent system that autonomously upgrades legacy web applications to the latest versions. The system distributes tasks across multiple phases, updating all relevant files. To evaluate its effectiveness, we employed Zero-Shot Learning (ZSL) and One-Shot Learning (OSL) prompts, applying identical instructions in both cases. The evaluation involved updating view files and measuring the number and types of errors in the output. For complex tasks, we counted the successfully met requirements. The experiments compared the proposed system with standalone LLM execution, repeated multiple times to account for stochastic behavior. Results indicate that our system maintains context across tasks and agents, improving solution quality over the base model in some cases. This study provides a foundation for future model implementations in legacy code updates. Additionally, findings highlight LLMs' ability to update small outdated files with high precision, even with basic prompts. The source code is publicly available on GitHub: https://github.com/alasalm1/Multi-agent-pipeline. Valtteri Ala-Salmi, Zeeshan Rasheed 0001, Malik Abdul Sami, Zheying Zhang, Kai-Kristian Kemell, Jussi Rasku, Shahbaz Siddeeq, Mika Saari, Pekka Abrahamsson |
ENASE | 8 |
| 2025 | Engineering RAG Systems for Real-World Applications: Design, Development, and Evaluation
Md Toufique Hasan, Muhammad Waseem 0011, Kai-Kristian Kemell, Ayman Asad Khan, Mika Saari, Pekka Abrahamsson |
SEAA (2) | 5 |
| 2025 | LLM-Based Multi-agent System for Intelligent Refactoring of Haskell Code
Shahbaz Siddeeq, Muhammad Waseem 0011, Zeeshan Rasheed 0001, Md Mahade Hasan, Jussi Rasku, Mika Saari, Henri Terho, Kalle Mäkelä, Kai-Kristian Kemell, Pekka Abrahamsson |
PROFES | 6 |
| 2025 | Generative Artificial Intelligence for Software Engineering - A Research AgendaabstractABSTRACT Context Generative artificial intelligence (GenAI) tools have become increasingly prevalent in software development, offering assistance to various managerial and technical project activities. Notable examples of these tools include OpenAI's ChatGPT, GitHub Copilot, and Amazon CodeWhisperer. Objective Although many recent publications have explored and evaluated the application of GenAI, a comprehensive understanding of the current development, applications, limitations, and open challenges remains unclear to many. Particularly, we do not have an overall picture of the current state of GenAI technology in practical software engineering usage scenarios. Method We conducted a literature review and focus groups for a duration of five months to develop a research agenda on GenAI for software engineering. Results We identified 78 open research questions (RQs) in 11 areas of software engineering. Our results show that it is possible to explore the adoption of GenAI in partial automation and support decision‐making in all software development activities. While the current literature is skewed toward software implementation, quality assurance and software maintenance, other areas, such as requirements engineering, software design, and software engineering education, would need further research attention. Common considerations when implementing GenAI include industry‐level assessment, dependability and accuracy, data accessibility, transparency, and sustainability aspects associated with the technology. Conclusions GenAI is bringing significant changes to the field of software engineering. Nevertheless, the state of research on the topic still remains immature. We believe that this research agenda holds significance and practical value for informing both researchers and practitioners about current applications and guiding future research. Anh Nguyen-Duc 0001, Beatriz Cabrero-Daniel, Adam Przybylek, Chetan Arora 0002, Dron Khanna, Tomas Herda, Usman Rafiq, Jorge Melegati, Eduardo Guerra 0001, Kai-Kristian Kemell, Mika Saari, Zheying Zhang, Thanh Tho Quan, Pekka Abrahamsson |
Softw. Pract. Exp. | 11 |
| 2024 | From Data to Documentation: Exploring the Use of ChatGPT's Custom Instructions for Report GenerationabstractGenerative artificial intelligence has attracted global attention and interest in software engineering research, practical applications, and in business, especially in the past two years. Tools like Gemini, Copilot, and ChatGPT have been widely studied in various professional contexts for their exceptional ability to produce human-like content. Furthermore, the number of these artificial intelligence tools is increasing explosively. Within this context we have picked a very specific use case: namely, converting notes into a full, uniform report using ChatGPT's custom instructions. In this paper, we present the process of creating custom instructions for a certain task, how the instructions were tested, and the results of our study. The study shows that even though using custom instructions alone is perhaps not quite sufficient yet for fully automatic report generation, the approach can still save a considerable amount of time and resources. This study also highlights the pitfalls and obstacles faced, and should help those who are planning to embark on a similar quest. Janne Harjamäki, Pekka Sillberg, Mika Saari, Petri Rantanen, Jari Soini, Pekka Abrahamsson |
IS | 3 |
| 2023 | Enhancing Collaborative Prototype Development: An Evaluation of the Descriptive Model for Prototyping ProcessabstractIn this article, the ongoing research on collaborative prototype development between university and enterprises is presented. The study of project featured numerous pilot cases and prototypes, executed in collaboration with organizations to address real-world challenges. This article assesses the appropriateness of the Descriptive Model for Prototyping Process (DMPP) for research project applications. We delve into two primary facets: the synergy between universities and enterprises, and the potential for artifact reusability within the DMPP. The article presents various pilot cases from the KIEMI project, highlighting the DMPP’s role in each. Furthermore, the paper evaluates the model, sets forward the challenges faced, and, finally, discusses topics for future research. Janne Harjamäki, Mika Saari, Mikko Nurminen, Petri Rantanen, Jari Soini, David Hästbacka |
EJC | 2 |
| 2020 | Modeling the Software Prototyping Process in a Research ContextabstractThe paper examines the Third Mission of universities from the point of view of company collaboration in the prototype development process. The paper presents an implementation of university-enterprise collaboration in prototype development described by means of process modeling notation. In this article, the focus is on modeling the software prototyping process in a research context. This research paper introduces prototype development in a university environment. The prototypes are made in collaboration with companies, which offered real-world use cases. The prototype development process is introduced by a modeling procedure with four example prototype cases. The research method used is an eight-step process modeling approach. The goal was to find instances of activity, artifact, resource, and role. The results of modeling are presented using textual and graphical notation. This paper describes the data elicitation, where the process knowledge is collected using stickers-on-the-wall technique, and the creation of the model is described. Finally, the shortcomings found in our existing practices and possibilities for improving our prototype development processes and practices are discussed. Mika Saari, Jari Soini, Jere Grönman, Petri Rantanen, Timo Mäkinen, Pekka Sillberg |
EJC | 1 |
| 2019 | A Study on an Evolution of a Data Collection System for Knowledge RepresentationabstractIn this article the focus is on software evolution, which is an important part of software engineering. In practice, software development does not stop when a system is delivered but continues throughout the lifetime of the system. After the system has been deployed, external pressure for change can generate new requirements for the existing software. This change aspect, which is a characteristic of software engineering, should be taken into consideration when developing and modeling new software systems. In this paper the theme was studied using experience gained from the piloting of a reference system developed in an earlier research project carried out by T ampere University of Technology. Software evaluation is examined from the point of view of system developers, administrators (maintenance), and end users based on a concrete long-term piloting period. Jari Soini, Markku Kuusisto, Petri Rantanen, Mika Saari, Pekka Sillberg |
EJC | 4 |
| 2016 | Low-energy algorithm for self-controlled Wireless Sensor NodesabstractIn Internet of Things (IoT), the lifespan of Wireless Sensor Networks (WSN) has often become an issue. Sensor nodes are typically battery powered. However, high energy consumption by Radio Frequency (RF) module limits the lifespan of sensor nodes. In conventional WSN, the frequency of data transmission is normally fixed or adjusted according to requests from the gateway. In this paper, we present a WSN system for intelligent sensing. We propose a low-energy algorithm for sensor data transmission from sensor nodes for such system. In this algorithm, the sensor nodes are able to self-control their data transmission according to the trends of data. We adopt Adaptive Duty Cycle for adjustment of data transmission frequency and Compressive Sensing (CS) for sensor data compression. The simulation results show that Collective Transmission with CS-based data compression achieves 83.34% of RF energy reduction for the best-case transmission and 83.31% of RF energy reduction in the worst-case transmission, compared to the Continuous Transmission. Ahmad Muzaffar bin Baharudin, Mika Saari, Pekka Sillberg, Petri Rantanen, Jari Soini, Tadahiro Kuroda |
WINCOM | 2 |