Siamak Farshidi

dblp:208/0953 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 15 · 5 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Advancing research software engineering with AI: a research framework
abstract
Abstract The rapid adoption of Artificial Intelligence (AI) and Generative AI (GenAI) tools is transforming the creation, maintenance, and dissemination of research software. Despite their growing prevalence, the implications of these technologies for Research Software Engineering (RSE) practices remain underexplored. This work introduces AI4RSE , an emerging research domain focused on the integration of AI into the development lifecycle of research software. To investigate current trends in AI-augmented RSE, we conducted an empirical study of more than 1,500 open-source research software repositories hosted on Zenodo. Each repository was assessed using a quadrant-based typology defined by two key dimensions: software engineering maturity and the level of AI integration. Our analysis combined static and semantic code inspection, evaluation of alignment with the FAIR Principles for Research Software (FAIR4RS), and heuristic classification of generative AI usage and MLOps adoption. Repositories are categorized into four development modes: Exploratory Coding , Vibe Coding , RSE , and AI4RSE , which reflect different levels of process rigor and AI tool integration. While many projects exhibit informal development patterns, a growing subset demonstrates mature, AI-assisted workflows. This landscape reveals key challenges, such as reproducibility risks and licensing ambiguity, while also highlighting emerging opportunities, including AI-assisted testing and intelligent documentation generation. The findings support a research agenda for AI4RSE, outlining benchmarks, guidelines, and community standards to promote responsible, reproducible, and scalable adoption of AI in scientific software development.
Siamak Farshidi, Kwabena Ebo Bennin, Önder Babur, June Sallou, Ayalew Kassahun, Bedir Tekinerdogan
Autom. Softw. Eng.1
2026 A decision model for selecting FaaS platforms
abstract
Serverless computing is an umbrella paradigm that encompasses multiple service categories, including Function-as-a-Service (FaaS) and Backend-as-a-Service (BaaS). It has reshaped cloud application development by abstracting infrastructure management tasks such as provisioning, scaling, and operational maintenance from developers, thereby enabling greater focus on application logic. To realize these benefits, organizations need to select among various FaaS platforms such as AWS Lambda, Google Cloud Functions, Azure Functions, and Apache OpenWhisk to develop their serverless applications. Nevertheless, the increasing number of available FaaS platforms and their heterogeneous characteristics (e.g., timeout constraints, cold-start behavior, memory and package size limits, and event-source integration) make platform selection complex and knowledge-intensive for the decision makers. The main objective of this study is to support decision-makers in selecting appropriate FaaS platforms by designing an effective and systematic decision model. The model aims to simplify the selection process, reduce time and effort, and provide deeper insights into platform suitability based on specific organizational requirements. We employed a mixed-method research design to develop a decision model for the FaaS platforms selection problem. The model contains a mapping of 219 features across 16 FaaS platforms. The model was evaluated through five real-world case studies conducted at different software development companies. It suggests and prioritizes more than one FaaS platform based on the participants’ requirements. The case study participants reported that the model offered valuable insights and significantly simplified the selection process by reducing the time and costs associated with the decision-making process. We observe in the empirical evidence that decision-makers can make more rational, efficient, and effective decisions with the decision model. Additionally, the model provides reusable insights that can support future research, such as developing new frameworks and solutions for emerging challenges in serverless computing.
Siamak Farshidi, Muhammad Azeem Akbar, Rafael Capilla, Slinger Jansen, Kari Smolander
Inf. Softw. Technol.2
2026 Introduction to the special issue - software and society: ethics, equity, and sustainability in software
Antti Knutas, Sonja Hyrynsalmi, Nicolas Jullien, Slinger Jansen, Siamak Farshidi
Inf. Softw. Technol.5
2026 A data-driven decision model for selecting ML models in research software
abstract
Context: The process of selecting machine learning models is complex for research software engineers, requiring careful consideration of factors like trainability and comprehensibility to ensure long-term usability and success. Objective: This study aims to develop and evaluate a data-driven decision model that supports research software engineers in systematically selecting suitable ML models for integration into research software. Method: A meta-model was created to guide model selection, drawing from systematic literature reviews, expert interviews, case studies, and design science. Each phase contributed valuable insights and refined the decision-making framework. Results: The study analyzed 43 models across 72 attributes, resulting in a taxonomy of ML paradigms, approaches, and domains. Key findings include trends in model selection, combinations, evaluation metrics, and datasets. The decision model was further refined through expert feedback and validated with 11 case studies. Contribution: This data-driven decision model supports research software engineers in selecting optimal ML models for integration into research software. Continued development is recommended to enhance its accuracy and applicability across varied research scenarios.
Elena Baninemeh, Lex Steffens, Slinger Jansen, Siamak Farshidi
J. Syst. Softw.4
2025 D-VRE: From a Jupyter-enabled private research environment to decentralized collaborative research ecosystem
abstract
Today, scientific research is increasingly becoming data-centric and compute-intensive, relying on data and models across distributed sources. However, challenges still exist in the traditional cooperation mode, given the high storage and computing costs, geolocation barriers, and local confidentiality regulations. The Jupyter environment has recently emerged and evolved into a vital virtual research environment for scientific computing, which researchers can use to scale computational analyses up to larger datasets and high-performance computing resources. Nevertheless, existing approaches lack robust support of a decentralized cooperation mode to unlock the full potential of decentralized collaborative scientific research, e.g., seamlessly secure data sharing. In this work, we change the basic structure and legacy norms of current research environments via the seamless integration of Jupyter with Ethereum blockchain capabilities. As such, it creates a Decentralized Virtual Research Environment (D-VRE) from private computational notebooks to a decentralized collaborative research ecosystem. We propose a novel architecture for the D-VRE and prototype some essential D-VRE elements for enabling secure data sharing with decentralized identity, user-centric agreement-making, membership, and research asset management. To validate our method, we conduct an experimental study to test all functionalities of D-VRE smart contracts and their gas consumption. In addition, we deploy the D-VRE prototype on a test net of the Ethereum blockchain for demonstration. The feedback from the studies showcases the current prototype's usability, ease of use, and potential, and suggests further improvements.
Yuandou Wang, Sheejan Tripathi, Siamak Farshidi, Zhiming Zhao
Blockchain Res. Appl.3
2025 Search Multiple Types of Research Assets From Jupyter Notebook
abstract
ABSTRACT Objective Data science and machine learning methodologies are essential to address complex scientific challenges across various domains. These advancements generate numerous research assets such as datasets, software tools, and workflows, which are shared within the open science community. Concurrently, computational notebook environments like Jupyter Notebook, along with platforms like Google Colab and Kaggle Kernel, facilitate data science research and machine learning workflows, transforming data analysis, model development, and knowledge sharing processes. The proliferation of computational notebooks has further enriched the pool of valuable research assets. Researchers frequently require efficient access to these assets to advance their work, yet current tools often require navigating multiple websites and portals, leading to inefficiency and information overload. The challenge is compounded when relying on general web search engines that might not adequately highlight niche scientific resources. Methods To address these issues, we propose the development of an innovative Multiple Research Asset Search (MRAS) system designed to index diverse research assets from heterogeneous sources, offering a unified search interface for researchers. Our system aims to significantly improve the discovery of computational notebooks and datasets, facilitating data‐driven research. Results We developed a pipeline for data extraction and indexing, reviewed and applied state‐of‐the‐art ranking algorithms, enhanced indexing documents with content analysis, and created a Jupyter extension for asset discovery within the working environment. Conclusion This work is structured to detail our approach, literature review, system development, empirical validation, results, and conclusions, illustrating the potential impact of our MRAS system on scientific research efficiency.
Siamak Farshidi, Zhiming Zhao
Softw. Pract. Exp.2
2025 AE-MCDM: an autoencoder-based multi-criteria decision-making approach for unsupervised feature selection
Amin Hashemi, Mohammad Bagher Dowlatshahi, Siamak Farshidi, Parham Moradi
J. Supercomput.3
2025 Decision Support Model for Selecting the Optimal Blockchain Oracle Platform: An Evaluation of Key Factors
abstract
Smart contract-based applications are executed in a blockchain environment, and they cannot directly access data from external systems, which is required for the service provision of these applications. Instead, smart contracts use agents known as blockchain oracles to collect and provide data feeds to the contracts. The functionality and compatibility with smart contract applications need to be considered when selecting the best-fit oracle platform. As the number of oracle alternatives and their features increases, the decision-making process becomes increasingly complex. Selecting the wrong or sub-optimal oracle is costly and may lead to severe security risks. This article provides a decision support model for the oracle selection problem. The model supports smart contract decision-makers in selecting a secure, cost-effective, and feasible oracle platform for their applications. We interviewed oracle co-founders and smart contracts experts to refine and validate the decision model. Two real-world smart contract application case studies were used to evaluate the model. Our model prioritises and suggests more than one possible oracle platform based on the developer’s required criteria, security assessment and cost analysis. Moreover, this guided decision model serves to reveal issues that may go unnoticed if done haphazardly, reduce decision-making efforts and provide a cost-effective solution.
Sabreen Ahmadjee, Carlos Joseph Mera-Gómez, Siamak Farshidi, Rami Bahsoon, Rick Kazman
ACM Trans. Softw. Eng. Methodol.3
2024 Evaluating Software Quality Through User Reviews: The ISOftSentiment Tool
Fang Hou 0001, Siamak Farshidi, Slinger Jansen
PROFES3
2024 Business process modeling language selection for research modelers
abstract
Abstract Business process modeling is a crucial aspect of domains such as Business Process Management and Software Engineering. The availability of various BPM languages in the market makes it challenging for process modelers to select the best-fit BPM language for a specific process modeling task. A decision model is necessary to systematically capture and make scattered knowledge on BPM languages available for reuse by process modelers and academics. This paper presents a decision model for the BPM language selection problem in research projects. The model contains mappings of 72 BPM features to 23 BPM languages. We validated and refined the decision model through 10 expert interviews with domain experts from various organizations. We evaluated the efficiency, validity, and generality of the decision model by conducting four case studies of academic research projects with their original researchers. The results confirmed that the decision model supports process modelers in the selection process by providing more insights into the decision process. Based on the empirical evidence from the case studies and domain expert feedback, we conclude that having the knowledge readily available in the decision model supports academics in making more informed decisions that align with their preferences and prioritized requirements. Furthermore, the captured knowledge provides a comprehensive overview of BPM languages, features, and quality characteristics that other researchers can employ to tackle future research challenges. Our observations indicate that BPMN is a commonly used modeling language for process modeling. Therefore, it is more sensible for academics to explain why they did not select BPMN than to discuss why they chose it for their research project(s).
Siamak Farshidi, Izaak Beer Kwantes, Slinger Jansen
Softw. Syst. Model.1
2024 Understanding user intent modeling for conversational recommender systems: a systematic literature review
abstract
Abstract User intent modeling in natural language processing deciphers user requests to allow for personalized responses. The substantial volume of research (exceeding 13,000 publications in the last decade) underscores the significance of understanding prevalent models in AI systems, with a focus on conversational recommender systems. We conducted a systematic literature review to identify models frequently employed for intent modeling in conversational recommender systems. From the collected data, we developed a decision model to assist researchers in selecting the most suitable models for their systems. Furthermore, we conducted two case studies to assess the utility of our proposed decision model in guiding research modelers in selecting user intent modeling models for developing their conversational recommender systems. Our study analyzed 59 distinct models and identified 74 commonly used features. We provided insights into potential model combinations, trends in model selection, quality concerns, evaluation measures, and frequently used datasets for training and evaluating these models. The study offers practical insights into the domain of user intent modeling, specifically enhancing the development of conversational recommender systems. The introduced decision model provides a structured framework, enabling researchers to navigate the selection of the most apt intent modeling methods for conversational recommender systems.
Siamak Farshidi, Kiyan Rezaee, Sara Mazaheri, Amir Hossein Rahimi, Ali Dadashzadeh, Morteza Ziabakhsh, Sadegh Eskandari, Slinger Jansen
User Model. User Adapt. Interact.1
2023 FAIRSECO: An Extensible Framework for Impact Measurement of Research Software
abstract
The growing usage of research software in the research community has highlighted the need to recognize and acknowledge the contributions made not only by researchers but also by Research Software Engineers. However, the existing methods for crediting research software and Research Software Engineers have proven to be insufficient. In response, we have developed FAIRSECO, an extensible open source framework with the objective of assessing the impact of research software in research through the evaluation of various factors. The FAIRSECO framework addresses two critical information needs: firstly, it provides potential users of research software with metrics related to software quality and FAIRness. Secondly, the framework provides information for those who wish to measure the success of a project by offering impact data. By exploring the quality and impact of research software, our aim is to ensure that Research Software Engineers receive the recognition they deserve for their valuable contributions.
Deekshitha, Siamak Farshidi, Jason Maassen, Rena Bakhshi, Rob van Nieuwpoort, Slinger Jansen
e-Science2
2023 A decision model for decentralized autonomous organization platform selection: Three industry case studies
abstract
Decentralized autonomous organizations are a new form of smart contract based governance. Decentralized autonomous organization platforms, which support the creation of such organizations, are becoming increasingly popular, such as Aragon and Colony. Selecting the best fitting platform is challenging for organizations, as a significant number of decision criteria, such as popularity, developer availability, governance issues, and consistent documentation of such platforms, should be considered. Additionally, decision-makers at the organizations are not experts in every domain, so they must continuously acquire volatile knowledge regarding such platforms. Supporting decision-makers in selecting the right decentralized autonomous organizations by designing an effective decision model is the main objective of this study. We aim to provide more insight into their selection process and reduce time and effort significantly by designing a decision model. This study presents a decision model for the decentralized autonomous organization platform selection problem. The decision model captures knowledge regarding such platforms and concepts systematically. The decision model is based on an existing theoretical framework that assists software engineers with a set of Multi-Criteria Decision-Making problems in software production. We conducted three industry case studies in the context of three decentralized autonomous organizations to evaluate the effectiveness and efficiency of the decision model in assisting decision-makers. The case study participants declared that the decision model provides significantly more insight into their selection process and reduces time and effort. We observe in the empirical evidence from the case studies that decision-makers can make more rational, efficient, and effective decisions with the decision model. Furthermore, the reusable form of captured knowledge regarding Decentralized Autonomous Organization Platforms can be employed by other researchers in their future investigations.
Elena Baninemeh, Siamak Farshidi, Slinger Jansen
Blockchain Res. Appl.2
2022 Context-Aware Notebook Search in a Jupyter-Based Virtual Research Environment
abstract
Computational notebook environments such as the Jupyter play an increasingly important role in data-centric research for prototyping computational experiments, documenting code implementations, and sharing scientific results. Effectively discovering and reusing notebooks available on the web can reduce repetitive work and facilitate scientific innovations. However, general-purpose web search engines (e.g., Google Search) do not explicitly index the contents of notebooks, and notebook repositories (e.g., Kaggle and GitHub) require users to create domain-specific queries based on the metadata in the notebook catalogs, which fail to capture the working contexts in the notebook environment. This poster presents a Context-aware Notebook Search Framework (CANSF) to enable a researcher to seamlessly discover external notebooks based on semantic contexts of the literate programming activities in the Jupyter environment.
Siamak Farshidi, Riccardo Bianchi, Spiros Koulouzis, Zhiming Zhao
e-Science2
2022 An Adaptable Indexing Pipeline for Enriching Meta Information of Datasets from Heterogeneous Repositories
Siamak Farshidi, Zhiming Zhao
PAKDD (2)1
2022 Featured Cover
abstract
The cover image is based on the Research Article Notebook-as-a-VRE (NaaVRE): From private notebooks to a collaborative cloud virtual research environment by Zhiming Zhao et al., https://doi.org/10.1002/spe.3098.
Zhiming Zhao, Spiros Koulouzis, Riccardo Bianchi, Siamak Farshidi, Zeshun Shi, Ruyue Xin, Yuandou Wang, Yifang Shi 0002, Joris Timmermans, W. Daniel Kissling
Softw. Pract. Exp.4
2022 Notebook-as-a-VRE (NaaVRE): From private notebooks to a collaborative cloud virtual research environment
abstract
Abstract Virtual research environments (VREs) provide user‐centric support in the lifecycle of research activities, for example, discovering and accessing research assets or composing and executing application workflows. A typical VRE is often implemented as an integrated environment, including a catalog of research assets, a workflow management system, a data management framework, and tools for enabling user collaboration. In contrast, notebook environments like Jupyter allow researchers to rapidly prototype scientific code and share their experiments as online accessible notebooks. Jupyter can support several popular languages used by data scientists, such as Python, R, and Julia. However, such notebook environments do not have seamless support for running heavy computations on remote infrastructure or finding and accessing collaborative software code inside notebooks. This article investigates the gap between a notebook environment and a VRE and proposes an embedded VRE solution for the Jupyter environment called Notebook‐as‐a‐VRE (NaaVRE). The NaaVRE solution provides functional components via a component marketplace and allows users to create a customized VRE on top of the Jupyter environment. From the VRE, a user can search research assets (data, software, and algorithms), compose workflows, manage the lifecycle of an experiment, and share the results among users in the community. We demonstrate how such a solution can enhance a legacy workflow that uses Light Detection and Ranging (LiDAR) data from country‐wide airborne laser scanning surveys for deriving geospatial data products of ecosystem structure at high resolution over broad spatial extents. This enables users to scale out the processing of multi‐terabyte LiDAR point clouds for ecological applications to more data sources in a distributed cloud environment. Similar applications could be developed for workflows producing other essential biodiversity variables.
Zhiming Zhao, Spiros Koulouzis, Riccardo Bianchi, Siamak Farshidi, Zeshun Shi, Ruyue Xin, Yuandou Wang, Yifang Shi 0002, Joris Timmermans, W. Daniel Kissling
Softw. Pract. Exp.4
2021 Unsupervised Anomaly Detection in Data Quality Control
abstract
Data is one of the most valuable assets of an organization and has a tremendous impact on its long-term success and decision-making processes. Typically, organizational data error and outlier detection processes perform manually and reactively, making them time-consuming and prone to human errors. Additionally, rich data types, unlabeled data, and increased volume have made such data more complex. Accordingly, an automated anomaly detection approach is required to improve data management and quality control processes. This study introduces an unsupervised anomaly detection approach based on models comparison, consensus learning, and a combination of rules of thumb with iterative hyper-parameter tuning to increase data quality. Furthermore, a domain expert is considered a human in the loop to evaluate and check the data quality and to judge the output of the unsupervised model. An experiment has been conducted to assess the proposed approach in the context of a case study. The experiment results confirm that the proposed approach can improve the quality of organizational data and facilitate anomaly detection processes.
Lex Poon, Siamak Farshidi, Zhiming Zhao
IEEE BigData2
2021 A decision model for programming language ecosystem selection: Seven industry case studies
abstract
Software development is a continuous decision-making process that mainly relies on the software engineer’s experience and intuition. One of the essential decisions in the early stages of the process is selecting the best fitting programming language ecosystem based on the project requirements. A significant number of criteria, such as developer availability and consistent documentation, in addition to the number of available options in the market, lead to a challenging decision-making process. As the selection of programming language ecosystems depends on the application to be developed and its environment, a decision model is required to analyze the selection problem using systematic identification and evaluation of potential alternatives for a development project. Recently, we introduced a framework to build decision models for technology selection problems in software production. Furthermore, we designed and implemented a decision support system that uses such decision models to support software engineers with their decision-making problems. This study presents a decision model based on the framework for the programming language ecosystem selection problem. The decision model has been evaluated through seven real-world case studies at seven software development companies. The case study participants declared that the approach provides significantly more insight into the programming language ecosystem selection process and decreases the decision-making process’s time and cost. With the decision model, software engineers can more rapidly evaluate and select programming language ecosystems. Having the knowledge in the decision model readily available supports software engineers in making more efficient and effective decisions that meet their requirements and priorities. Furthermore, such reusable knowledge can be employed by other researchers to develop new concepts and solutions for future challenges.
Siamak Farshidi, Slinger Jansen, Mahdi Deldar
Inf. Softw. Technol.1
2021 Model-driven development platform selection: four industry case studies
abstract
Abstract Model-driven development platforms shift the focus of software development activity from coding to modeling for enterprises. A significant number of such platforms are available in the market. Selecting the best fitting platform is challenging, as domain experts are not typically model-driven deployment platform experts and have limited time for acquiring the needed knowledge. We model the problem as a multi-criteria decision-making problem and capture knowledge systematically about the features and qualities of 30 alternative platforms. Through four industry case studies, we confirm that the model supports decision-makers with the selection problem by reducing the time and cost of the decision-making process and by providing a richer list of options than the enterprises considered initially. We show that having decision knowledge readily available supports decision-makers in making more rational, efficient, and effective decisions. The study’s theoretical contribution is the observation that the decision framework provides a reliable approach for creating decision models in software production.
Siamak Farshidi, Slinger Jansen, Sven Fortuin
Softw. Syst. Model.1
2020 Capturing software architecture knowledge for pattern-driven design
abstract
Software architecture is a knowledge-intensive field. One mechanism for storing architecture knowledge is the recognition and description of architectural patterns. Selecting architectural patterns is a challenging task for software architects, as knowledge about these patterns is scattered among a wide range of literature. We report on a systematic literature review, intending to build a decision model for the architectural pattern selection problem. Moreover, twelve experienced practitioners at software-producing organizations evaluated the usability and usefulness of the extracted knowledge. An overview is provided of 29 patterns and their effects on 40 quality attributes. Furthermore, we report in which systems the 29 patterns are applied and in which combinations. The practitioners confirmed that architectural knowledge supports software architects with their decision-making process to select a set of patterns for a new problem. We investigate the potential trends among architects to select patterns. With the knowledge available, architects can more rapidly select and eliminate combinations of patterns to design solutions. Having this knowledge readily available supports software architects in making more efficient and effective design decisions that meet their quality concerns.
Siamak Farshidi, Slinger Jansen, Jan Martijn E. M. van der Werf
J. Syst. Softw.1