Valentina Golendukhina

dblp:317/1002 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-8274-5924ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Visualization Tools for Machine Learning Pipelines: A Review
abstract
The increasing complexity and adoption of machine learning (ML) pipelines has led to a rising demand for effective visualization tools. This paper presents a comprehensive review of existing tools for visualizing data flow in machine learning (ML) pipelines. We highlight the tools’ purposes, integration methods, and visualization techniques. We collected and analyzed 22 open-source tools and concepts, analyzing their features and classifying them based on their primary purpose. Our analysis revealed five main purposes of visualization tools: exploration, explanation, visual development, comparison, monitoring. We provide an analysis of their integration methods, from standalone visual interfaces to code-level libraries, as well as a review of various visualization techniques, including Directed Acyclic Graphs (DAGs), pipeline matrices, and annotated visualizations. Our findings highlight the importance of visualization in enhancing the interpretability and efficiency of ML workflows. Moreover, the paper provides key limitations and challenges in current visualization methods to promote future research directions enhancing the usability and functionality of ML pipeline visualization tools.
Valentina Golendukhina, Michael Felderer, Lisa Sonnleithner
PacificVis1
2024 A Review of Publicly Available Datasets from Manufacturing Systems
abstract
Smart manufacturing systems adapt to changes in the environment, which are detected via sensors. Decisions are then made by the control software of such manufacturing systems, e.g., by integrated AI-based algorithms. The design and evaluation of such algorithms require the availability of high-quality datasets. This paper provides an overview of the existing publicly available manufacturing datasets, offering a detailed exploration of the current landscape of shared data resources in the manufacturing sector and highlighting the utility of these datasets. The review identifies nine notable datasets with extensive documentation and comprehensive data coverage. Each of these datasets is described in detail and can be used for developing intelligent manufacturing systems, assessing their quality, and reporting on open gaps.
Valentina Golendukhina, Bianca Wiesmayr, Michael Felderer
ETFA1
2024 Unveiling Data Preprocessing Patterns in Computational Notebooks
abstract
Data preprocessing, which includes data integration, cleaning, and transformation, is often a time and effort-intensive step due to its fundamental importance. This crucial phase is integral for ensuring the quality and suitability of data for sub-sequent stages, such as feature engineering and model training in Machine Learning-enabled and data-driven systems. This paper provides an extensive overview of data preprocessing functions in Python and examines their application and prevalence in computational notebooks by analyzing 149,048 computational notebooks collected from Kaggle. Despite the crucial role played by data preprocessing in model performance, our results expose a significant lack of emphasis on data preprocessing activities in the examined notebooks. Notably, users holding the highest rankings tend to skip data preprocessing steps and focus on model-related activities. Although other users exhibit more frequent incorporation of data preprocessing methods, the overall prevalence remains relatively limited. We discovered that data preparation practices such as missing values are present in 20 % to 60 % of the notebooks depending on the competition, whereas outliers handling is only present in less than 20% of the analyzed scripts. The most frequently and consistently applied practices are the data transformation methods.
Valentina Golendukhina, Michael Felderer
SEAA1
2024 Data pipeline quality: Influencing factors, root causes of data-related issues, and processing problem areas for developers
abstract
Data pipelines are an integral part of various modern data-driven systems. However, despite their importance, they are often unreliable and deliver poor-quality data. A critical step toward improving this situation is a solid understanding of the aspects contributing to the quality of data pipelines. Therefore, this article first introduces a taxonomy of 41 factors that influence the ability of data pipelines to provide quality data. The taxonomy is based on a multivocal literature review and validated by eight interviews with experts from the data engineering domain. Data, infrastructure, life cycle management, development & deployment, and processing were found to be the main influencing themes. Second, we investigate the root causes of data-related issues, their location in data pipelines, and the main topics of data pipeline processing issues for developers by mining GitHub projects and Stack Overflow posts. We found data-related issues to be primarily caused by incorrect data types (33%), mainly occurring in the data cleaning stage of pipelines (35%). Data integration and ingestion tasks were found to be the most asked topics of developers, accounting for nearly half (47%) of all questions. Compatibility issues were found to be a separate problem area in addition to issues corresponding to the usual data pipeline processing areas (i.e., data loading, ingestion, integration, cleaning, and transformation). These findings suggest that future research efforts should focus on analyzing compatibility and data type issues in more depth and assisting developers in data integration and ingestion tasks. The proposed taxonomy is valuable to practitioners in the context of quality assurance activities and fosters future research into data pipeline quality.
Harald Foidl, Valentina Golendukhina, Rudolf Ramler, Michael Felderer
J. Syst. Softw.2
2023 Automation and Development Effort in Continuous AI Development: A Practitioners' Survey
abstract
The widespread adoption of AI-enabled systems and their required continuous development and deployment (MLOps) sparks research interest due to the added intricacy of automatically handling data, code, and the model itself. A better understanding of the stages for the continuous development of AI, namely Data Handling, Model Learning, Software Development, and System Operations, and the respective tasks can help to optimize and improve their effectiveness.Thus, this paper explores the degree of automation, development effort, importance, utilization of computing resources, and factors contributing to automation throughout these stages and tasks. We conducted a questionnaire-based global survey to explore these topics by analyzing 150 responses from experienced AI, data, and MLOps engineers.The results determined that the stage System Operations is mainly automated. Whereas several tasks from the other three stages (e.g., data cleaning, data quality assurance, model design, model improvement, and system level quality assurance) are more often partially automated than automated, and documentation-related tasks are mostly not automated or developed. Participants required the highest development effort for the stage Data Handling. Furthermore, the study reveals a negative correlation between automation and the perceived development effort, whereas the importance of the tasks does not seem to affect automation. 93% of participants consider the availability of computing resources, with model training, data transformation, and data cleaning ranked as the most resource-intensive tasks.
Monika Steidl, Valentina Golendukhina, Michael Felderer, Rudolf Ramler
SEAA2
2022 What is software quality for AI engineers?: towards a thinning of the fog
abstract
It is often overseen that AI-enabled systems are also software systems and therefore rely on software quality assurance (SQA). Thus, the goal of this study is to investigate the software quality assurance strategies adopted during the development, integration, and maintenance of AI/ML components and code. We conducted semi-structured interviews with representatives of ten Austrian SMEs that develop AI-enabled systems. A qualitative analysis of the interview data identified 12 issues in the development of AI/ML components. Furthermore, we identified when quality issues arise in AI/ML components and how they are detected. The results of this study should guide future work on software quality assurance processes and techniques for AI/ML components.
Valentina Golendukhina, Valentina Lenarduzzi, Michael Felderer
CAIN1