Cristina Tavares

dblp:25/851 · DBLP profile ↗
← Back
3ranked-venue papers in the field
2as first author
3since 2021 · last 2023
0009-0008-6970-0108ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3 (2 first)
YearPublicationVenuePosition
2023 GPT in Data Science: A Practical Exploration of Model Selection
abstract
There is an increasing interest in leveraging Large Language Models (LLMs) for managing structured data and enhancing data science processes. Despite the potential benefits, this integration poses significant questions regarding their reliability and decision-making methodologies. Our objective is to elucidate and express the factors and assumptions guiding GPT-4’s model selection recommendations. It highlights the importance of various factors in the model selection process, including the nature of the data, problem type, performance metrics, computational resources, interpretability vs accuracy, assumptions about data, and ethical considerations. We employ a variability model to depict these factors and use toy datasets to evaluate both the model and the implementation of the identified heuristics. By contrasting these outcomes with heuristics from other platforms, our aim is to determine the effectiveness and distinctiveness of GPT-4’s methodology. This research is committed to advancing our comprehension of AI decision-making processes, especially in the realm of model selection within data science. Our efforts are directed towards creating AI systems that are more transparent and comprehensible, contributing to a more responsible and efficient practice in data science.
Nathalia Moraes do Nascimento, Cristina Tavares, Paulo S. C. Alencar, Donald D. Cowan
IEEE Big Data2
2023 Extending Variability-Aware Model Selection with Bias Detection in Machine Learning Projects
abstract
Data science projects often involve various machine learning (ML) methods that depend on data, code, and models. One of the key activities in these projects is the selection of a model or algorithm that is appropriate for the data analysis at hand. ML model selection depends on several factors, which include data-related attributes such as sample size, functional requirements such as the prediction algorithm type, and nonfunctional requirements such as performance and bias. However, the factors that influence such selection are often not well understood and explicitly represented. This paper describes ongoing work on extending an adaptive variability-aware model selection method with bias detection in ML projects. The method involves: (i) modeling the variability of the factors that affect model selection using feature models based on heuristics proposed in the literature; (ii) instantiating our variability model with added features related to bias (e.g., bias-related metrics); and (iii) conducting experiments that illustrate the method in a specific case study to illustrate our approach based on a heart failure prediction project. The proposed approach aims to advance the state of the art by making explicit factors that influence model selection, particularly those related to bias, as well as their interactions. The provided representations can transform model selection in ML projects into a non ad hoc, adaptive, and explainable process.
Cristina Tavares, Nathalia Moraes do Nascimento, Paulo S. C. Alencar, Donald D. Cowan
IEEE Big Data1
2022 Adaptive Method for Machine Learning Model Selection in Data Science Projects
abstract
Data science projects involve a machine learning (ML) process based on data, code, and models that change over time. For example, the datasets may increase in size and allow an ML model that requires larger datasets to be applied. However, the dynamic factors that influence model selection are not well understood and explicitly represented. This paper presents ongoing work on an adaptive method for ML model selection in big data science projects. The proposed method involves (i) identifying the factors that affect model selection based on heuristics proposed in the literature; and (ii) modeling the variability of these factors using a feature diagram and constraints that trigger adaptive reconfiguration, that is, changes in model selection due to changes in the variability factors. The applicability of the method is demonstrated through an illustrative use case. The proposed method can lead to an improved understanding of dynamic factors that influence model selection, how these factors explicitly affect the selection, and how the adaptive factors can be represented and automated. This improved understanding can result in a project model selection process that is less implicit and more efficient, more adaptive and explainable, and ultimately constitute a foundation for the creation of novel dynamic software product lines to support this process.
Cristina Tavares, Nathalia Moraes do Nascimento, Paulo S. C. Alencar, Donald D. Cowan
IEEE Big Data1