EDBT 2026 Demo / reviewers in the wild / expert
Jeffrey S. Saltz
dblp:61/1466 · also Jeff Saltz 0001
· DBLP profile ↗
21ranked-venue papers in the field
14as first author
8since 2021 · last 2024
0000-0002-8913-1095ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 18 (12 first)Information Retrieval & Web Search · 2 (1 first)Database Systems & Data Management · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | GenAI Tools to Improve Data Science Project OutcomesabstractThe introduction of Generative AI (GenAI) has significantly impacted data science, offering powerful tools that enhance project outcomes through automated analysis, decision support, and personalized guidance. This study investigates the features of GenAI-powered tools designed to support both individuals and teams in data science projects. Using a qualitative approach, this study identifies essential features for supporting individuals and project teams. Key findings suggest that GenAI tools should include tailored learning aids, automated data processing capabilities, and collaborative project management features that facilitate workflow efficiency. For tool features, teams prioritize resource sharing and collaborative progress, while individuals focus on personalized support and timeline management. The study also emphasizes the advantages of domain-specific GenAI tools, offering project-specific guidance and management that surpass the capabilities of generalized solutions. Akit Kumar, M. S. Lakshmi Devi, Jeffrey S. Saltz |
IEEE Big Data | 3 |
| 2023 | Bridging the Gap in AI-Driven Workflows: The Case for Domain-Specific Generative BotsabstractThe widespread adoption of generative AI tools, such as ChatGPT, has resulted in its extensive use in a broad range of situations. However, language models often generate inaccurate or misleading responses, negatively impacting its use. Developing domain-specific bots for specific work situations could enhance accuracy and robustness, enabling more effective use of Generative AI in a work context. To help explore this possibility, we developed a data science process-expert generative AI assistant (bot) and evaluated its efficacy. We observed that the bot significantly improved efficiency, guided the exploration of new concepts within data science project management, and fostered creativity. Moreover, the constant availability of the bot allowed access to expertise whenever needed. Furthermore, responses indicated people viewed the bot as a collaborative tool that enabled communication and comprehension of complex questions. In addition, a Likert-scale analysis showed that the bot has the potential to impact the data science field positively. In summary, this research underscores the value of domain-specific bots and the potential impact on data science project management, as well as in other domains. Akit Kumar, M. S. Lakshmi Devi, Jeffrey S. Saltz |
IEEE Big Data | 3 |
| 2023 | Data Science Failure: A Literature ReviewabstractData science is a multifaceted field that integrates statistics, computer science, social science, and other domains to generate valuable insights from data. Despite unprecedented development, many data science projects fail to achieve desired outcomes. This paper presents a work-in-progress systematic literature review of grey literature to explore the opinions of industry practitioners on data science failure. Specifically, this study reviews trade journals, news articles, blogs, and industry reports published from 2018-2023 to identify common data science failure themes outside of traditional academic literature. Initial findings reveal that technical, process, people, financial, and organizational frictions frequently undermine data science projects. Furthermore, risks related to AI governance, ethical considerations, CRM strategies, data quality, access, and team skills also contribute to data science failure. The analysis highlights the contextual nature of “failure,” emphasizing the importance of critical thinking that must align with data science goals and business needs. In short, the results suggest that grey literature provides unique perspectives into data science failure, which can be complementary to peer-reviewed scholarship. Sucheta Lahiri, Jeffrey S. Saltz |
IEEE Big Data | 2 |
| 2023 | Data science curriculum in the iFieldabstractMany disciplines, including the broad Field of Information (iField), have been offering Data Science (DS) programs. There have been significant efforts exploring an individual discipline's identity and unique contributions to the broader DS education landscape. To advance DS education in the iField, the iSchool Data Science Curriculum Committee (iDSCC) was formed and charged with building and recommending a DS education framework for iSchools. This paper reports on the research process and findings of a series of studies to address important questions: What is the iField identity in the multidisciplinary DS education landscape? What is the status of DS education in iField schools? What knowledge and skills should be included in the core curriculum for iField DS education? What are the jobs available for DS graduates from the iField? What are the differences between graduate-level and undergraduate-level DS education? Answers to these questions will not only distinguish an iField approach to DS education but also define critical components of DS curriculum. The results will inform individual DS programs in the iField to develop curriculum to support undergraduate and graduate DS education in their local context. Yin Zhang 0007, Dan Wu 0003, Loni Hagen, Il-Yeol Song, Javed Mostafa, Sam Gyun Oh, Theresa Dirndorfer Anderson, Chirag Shah 0001, Bradley Wade Bishop, Frank Hopfgartner, Kai Eckert 0001, Lisa Federer, Jeffrey S. Saltz |
J. Assoc. Inf. Sci. Technol. | 13 |
| 2022 | Nine Questions to Evaluate a Data Science Team's Process: Exploring a Big Data Science Team Process Evaluation Framework Via a Delphi StudyabstractWhile the lack of an effective team process is often noted as one of the key drivers for data science project inefficiencies and failures, there has been minimal research on how to evaluate a data science team’s process. Without an evaluation framework, it is difficult for data science teams to understand their team process strengths and weaknesses. To help address this challenge, this exploratory research, via a Delpha study, identified nine key questions a data science team could answer to help evaluate their process. In short, the study identified questions evaluating the team’s communication (within the team and with stakeholders). The study also identified team process questions (e.g., the use of iterations, life cycles and a prioritization process for potential tasks). Future research could explore how data science teams can best improve their process by leveraging and refining these questions as well as defining an overall data science project management evaluation framework. Jeffrey S. Saltz |
IEEE Big Data | 1 |
| 2022 | Analyzing a Data Science Online Practitioner Community: Trends and Implications for Data Science Project ManagementabstractThe overarching goal of this research was to gain an understanding of what the data science Reddit online community discussed before, during, and after COVID-19. We used a publicly available Reddit API to harvest the r/datascience subreddit first level post data. We then performed manual annotation to explore the taxonomy of trends and themes discussed by the practitioners who belonged to reddit data science community. Then, we augmented the manually annotated data using a BERT model with topic modeling. In short, the key discussion themes, in order of frequency, were: Education, Jobs, Methods (of data science), Hardware and data collection, Data visualization, and Quality. The Quality theme includes discussions on bias, transparency, and fairness. Hence, a key finding was that there were very few discussions on data science project quality, especially trying to minimize the risk of machine learning bias. As discussions on bias are not yet common, data science teams should proactively identify and address potential questions and concerns that might arise in data science projects, especially the need to increase the team’s focus on potential bias and fairness. Zhasmina Tacheva, Sucheta Lahiri, Jeffrey S. Saltz |
IEEE Big Data | 3 |
| 2021 | CRISP-DM for Data Science: Strengths, Weaknesses and Potential Next StepsabstractThis paper explores the strengths and weaknesses of CRISP-DM when used for data science projects. The paper then explores what key actions data science teams using CRISP-DM should consider that addresses CRISP-DM’s weaknesses. In brief, CRISP-DM, which is the most popular framework teams use to execute data science projects, provides an easy to understand description of the data science project workflow (i.e., the data science life cycle). However, CRISP-DM’s project phases miss some key aspects of the data science project life cycle. In addition, CRISP-DM’s task-focused approach fails to address how a team should prioritize tasks, and in general, collaborate and communicate. Hence, this paper also describes how CRISP-DM could be combined with a team coordination framework, such as Scrum or Data Driven Scrum, which is a newer collaboration framework developed to address the unique data science coordination challenges. Jeffrey S. Saltz |
IEEE BigData | 1 |
| 2021 | Identifying and Addressing 6 Key Questions when Using Data Driven ScrumabstractData Driven Scrum (DDS) enables lean data science project agility and addresses the key challenges that have been identified when using Scrum in a data science context. However, little has been written with respect to the questions or challenges teams might encounter when trying to use DDS. Based on a survey of 18 team leads trying to use DDS, this paper describes six common questions teams might encounter when trying to implement DDS, as well as how to address these challenges. Jeffrey S. Saltz, Alex Sutherland, Thibaut Jombart |
IEEE BigData | 1 |
| 2020 | Identifying the most Common Frameworks Data Science Teams Use to Structure and Coordinate their ProjectsabstractThis paper presents the results of a study focused on exploring which framework, if any, teams use to execute data science projects. The study consisted of a survey of 109 industry professionals, as well as an evaluation of relevant framework terms searched at Google. Overall, CRISP-DM was the most commonly used framework, with Scrum and Kanban being the second and third most frequently used. We note that CRISP-DM is a life cycle framework, whereas Scrum and Kanban are team coordination frameworks. Hence, this research also notes the potential demand for a framework that integrates both life cycle and team coordination aspects of leading a data science project. Jeffrey S. Saltz, Nicholas Hotz |
IEEE BigData | 1 |
| 2020 | The Need for an Enterprise Risk Management Framework for Big Data Science Projects
Jeffrey S. Saltz, Sucheta Lahiri |
DATA | 1 |
| 2019 | SKI: An Agile Framework for Data ScienceabstractThis paper explores data science project management by first noting the need for a new process management framework and then defines a process framework that effectively supports the needs of a data science team. The paper also reports on a pilot study of teams using the framework. The framework adheres to the lean Kanban philosophy but augments Kanban by providing a structured iteration process for teams to incrementally explore and learn via lean hypothesis testing. Specifically, the Structured Kanban Iteration (SKI) framework focuses on having teams define capability-based iterations (as opposed to Kanban-like no iterations or Scrumlike time-based sprints). Furthermore, unlike Kanban, the framework leverages Scrum best practices to define roles, meetings and artifacts. Thus, SKI implements the Kanban process, but with a more repeatable and structured approach. Jeffrey S. Saltz, Alex Sutherland |
IEEE BigData | 1 |
| 2019 | Achieving Agile Big Data Science: The Evolution of a Team's Agile Process MethodologyabstractWhile there has been a rapid increase in the use of data science and the related field of big data, there has been minimal discussion on how teams using these techniques should best plan, coordinate and communicate their activities. To help address this gap, this paper reports on a mixed method qualitative study exploring how a big data science team within a Fortune 500 organization used two different agile process methodologies. The study helps clarify the concept of agility within a big data science project, as well as the key process challenges teams encounter when executing a big data science project. Specifically, three key issues were identified: (a) the challenge in task duration estimation, (b) how to account for team members that might be pulled onto other tasks for short bursts and (c) coordination challenges across the different groups within the big data science team. Our findings help explain how different process methodologies might mitigate or exacerbate these challenges and supports previous research showing that big data science teams would benefit from an increased focus on their process methodology and that adopting an Agile Kanban methodology, which focuses on minimizing work-in-progress, could prove beneficial for many big data science teams. Jeffrey S. Saltz, Ivan Shamshurin |
IEEE BigData | 1 |
| 2018 | Will Deep Learning Change How Teams Execute Big Data Projects?abstractAs data continues to be produced in ever increasing quantities, and technologies such as high performance computing continue to be enhanced, the number of big data projects using advanced neural network machine learning, often referred to as deep learning, continues to increase. Unfortunately, while much has been written on the use of deep learning algorithms in terms of generating insightful analysis, much less has been written about the project management process methodologies that could enable teams to more effectively and efficiently "do" big data deep learning projects. Specifically, the rapid growth in the use of deep learning techniques might introduce new challenges with respect to how to execute a big data deep learning project, due to how deep learning models can learn features automatically. For example, feature engineering and model evaluation phases of big data projects might grow in importance, while other areas, such as model selection, might decrease in importance. Hence, this paper discusses the key research questions relating the potential impact of the use of deep learning on how teams should execute big data projects. Ivan Shamshurin, Jeffrey S. Saltz |
IEEE BigData | 2 |
| 2018 | Improving Data Science Projects by Enriching Analytical Models with Domain KnowledgeabstractDomain knowledge is very important to support the development of analytic models. However, in today's data science projects, domain knowledge is typically documented, but not captured and integrated with the actual analytic model. This raises problems in interoperability and traceability of the relevant domain knowledge that is used to develop an analytic model. To address this challenge, this paper proposes a Knowledge Enriched Analytic Model (KEAM) to enrich analytic models with domain knowledge. To explore the proposed methodology and its benefits, a case study explores the utilization of KEAM to support the development of a Bayesian Network model within the smart manufacturing domain. The case study shows that the efficiency in developing an analytic model is improved by using the proposed KEAM. Heng Zhang 0010, Utpal Roy, Jeffrey S. Saltz |
IEEE BigData | 3 |
| 2017 | The ambiguity of data science team roles and the need for a data science workforce frameworkabstractThis paper first reviews the benefits of well-defined roles and then discusses the current lack of standardized roles within the data science community, perhaps due to the newness of the field. Specifically, the paper reports on five case studies exploring five different attempts to define a standard set of roles. These case studies explore the usage of roles from an industry perspective as well as from national standard big data committee efforts. The paper then leverages the results of these case studies to explore the use of data science roles within online job postings. While some roles appeared frequently, such as data scientist and data engineer, no role was consistently used across all five case studies. Hence, the paper concludes by noting the need to create a data science workforce framework that could be used by students, employers, and academic institutions. This framework would enable organizations to staff their data science teams more accurately with the desired skillsets. Jeffrey S. Saltz, Nancy W. Grady |
IEEE BigData | 1 |
| 2017 | Does pair programming work in a data science context? An initial case studyabstractWhile pair programming has been studied extensively for software programmers, very little has been reported with respect to pair programming in a data science project. This paper reports on a case study evaluating the effectiveness of pair programming within a data science / big data context. Our findings show that pair programming can be useful for data science teams. In addition, while the driver role was similar to what has been described for software programmers, we note that the observer role had an expanded set of responsibilities, which we termed researcher activities. Further exploration is required to explore if these expanded roles are specific to data science pair programming. Jeffrey S. Saltz, Ivan Shamshurin |
IEEE BigData | 1 |
| 2017 | Predicting data science sociotechnical execution challenges by categorizing data science projectsabstractThe challenge in executing a data science project is more than just identifying the best algorithm and tool set to use. Additional sociotechnical challenges include items such as how to define the project goals and how to ensure the project is effectively managed. This paper reports on a set of case studies where researchers were embedded within data science teams and where the researcher observations and analysis was focused on the attributes that can help describe data science projects and the challenges faced by the teams executing these projects, as opposed to the algorithms and technologies that were used to perform the analytics. Based on our case studies, we identified 14 characteristics that can help describe a data science project. We then used these characteristics to create a model that defines two key dimensions of the project. Finally, by clustering the projects within these two dimensions, we identified four types of data science projects, and based on the type of project, we identified some of the sociotechnical challenges that project teams should expect to encounter when executing data science projects. Jeffrey S. Saltz, Ivan Shamshurin, Colin Connors |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2016 | Big data team process methodologies: A literature review and the identification of key factors for a project's successabstractThis paper reports on our review of published research relating to how teams work together to execute Big Data projects. Our findings suggest that there is no agreed upon standard for executing these projects but that there is a growing research focus in this area and that an improved process methodology would be useful. In addition, our synthesis also provides useful suggestions to help practitioners execute their projects, specifically our identified list of 33 important success factors for executing Big Data efforts, which are grouped by our six identified characteristics of a mature Big Data organization. Jeffrey S. Saltz, Ivan Shamshurin |
IEEE BigData | 1 |
| 2016 | Not all software engineers can become good data engineersabstractThe amount of data that businesses collect and analyze has been rapidly increasing, which has triggered an increase in big data teams. With the growth of both the number and size of big data teams, specialized roles are starting to be defined. One such role is the data engineer, who focuses on ensuring that the data is easily available for advanced analytics. Via a case study, this paper explores the role of the data engineer and the key characteristics that enable someone to be a good data engineer. The paper also explores if good software engineers could become good data engineers. Our findings show that the knowledge and skills required to be a data engineer are significantly different from those required to be a software engineer. Hence, not surprisingly, we found that that not all software engineers could become good data engineers. Jeffrey S. Saltz, Sibel Yilmazel, Özgür Yilmazel |
IEEE BigData | 1 |
| 2015 | The need for new processes, methodologies and tools to support big data teams and improve big data project effectivenessabstractAs data continues to be produced in massive amounts, with increasing volume, velocity and variety, big data projects are growing in frequency and importance. However, the growth in the use of big data has outstripped the knowledge of how to support teams that need to do big data projects. In fact, while much has been written in terms of the use of algorithms that can help generate insightful analysis, much less has been written about methodologies, tools and frameworks that could enable teams to more effectively and efficiently "do" big data projects. Hence, this paper discusses the key research questions relating methodologies, tools and frameworks to improve big data team effectiveness as well as the potential goals for a big data process methodology. Finally, the paper also discusses related domains, such as software development, operations research and business intelligence, since these fields might provide insight into how to define a big data process methodology. Jeffrey S. Saltz |
IEEE BigData | 1 |
| 2015 | Exploring the process of doing data science via an ethnographic study of a media advertising companyabstractThis paper presents the results of an ethnographic study focused on how data science projects were conducted within a global media advertising company. Observations, via embedding a researcher within the team, as well as more structured interviews and surveys, are documented. Recommendations to improve the current data science methodology within the company are also discussed. Overall, there had been little focus on the team's process methodology and the suggested process improvements would result in the company's data science projects having less risk and shorter timelines. Other big data teams might also benefit from reviewing and refining their work processes, but more work needs to be done to validate this assumption. Jeffrey S. Saltz, Ivan Shamshurin |
IEEE BigData | 1 |