Kerk F. Kee

dblp:09/1817 · DBLP profile ↗
← Back
7ranked-venue papers in the field
5as first author
4since 2021 · last 2024
0000-0002-0543-5009ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 7 (5 first)
YearPublicationVenuePosition
2024 Leveraging Race Prediction Algorithms to Enhance Team Composition in Big Data Science Teams
abstract
As big data science projects scale in complexity, optimizing team composition has become vital for improving creativity, productivity, and project success. We explore the possibility of incorporating race prediction algorithms for enhancing racial diversity in team composition in big data science projects. This paper evaluates five race prediction algorithms—wru, ethnicolr, ethnicolr2, pyethnicity, and rethnicity—and then discuss their potential in supporting racially diverse team assembly in big data projects. Utilizing three datasets, we assess algorithm performance and applicability, emphasizing their role in building balanced teams that enhance agility, inclusivity, and bias mitigation. We present an actionable methodology for integrating demographic insights into team management. In addition, we propose ethical safeguards to ensure responsible race prediction use, recommending data privacy measures, aggregate-only data handling, and transparency in communication. We argue that when used within ethical constraints, race prediction can support robust team processes, reduce reliance on less diverse teams, and ultimately facilitate more creative and equitable big data project outcomes.
Thanathip Chumthong, Kulsawasd Jitkajornwanich, Obada Kraishan, Kerk F. Kee, Akan Narabin
IEEE Big Data4
2023 A Capacity Framework of Community Readiness for Supporting Big Data Science Projects during Cyberinfrastructure Diffusion
abstract
This study presents a capacity framework for measuring community readiness for supporting big data science projects during cyberinfrastructure (CI) diffusion. CI projects are academic big data science projects driven by data-intensive research efforts. CI projects are an interesting example of understanding big data science projects from a scientific and academic perspective. In this paper, we advanced the argument that in order for CI projects to succeed, they need to draw from five dimensions of community readiness. More specifically, we present a capacity framework of community readiness consisting of the five dimensions of national support networks, peer-to-peer support networks, CI opinion leaders, curriculum and training, and student workforce. We proposed composite scale items developed to quantitatively define and measure these five dimensions, which can be administered using a questionnaire. The overall average score and the composite scores of the five dimensions can be utilized as reflexive assessment and feedback for CI projects about the academic and professional community they belong to. Future research could statistically validate the framework via factor analyses.
Kerk F. Kee, Alex Olshansky, Shan Xu 0001, Kulsawasd Jitkajornwanich
IEEE Big Data1
2022 An Organizational Framework of Institutional Stakeholder Engagement for Capacity to Support Big Data Science Teams Towards Cyberinfrastructure Diffusion
abstract
This paper presents an organizational framework for measuring institutional stakeholder engagement for big data science teams toward cyberinfrastructure (CI) diffusion. CI projects are an academic example of big data science projects in data-intensive research efforts. CI projects provide a unique case for understanding big data science teams from an important context, which is scientific and academic in nature. We argue that the capacity of a big data science team needs to take into consideration several institutional stakeholder engagement factors, such as having a pro-CI administration, institutional CI investments, campus CI tech support, and a non-traditional research culture. We proposed composite scale items designed to quantitatively measure these four dimensions, using a self-reported questionnaire. The overall mean score and the four individual composite scores of the main dimensions can be used as reflexive feedback and assessment for teams about the macro institutional environment in which they are embedded. Future research should statistically validate the framework via (exploratory and confirmatory) factor analyses.
Kerk F. Kee, Alex Olshansky, Shan Xu 0001
IEEE Big Data1
2021 A Socio-Technical Framework for Measuring Organizational Capacity During Cyberinfrastructure Diffusion
abstract
This paper presents a socio-technical framework for measuring organizational capacity for cyberinfrastructure (CI) implementation, adoption, and diffusion at the team’s level. CI implementation is an example of big data science project in data-intensive projects funded by the US National Science Foundation (NSF), providing a unique case for understanding big data science teams from an important field that is academic and scientific in nature. We argue that organizational capacity can be defined by the three dimensions of foundational technical expertise, daily social interactions, and enduring organizational qualities. We provide scale items for measuring these three dimensions, using a questionnaire in a self-reported and self-reflexive fashion. The overall average score and the individual composite scores of the three dimensions (and their sub-dimensions) can be used as feedback and capacity building activities as intervention strategies. Future research will statistically validate the framework using exploratory and confirmatory factor analyses.
Kerk F. Kee, Alex Olshansky, Shan Xu 0001
IEEE BigData1
2018 Fuzzy-Based Conversational Recommender for Data-intensive Science Gateway Applications
abstract
Neuro-scientists are increasingly relying on parallel and distributed computing resources for analysis and visualization of their neuron simulations. Although science gateways have democratized relevant high performance/throughput resources, users require expert knowledge about programming and infrastructure configuration that is beyond the repertoire of most neuroscience programs. These factors become deterrents for the successful adoption and the ultimate diffusion (i.e., systemic spread) of science gateways in the neuroscience community. In this paper, we present a novel intuitionistic fuzzy logic based conversational recommender that can provide guidance to users when using science gateways for research and education workflows. The users interact with a context-aware chatbot that is embedded within custom web-portals to obtain simulation tools/resources to accomplish their goals. In order to ensure user goals are met, the chatbot profiles a user's cyberinfrastructure and neuroscience domain proficiency level using a `usability quadrant' approach. Simulation of user queries for an exemplary neuroscience use case demonstrates that our chatbot can provide step-by-step navigational support and generate distinct responses based on user proficiency.
Arjun Ankathatti Chandrashekara, Radha Krishna Murthy Talluri, Sai Swathi Sivarathri, Reshmi Mitra, Prasad Calyam, Kerk F. Kee, Satish S. Nair
IEEE BigData6
2018 What is Good Feedback in Big Data Projects for Cyberinfrastructure Diffusion in e-Science?
abstract
This paper investigates the role of feedback in big data projects for cyberinfrastructure (CI) diffusion in e-science. For many of these projects, large-scale and heterogeneous datasets, multidisciplinary and dispersed experts, and advanced technologies are brought together to harness analytic insights. However, without effective CI and computational tools, the accuracy and meaningfulness of analytics results are compromised. In fact, without CI tools, raw data remain raw with hidden insights, as data analytics cannot be executed at all. In order to improve such tools for meaningful results, we argue to conceptualize the communication mechanism of `feedback' in agile software development, with the goal of producing CI tools that are responsive to users. Based on a grounded analysis of interview data, we concluded that feedback helps developers in big data projects understand users' needs, makes tools user-friendly, prevents emergencies, and is better for developers than no feedback. Furthermore, good feedback is often structured, specific, actionable, timely, generalizable, and delivered in a tactful way. Despite the limitation of the findings being exploratory and yet to be evaluated experimentally, we argued that they still can motivate developers to be proactive seekers of feedback for their tools, productively guide developers' communication with users, and ultimately promote further adoption and diffusion of CI tools in e-science.
Kerk F. Kee, Jamie C. McCain
IEEE BigData1
2015 Three critical matters in big data projects for e-science: Different user groups, the mutually constitutive perspective, and virtual organizational capacity
abstract
This paper discusses three critical matters in big data projects for e-science, in which computer simulations, computational visualizations, and mathematical modeling are practiced as powerful methods to advance science. More specifically, three different user groups are identified, and their unique situations elaborated for developing project methodologies sensitive to their differences. Furthermore, `technology development and technology use,' as well as `technology and organizing' are two parallel processes argued as mutually constitutive within the individual pairs. Finally, project methodologies should include organizational capacity assessment and capacity building as strategic components to increase big data projects' likelihood of success in producing intended outcomes.
Kerk F. Kee
IEEE BigData1