Michael Smit

dblp:19/86 · DBLP profile ↗
← Back
7ranked-venue papers in the field
1as first author
2since 2021 · last 2023
0000-0002-2028-4317ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 4 (1 first)Data Mining & Knowledge Discovery · 3
YearPublicationVenuePosition
2023 Designing a Natural Language Processing System to Support Social Science Research
abstract
The rapid development of machine learning has delivered new approaches, methods, and tools to multiple domains. I see potential for these developments, specifically natural language processing (NLP), to provide new insights, novel methods, and larger scale to social science research. However, novel NLP methods require substantial technical skills to implement. Some of the highest adoption of novel technical tools is in the area of social media analysis, where the volume of source material can overwhelm methods that rely on human capacity. My PhD dissertation aims to bridge the gap between NLP technologies and the unique needs of social science research by contributing to the development of an open-source NLP tool specifically tailored for social science researchers that reduces barriers to entry. The goal is to empower social science researchers by providing more opportunities to explore data in novel ways. This paper outlines the objectives, methodology, and expected outcomes of the proposed research study, which includes designing the development process, requirement analysis, prototyping an NLP tool, evaluating its usability and performance, and providing support for its integration into the research workflow.
Keshava Pallavi Gone, Michael Smit
ASONAM2
2023 Natural Language Processing to Understand Human Activities Impacted by Hydroelectric Energy Projects
abstract
Sustainability and large infrastructure projects are intricately linked. When considering large and disruptive infrastructure projects, proponents, regulators, and stakeholders consider a variety of information sources to understand how the project and its environmental changes will impact people living in the region. This is often done using small data, through interviews, town halls, and surveys. Big Data approaches have potential to augment these methods; for example, novel approaches natural language processing (NLP) on social media data can reveal human-nature interactions in impacted areas. Existing literature suggested that sentiment and frequency analyses can help us understand public attitudes. As a case study to illustrate this potential, we have examined public engagement with landscapes impacted by large hydroelectric development. We examined S6,122 geotagged image captions within 35 well-defined geographical locations associated with three dam projects in Canada: Mactaquac dam (New Brunswick), Oldman dam (Alberta), and Site C dam (British Columbia) on Instagram. The sentiment results of the study expressed a high positive attitude towards the three landscapes and frequency analysis exposed that the people are more likely to engage with the landscape in the context of fun, recreation, exercise activities and vacations. This can inform planning processes to ensure that impacts to these important interactions are examined and mitigated, achieving long-term environmental, social, and economic well-being.
Keshava Pallavi Gone, Michael Smit
IEEE Big Data3
2019 Code Convention Adherence in Research Data Infrastructure Software: An Exploratory Study
abstract
Science is rapidly evolving, incorporating technology like autonomous vehicles, high-throughput scientific instruments, high-fidelity numerical models, and sensor networks, all generating data with increasing frequency, variety, and volume. Scientists committed to open science are interested in sharing this data, which requires research data infrastructure (RDI). The software underlying RDI is often created and/or deployed by people who have not received formal training in software engineering, or at organizations with primary mandates that do not include software development. Our understanding of software engineering as a field and practice does not universally translate to this software. As RDI software is pushed to handle larger data sets, and used to share data more widely, it is important to understand the maintainability, the resilience of the development community, and other indicators of long-term software project health. While there is a body of research on scientific software, and on free and open source software, it is not known if existing approaches to assessing these properties are effective for RDI software. In this exploratory study, we calculate one proxy measure for maintainability (code convention adherence) for a popular ocean data management system, and compare the results with four open source projects, and with the apparent experience of users as captured in public mailing lists and an issue tracker. The results advance our limited understanding of this type of software, and inform hypothesis generation and future research design.
Michael Smit
IEEE BigData1
2017 Identifying and mitigating risks to the quality of open data in the post-truth era
abstract
Big Data analysis often relies on open data, integrating it with large private data sets, using it as ground truth information, or providing it as part of the input to large simulations. Data can be released openly by governments to achieve various objectives: transparency, informing citizen engagement, or supporting private enterprise, to name a few. To the latter objective, Big Data analytics algorithms rely on high-quality, timely access to various data sources, including open data. Examples include retail analytics drawing on open demographic data and weather forecast systems drawing on open weather and climate data. In this paper, we describe the rise of post-truth in society, and the risks this poses to the quality, integrity, and authenticity of open data. We also discuss approaches to identifying, assessing, and mitigating these risks, and suggest future steps to manage this data quality concern.
Adrienne Colborne, Michael Smit
IEEE BigData2
2017 Preparing data managers to support open ocean science: Required competencies, assessed gaps, and the role of experiential learning
abstract
Ocean science is experiencing an explosion of data as researchers employ a widening variety of sensors, operating at higher fidelity and frequency, to inform our understanding of the global ocean. This is further complicated by the increasing integration of open science data from other disciplines to analyze complex systems, like climate change, animal migration, and sea/air interaction. This shift has been unplanned, chaotic, and emergent, and has placed the onus on researchers to stay current with best practices for managing, analyzing, and sharing data. Ocean scientists who do not have the technical skill to manage this data are turning to technologists, on the assumption they have the expertise required to help. To test this assumption, we examined an experiential learning program that placed technologists at ocean data centres in Canada, conducting interviews with students and employers to identify the competencies they believed were required to manage ocean data, which were missing in students' education up to that point, and which students gained during the work term placement.
Lee Wilson, Adrienne Colborne, Michael Smit
IEEE BigData3
2016 Toward understanding how users respond to rumours in social media
abstract
As the spread of rumours has been increasing every day in online social networks (OSNs), it is important to analyze and understand this phenomenon. Damage caused by the spread of rumours is difficult to handle without a full understanding of the dynamics behind it. One of the central steps of understanding rumour spread is to analyze who spread rumours online, why, and how. In this research, we focus on the steps who and why by describing, implementing, and evaluating an approach that studies whether or not a group of users is actively involved in rumour discussions, and assesses rumour-spreading personality types in OSNs. We implement this general approach using Reddit data, and demonstrate its use by determining which users engage with a recurring rumour, and analyzing their comments using qualitative methods. We find that we can reliably classify users into one of three categories: (1) “Generally support a false rumour”, (2) “Generally refute a false rumour”, or (3) “Generally joke about a false rumour”. Combining text mining techniques, such as text classification, sentiment analysis, and social network analysis, we aim to identify and classify those rumour-spreading user categories automatically and provide a more holistic view of rumour spread in OSNs.
Anh Dang, Michael Smit, Abidalrahman Mohammad, Rosane Minghim, Evangelos E. Milios
ASONAM2
2011 Learning Actions in Complex Software Systems
Koosha Golmohammadi, Michael Smit, Osmar R. Zaïane
DaWaK2