David M. Bourrie

dblp:124/4749 · also David Bourrie · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0002-5009-6843ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Security and privacy · 1
YearPublicationVenuePosition
2024 I'm not fluent: How linguistic fluency, new media literacy, and personality traits influence fake news engagement behavior on social media
Stacy Miller, Philip Menard, David M. Bourrie
Inf. Manag.3
2021 SURFR: A Real-Time Platform for Non-Coding RNA Fragmentation Analysis Using Wavelets
abstract
It is well known that microRNAs (miRNAs or miRs) are small (~18-25 nt) yet highly potent non-coding RNA-derived RNAs (ndRNAs), originating from pre-miRNA fragmentation, that have been shown to alter the post-transcriptional functionality of many messenger RNAs (mRNAs). Biologically, the identification and study of miRNAs is very critical due to their increasing significance as biomarkers for many types of cancers and other genetic diseases. While empirical evidence supporting the existence of several novel ndRNAs excised from other longer non coding RNAs (ncRNAs) is growing, recent evidence suggests the full extent of their prevalence is likely underappreciated. Although some computational methods have been designed to help domain experts identify and understand miRNAs by analyzing Next Generation Sequencing (NGS) datasets, there are some crucial challenges, such as efficiency, effectiveness, and generalizability, in the state-of-the-art in-silico methods. To address such problems, our group proposed a new algorithm to mine ndRNAs by applying wavelet-based signal processing techniques as opposed to the current string-based NGS sequence alignment/analysis. However, due to novelty of the approach, our initial version of the algorithm was focused specifically on mining miRNAs, snoRNA-derived RNAs (sdRNAs) and transfer RNA (tRNA) fragments (tRFs) because of their importance in the literature plus the availability of experimentally validated databases to confirm our findings. Despite the computational issues, we still lack a basic understanding of the existence and the range of ndRNA functionalities from a) ndRNAs other than miRs, sdRNAs & tRFs in humans, and b) all ndRNAs in millions of organisms other than humans. Hence, there is an urgent requirement to automate the extraction and experimentation of ndRNAs, especially considering the rate at which NGS data is being produced. Therefore, in the current article, we extended our algorithm to be applicable to ~500 organisms—including eukaryotes, plants, bacteria, fungi, and protists—along with all their ncRNAs available in the current NCBI annotation. We also constructed a real-time user-friendly platform, SURFR, available at salts.soc.southalabama.edu/surfr, to aid domain experts and the aspiring biomedical scientists to perform RNA-Seq experiments to study ndRNAs. Not only our platform is extremely efficient, but we are also capable of allowing the users to identify, analyze, visualize, and compare ndRNAs from up to 30 NGS files to perform rigorous experimentation. Moreover, access to NGS files from public databases like SRA, and ndRNAs from private databases like TCGA are made readily available to the users to further validate their novel findings. Finally, we provide theoretical validation to examine our platform’s effectiveness.
Mohan Vamsi Kasukurthi, Dominika Houserova, Dongqi Li, Jingwei Lin, Guanhuan Yang, Shaobo Tan, David M. Bourrie, Bin Ma 0003, Glen M. Borchert, Jingshan Huang
BIBM9
2021 A New Classification Algorithm and a New Oversampling Method of Mapping Common Data Elements to the BRIDG Model
abstract
The Common Data Elements (CDEs) standard of the International organization for Standardization (ISO) 11179 is commonly used in the field of clinical data processing. The Biomedical Research Integrated Domain Group (BRIDG) model is the framework for biomedical and clinical research. Mapping CDEs to BRIDG (also known as CDE classification) would help with interoperability and data analysis in the field of clinical research. That said, manually mapping CDEs to their corresponding BRIDG class is highly time-consuming and labor-intensive. In this paper we present a new classification algorithm along with a new oversampling method. Our algorithm uses the Term Frequency-Inverse Document Frequency (TF-IDF) as the feature representation method. By assigning different weights to various attributes, we enable more important attributes to perform more important roles during the mapping process. In addition, the oversampling method generates every new attribute in the minor class by picking the length and setting the word of the new attribute according to the existing training set. Our research outcomes demonstrate significant contributions to the field in the following ways: (1) Generation of a new CDE classification algorithm that outperforms existing algorithms in the literature, including the Random Forest Classifier, Linear Support Vector Classification (SVC), Multinomial Naive Bayes (NB), Logistic Regression, and Long Short-Term Memory (LSTM) networks, in terms of accuracy, precision, recall, and F-1 score measures. (2) Generation of a new oversampling method able to improve CDE classification accuracy for Random Forest and Multinomial NB. (3) Our classification algorithm employs two novel attributes, namely “Data Element Preferred Definition” and “Document,” which are more efficient at classifying CDEs than the six attributes traditionally selected by domain experts.
Mohan Vamsi Kasukurthi, Jiajie Yang, Dongqi Li, Guanhuan Yang, Jingwei Lin, Shaobo Tan, David M. Bourrie, Bin Ma 0003, Glen M. Borchert, Jingshan Huang
BIBM9
2021 RARE: Rare Action Rule Exploration
abstract
Action rule mining seeks to generate rules that indicate what changes can be made to move an object from one class (state) to another class. An action rule is composed of the changes, known as actions, that correlate with the change in the class value. The current work in action rule mining focuses on frequent, or highly occurring, action rules. The working assumption is that the end user is interested in class transitions that move a large number of objects from the initial class to the new (final) class with a high degree of confidence. Currently, very little work focuses on rare action rules; that is, classes which occur infrequently. In this paper, we provide a definition for a rare action rule and then propose a consequent-constraint based algorithm for generating these rare action rules.
Blake A. Johns, Ryan G. Benton, Tom Johnsten, David M. Bourrie
IEEE BigData4
2021 Precise Weather Parameter Predictions for Target Regions via Neural Networks
Yihe Zhang 0001, Xu Yuan 0001, Sytske K. Kimball, Eric Rappin, Li Chen 0019, Paul J. Darby III, Tom Johnsten, Lu Peng 0001, Boisy Pitre, David M. Bourrie, Nian-Feng Tzeng
ECML/PKDD (5)10
2020 DFS3: automated distributed file system storage state reconstruction
abstract
Distributed file systems present distinctive forensic challenges in comparison to traditional locally mounted file system volume. Storage device media can number in the thousands, and forensic investigations in this setting necessitate a tailored approach to data collection. The Hadoop Distributed File System (HFDS) produces and maintains partially persistent metadata that is pursuant with a logical volume, a file system, and file addresses on the centralized server. Hence, this research investigates the viability of using a residual central server digital artifact to generate a history model of the distributed file system. The history model affords an investigator a high-level perspective of low-level events to narrow investigative process obligations. The model is generated through set-theoretic relations of the file system essential data structure. Graph-theoretic ordering is applied to the events to provide a history model. The research contribution is a rapid reconstruction of the HDFS storage state transitions generating timelines for system events to forensically assess HDFS properties with conceptual similarity to traditional low-level file system forensic tool output. The results of this research provide a prototype tool, DFS3, for rapid and noninvasive data storage state timeline reconstruction in a big data distributed file system.
Edward Harshany, Ryan G. Benton, David M. Bourrie, William Glisson
ARES3
2019 Enhancing Itemset Tree Rules and Performance
abstract
Association mining is the process of discovering relationships between items in a data set, where a group of items forms an itemset. A problem with the typical association mining approach is a large number of the generated frequent itemsets typically do not contain any items of interest to the user. Targeted association mining solves this by only deriving itemsets that include the items specified within a user's query. A popular targeted approach is the Itemset Tree, which consists of an index tree structure and algorithms for search and rule generation. Numerous enhancements to improve the Itemset Tree efficiency or extend its capability have been proposed. However, two major problems exist. First, the itemset generation process utilized by the Itemset Tree does not leverage the Apriori Principle, resulting in the unnecessary and costly generation of infrequent itemsets. Second, the Itemset Tree rule generation process has restrictions that prevent some rules containing the user query from being generated; as a result, the user misses useful information. In this paper, we offer a redesigned generation process resulting in faster itemset generation (milliseconds versus minutes) and expanded rulesets.
Jay Lewis, Ryan G. Benton, David M. Bourrie, Jennifer Lavergne
IEEE BigData3