EDBT 2026 Demo / reviewers in the wild / expert
Xiao Luo 0002
dblp:50/1585-2
· DBLP profile ↗
36ranked-venue papers
13as first author
16since 2021 · last 2024
0000-0002-3649-9785ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 7 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 3 since 2021Computer networks · 3Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Zero-shot learning to extract assessment criteria and medical services from the preventive healthcare guidelines using large language modelsabstractOBJECTIVES: The integration of these preventive guidelines with Electronic Health Records (EHRs) systems, coupled with the generation of personalized preventive care recommendations, holds significant potential for improving healthcare outcomes. Our study investigates the feasibility of using Large Language Models (LLMs) to automate the assessment criteria and risk factors from the guidelines for future analysis against medical records in EHR. MATERIALS AND METHODS: We annotated the criteria, risk factors, and preventive medical services described in the adult guidelines published by United States Preventive Services Taskforce and evaluated 3 state-of-the-art LLMs on extracting information in these categories from the guidelines automatically. RESULTS: We included 24 guidelines in this study. The LLMs can automate the extraction of all criteria, risk factors, and medical services from 9 guidelines. All 3 LLMs perform well on extracting information regarding the demographic criteria or risk factors. Some LLMs perform better on extracting the social determinants of health, family history, and preventive counseling services than the others. DISCUSSION: While LLMs demonstrate the capability to handle lengthy preventive care guidelines, several challenges persist, including constraints related to the maximum length of input tokens and the tendency to generate content rather than adhering strictly to the original input. Moreover, the utilization of LLMs in real-world clinical settings necessitates careful ethical consideration. It is imperative that healthcare professionals meticulously validate the extracted information to mitigate biases, ensure completeness, and maintain accuracy. CONCLUSION: We developed a data structure to store the annotated preventive guidelines and make it publicly available. Employing state-of-the-art LLMs to extract preventive care criteria, risk factors, and preventive care services paves the way for the future integration of these guidelines into the EHR. Xiao Luo 0002, Fattah Muhammad Tahabi, Tressica Marc, Laura Ann Haunert, Susan Storey |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | Zero-shot Learning with Minimum Instruction to Extract Social Determinants and Family History from Clinical Notes using GPT ModelabstractDemographics, social determinants of health, and family history documented in the unstructured text within the electronic health records are increasingly being studied to understand how this information can be utilized with the structured data to improve healthcare outcomes. After the GPT models were released, many studies have applied GPT models to extract this information from the narrative clinical notes. Different from the existing work, our research focuses on investigating the zero-shot learning on extracting this information together by providing minimum information to the GPT model. We utilize de-identified real-world clinical notes annotated for demographics, various social determinants, and family history information. Given that the GPT model might provide text different from the text in the original data, we explore two sets of evaluation metrics, including the traditional NER evaluation metrics and semantic similarity evaluation metrics, to completely understand the performance. Our results show that the GPT-3.5 method achieved an average of 0.975 F1 on demographics extraction, 0.615 F1 on social determinants extraction, and 0.722 F1 on family history extraction. We believe these results can be further improved through model fine-tuning or few-shots learning. Through the case studies, we also identified the limitations of the GPT models, which need to be addressed in future research. Neel Bhate, Ansh Mittal, Zhe He 0001, Xiao Luo 0002 |
IEEE Big Data | 4 |
| 2023 | Discovering COVID-19 Coughing and Breathing Patterns from Unlabeled Data Using Contrastive Learning with Varying Pre-Training Domains
Jinjin Cai, Sudip Vhaduri, Xiao Luo 0002 |
INTERSPEECH | 3 |
| 2023 | Improving Perceptions of Underrepresented Students towards Computing Majors through MentoringabstractThe low sense of belonging and self-efficacy have been identified as key factors for underrepresented students not choosing computing careers and retaining in computing disciplines. Researchers advised that without adequate mentorship and role models, many of these students do not view computing as a viable option. Even though several works utilized mentoring, they do not provide detailed guidelines on how this may have been accomplished and how it could be applied to other settings. In this study, we explore the efficacy of incorporating mentoring into the curriculum of an introductory undergraduate computing course. In this work, we seek the answer to the following two research questions: (1) How do culturally diverse mentor-mentee relationships impact the sense of belonging, computing identity, and self-efficacy of underrepresented students in computing programs? (2) How does the integration of mentoring initiatives influence the perceptions of underrepresented students toward computing majors? We implemented mentoring practices in the Fall of 2022 and results show that our mentoring interventions were able to improve participants' sense of belonging and computing identity. Mentors and mentees also shared positive opinions toward our initiatives. Shamima Mithun, Xiao Luo 0002 |
ITiCSE (1) | 2 |
| 2022 | ReferEmo: A Referential Quasi-multimodal Model for Multilabel Emotion Classification
Alvar Esperanca, Xiao Luo 0002 |
DEXA (1) | 2 |
| 2022 | InAction: Interpretable Action Decision Making for Autonomous Driving
Taotao Jing, Haifeng Xia, Renran Tian, Xiao Luo 0002, Joshua E. Domeyer, Rini Sherony, Zhengming Ding |
ECCV (38) | 5 |
| 2022 | Learning Management System Analytics to Examine the Behavior of Students in High Enrollment STEM Courses During the Transition to Online InstructionabstractThe emergence of the COVID-19 pandemic resulted in the transition to near-total online instruction in early 2020. Several studies surveyed students about the impact of the pandemic on their behavior and engagement with their education; however, those studies may not include analysis of actual student behaviors. The field of learning analytics allows researchers to examine the records made while students interact with the educational technology tools that are commonly used to facilitate instruction in Institutes of Higher Education (IHEs). In most universities, delivery of many instructional materials is conducted via a Learning Management System (LMS). In this study, we describe our process to examine student interactions within the LMS to discover any measurable changes to student behavior during the pandemic.We examined the usage logs of the Canvas LMS at a large university in the midwestern US to examine the behavior of students' interactions with high-enrollment STEM courses in two semesters: one prior to and one during the pandemic. The log data was integrated with student demographic data so that the LMS behavior of subsets of students can be compared. Machine learning algorithms including clustering models and association rule mining were applied on the data. The results of this study demonstrate that the students' behaviors did change in the transition to online instruction. Students had more frequent sessions in the LMS on both computers and mobile devices, although the duration of their mobile sessions was shorter after courses were moved online. Further, students in historically underrepresented groups in STEM fields were found to use their mobile devices more frequently for academic work. The information uncovered in this study can be used to inform future instructional design practices with the LMS to promote an equitable experience for all students. Rob Elliott, Xiao Luo 0002 |
FIE | 2 |
| 2022 | Application of unsupervised deep learning algorithms for identification of specific clusters of chronic cough patients from EMR dataabstractBACKGROUND: Chronic cough affects approximately 10% of adults. The lack of ICD codes for chronic cough makes it challenging to apply supervised learning methods to predict the characteristics of chronic cough patients, thereby requiring the identification of chronic cough patients by other mechanisms. We developed a deep clustering algorithm with auto-encoder embedding (DCAE) to identify clusters of chronic cough patients based on data from a large cohort of 264,146 patients from the Electronic Medical Records (EMR) system. We constructed features using the diagnosis within the EMR, then built a clustering-oriented loss function directly on embedded features of the deep autoencoder to jointly perform feature refinement and cluster assignment. Lastly, we performed statistical analysis on the identified clusters to characterize the chronic cough patients compared to the non-chronic cough patients. RESULTS: The experimental results show that the DCAE model generated three chronic cough clusters and one non-chronic cough patient cluster. We found various diagnoses, medications, and lab tests highly associated with chronic cough patients by comparing the chronic cough cluster with the non-chronic cough cluster. Comparison of chronic cough clusters demonstrated that certain combinations of medications and diagnoses characterize some chronic cough clusters. CONCLUSIONS: To the best of our knowledge, this study is the first to test the potential of unsupervised deep learning methods for chronic cough investigation, which also shows a great advantage over existing algorithms for patient data clustering. Wei Shao 0005, Xiao Luo 0002, Zuoyi Zhang, Zhi Han, Vasu Chandrasekaran, Vladimir Turzhitsky, Vishal Bali, Anna R. Roberts, Megan Metzger, Jarod Baker, Carmen La Rosa, Jessica Weaver, Paul Richard Dexter, Kun Huang 0001 |
BMC Bioinform. | 2 |
| 2022 | Attention-based Unsupervised Keyphrase Extraction and Phrase Graph for COVID-19 Medical Literature RetrievalabstractSearching, reading, and finding information from the massive medical text collections are challenging. A typical biomedical search engine is not feasible to navigate each article to find critical information or keyphrases. Moreover, few tools provide a visualization of the relevant phrases to the query. However, there is a need to extract the keyphrases from each document for indexing and efficient search. The transformer-based neural networks—BERT has been used for various natural language processing tasks. The built-in self-attention mechanism can capture the associations between words and phrases in a sentence. This research investigates whether the self-attentions can be utilized to extract keyphrases from a document in an unsupervised manner and identify relevancy between phrases to construct a query relevancy phrase graph to visualize the search corpus phrases on their relevancy and importance. The comparison with six baseline methods shows that the self-attention-based unsupervised keyphrase extraction works well on a medical literature dataset. This unsupervised keyphrase extraction model can also be applied to other text data. The query relevancy graph model is applied to the COVID-19 literature dataset and to demonstrate that the attention-based phrase graph can successfully identify the medical phrases relevant to the query terms. Xiao Luo 0002 |
ACM Trans. Comput. Heal. | 2 |
| 2022 | A Deep Language Model for Symptom Extraction From Clinical Text and its Application to Extract COVID-19 Symptoms From Social MediaabstractPatients experience various symptoms when they haveeither acute or chronic diseases or undergo some treatments for diseases. Symptoms are often indicators of the severity of the disease and the need for hospitalization. Symptoms are often described in free text written as clinical notes in the Electronic Health Records (EHR) and are not integrated with other clinical factors for disease prediction and healthcare outcome management. In this research, we propose a novel deep language model to extract patient-reported symptoms from clinical text. The deep language model integrates syntactic and semantic analysis for symptom extraction and identifies the actual symptoms reported by patients and conditional or negation symptoms. The deep language model can extract both complex and straightforward symptom expressions. We used a real-world clinical notes dataset to evaluate our model and demonstrated that our model achieves superior performance compared to three other state-of-the-art symptom extraction models. We extensively analyzed our model to illustrate its effectiveness by examining each component's contribution to the model. Finally, we applied our model on a COVID-19 tweets data set to extract COVID-19 symptoms. The results show that our model can identify all the symptoms suggested by the Center for Disease Control (CDC) ahead of their timeline and many rare symptoms. Xiao Luo 0002, Priyanka Gandhi, Susan Storey, Kun Huang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Contextual and Behavior Factors Extraction from Pedestrian Encounter Scenes Using Deep Language Models
Jithesh Gugan Sreeram, Xiao Luo 0002, Renran Tian |
DaWaK | 2 |
| 2021 | AttentionRank: Unsupervised Keyphrase Extraction using Self and Cross AttentionsabstractKeyword or keyphrase extraction is to identify words or phrases presenting the main topics of a document.This paper proposes the AttentionRank, a hybrid attention model, to identify keyphrases from a document in an unsupervised manner.AttentionRank calculates self-attention and cross-attention using a pretrained language model.The self-attention is designed to determine the importance of a candidate within the context of a sentence.The cross-attention is calculated to identify the semantic relevance between a candidate and sentences within a document.We evaluate the AttentionRank on three publicly available datasets against seven baselines.The results show that the AttentionRank is an effective and robust unsupervised keyphrase extraction model on both long and short documents.Source code is available on Github 1 . Xiao Luo 0002 |
EMNLP (1) | 2 |
| 2021 | Evaluating Factors for Effective Flipped Classroom Instruction in an Advanced Data Management CourseabstractThe following Research-to-Practice full paper presents the outcomes of a remote synchronous flipped classroom implementation of a senior-level data management course. As flipped classroom models of instruction have gained popularity in higher education, it has prompted the need for an investigation into content-specific methods for flipped curriculum design. Our work identified and implemented flipped classroom design factors in an Advanced Database Design (CIT 44400) course to address a gap in research around flipped classroom models within undergraduate data-science courses. Through literature review, we identified a set of eight factors of effective flipped classroom instruction and evaluated them via survey, focus group, course evaluation, student performance, and interview data. We segmented designed course activities which incorporated these eight factors with a set of learning objectives that emphasized collaborative iterative practices to position students as designers and evaluators of solutions to industry-authentic problems. Instructional design factors were evaluated via course evaluations and student surveys followed by an activity-based qualitative analysis of focus group, instructor interview, and student free-response data to understand student perceptions of instructional approaches, intended versus practical outcomes of such activities, and guidelines for future course design iterations and research. We argue instructional design must be student-centered and consider student goals alongside those that are ‘scripted’ into the course structure to best serve, motivate, and engage students. During the Fall 2020 implementation of CIT 44400, we prioritized learner independence, peer collaboration, and critical thinking in the instructional design, selecting flipped methodologies with the intention of fostering these skills in senior students. In light of Covid-19, the course was adapted to be synchronous online. Regardless of these unforeseen constraints, student performance and course evaluation data indicate that with peer and instructor feedback, students were able to apply course content appropriately in their final independent project as evidenced by a 12% increase in assessment scores and an improvement in students average overall final course grades from a ‘B’ to an ‘A-’ as compared to the previously taught Fall 2019 lecture-based section of the same class. Course evaluation scores also improved from 3.22 to 3.66 from Fall 2019 to Fall 2020. Self-reported survey data from students indicate that (a) feedback from the instructor, (b) small group work, (c) revision of work based on feedback, and (d) solution evaluations were the most positively impactful instructional design elements throughout the course. In particular, students mentioned one-on-one scaffolding from the instructor as beneficial for their learning. Students also communicated challenges faced during the pandemic, complaints, and recommendations for future course iterations. Data from focus group discussions conveyed that students (a) generally had positive learning experiences despite the constraints of learning exclusively online and (b) developed the industry-specific collaborative practices which were designed into the forefront of our instructional model. The flipped classroom design and implementation process in this research is transformative and can be employed by other STEM disciplines to design a domain-specific flipped model for their classroom which considers the needs of one's students, the challenges of the current time, and the state of the field at large. Shamima Mithun, Morgan Vickery, Xiao Luo 0002 |
FIE | 3 |
| 2021 | Modelling and visualising SSH brute force attack behaviours through a hybrid learning framework
Xiao Luo 0002, Chengchao Yao, Nur Zincir-Heywood |
Int. J. Inf. Comput. Secur. | 1 |
| 2021 | HeteroGraphRec: A heterogeneous graph-based neural networks for social recommendations
Amirreza Salamat, Xiao Luo 0002 |
Knowl. Based Syst. | 2 |
| 2021 | A Computational Framework to Analyze the Associations Between Symptoms and Cancer Patient Attributes Post Chemotherapy Using EHR DataabstractPatients with cancer, such as breast and colorectal cancer, often experience different symptoms post-chemotherapy. The symptoms could be fatigue, gastrointestinal (nausea, vomiting, lack of appetite), psychoneurological symptoms (depressive symptoms, anxiety), or other types. Previous research focused on understanding the symptoms using survey data. In this research, we propose to utilize the data within the Electronic Health Record (EHR). A computational framework is developed to use a natural language processing (NLP) pipeline to extract the clinician-documented symptoms from clinical notes. Then, a patient clustering method is based on the symptom severity levels to group the patient in clusters. The association rule mining is used to analyze the associations between symptoms and patient attributes (smoking history, number of comorbidities, diabetes status, age at diagnosis) in the patient clusters. The results show that the various symptom types and severity levels have different associations between breast and colorectal cancers and different timeframes post-chemotherapy. The results also show that patients with breast or colorectal cancers, who smoke and have severe fatigue, likely have severe gastrointestinal symptoms six months after the chemotherapy. Our framework can be generalized to analyze symptoms or symptom clusters of other chronic diseases where symptom management is critical. Xiao Luo 0002, Priyanka Gandhi, Susan Storey, Zuoyi Zhang, Zhi Han, Kun Huang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Retrieving Lab Test Related Questions from Social Q&A Sites by Combining Shallow Features and Deep Representations
Xiao Luo 0002, Zhan Zhang 0008, Zhe He 0001 |
AMIA | 2 |
| 2020 | Attention Mechanism with BERT for Content Annotation and Categorization of Pregnancy-Related Questions on a Community Q&A SiteabstractIn recent years, the social web has been increasingly used for health information seeking, sharing, and subsequent health-related research. Women often use the Internet or social networking sites to seek information related to pregnancy in different stages. They may ask questions about birth control, trying to conceive, labor, or taking care of a newborn or baby. Classifying different types of questions about pregnancy information (e.g., before, during, and after pregnancy) can inform the design of social media and professional websites for pregnancy education and support. This research aims to investigate the attention mechanism built-in or added on top of the BERT model in classifying and annotating the pregnancy-related questions posted on a community Q&A site. We evaluated two BERT-based models and compared them against the traditional machine learning models for question classification. Most importantly, we investigated two attention mechanisms: the built-in self-attention mechanism of BERT and the additional attention layer on top of BERT for relevant term annotation. The classification performance showed that the BERT-based models worked better than the traditional models, and BERT with an additional attention layer can achieve higher overall precision than the basic BERT model. The results also showed that both attention mechanisms work differently on annotating relevant content, and they could serve as feature selection methods for text mining in general. Xiao Luo 0002, Matthew Tang, Priyanka Gandhi, Zhan Zhang 0008, Zhe He 0001 |
BIBM | 1 |
| 2020 | Demonstrating the Impact of International Collaborative Disciplinary Experiences on Student Global, International, and Intercultural CompetenciesabstractThis Work in Progress research paper describes a study that directly compares the impact of a globally themed Information Technology (IT) project on students' global, international, and intercultural (GII) competencies. The authors will compare the change in student competencies by analyzing the impact of a common project with an international theme integrated into three different undergraduate IT classroom modalities: (1) a traditional classroom course with no interaction with students from a foreign university, (2) a virtual exchange context where teams of local students and students from a foreign university collaborate via information and communication technologies (ICT), and (3) a hands-on version of the course project that is implemented by local and remote students collaborating at a foreign university during a short-term study abroad program.This paper directly compares the changes in student GII competencies after executing the classroom project in the first and third modalities: a classroom without international collaboration and in conjunction with students at a foreign university during a short-term study abroad program. Preliminary results suggest that student competencies are more significantly improved when students collaborate with their international peers. However, the results might be influenced by the demographic profiles of students in the various courses. Students who pursue opportunities to study abroad may already have an innate expectation or desire to improve their GII competencies. Rob Elliott, Xiao Luo 0002 |
FIE | 2 |
| 2020 | Design and Evaluate the Factors for Flipped Classrooms for Data Management CoursesabstractThis Research to Practice Full Paper presents a framework to evaluate and design flipped classroom activities for data science and management courses. Variants of flipped classrooms have been employed in STEM fields with great success in students' learning outcomes. Research shows that flipped classrooms would improve students' learning if it is implemented following rigorous procedures of an efficient instructional design. As a result, one of the critical focus of current flipped classroom research is what factors educators need to consider when designing a flipped learning environment. Currently, educators incorporate various factors such as "pre-recorded video lecture", "group activity" as a trial and error basis and adjust these factors based on their own experience and students' feedback. On the other hand, the emergence of big data expects a new graduate to demonstrate mastery of concepts and skills for data acquisition, management, and analysis of inference from data when they enter the workforce. Currently, there is no systematic approach available to design a flipped classroom that is for the data science and management courses. In this research, we develop a framework first to investigate and evaluate the flipped classroom factors mentioned in the literature and identify a few that are most relevant to the two data management courses at our institute. Then, we classify each course topics into broader categories. So that the flipped classroom model can be developed for each category. For the flipped classroom for each category, we identify the pre-class and in-class activities to meet a certain learning objective for that topic category for each course. To evaluate the effectiveness of different factors as well as our flipped classroom models, students' performance data, interviews, and surveys are conducted. This process is transformative and can be employed by other STEM disciplines to find the most influential factors to design effective flipped learning classrooms. Shamima Mithun, Xiao Luo 0002 |
FIE | 2 |
| 2020 | BalNode2Vec: Balanced Random Walk based Versatile Feature Learning for NetworksabstractResearch on social networks and understanding the interactions of the users can be modeled as a task of graph mining, such as predicting nodes and edges in networks. The challenges of building the graph representation include engineering features used by learning algorithms. Recent research in representation learning can automate the prediction by learning the features themselves. Many types of research have performed graph sampling using random walks or it's derivatives. However, the random walk sometimes can not represent the features of the graph accurately enough. In this research, we propose BalNode2Vec - a new sampling algorithm for learning feature representations for nodes in networks by using balanced random walks. We define a notion of a nodes network neighborhood and design a balanced random walk procedure, which adapts to the graph topology. We show that through exploring the graph through a balanced random walk can generate richer representations. Efficacy of BalNode2vec over existing state-of-the-art techniques on link prediction is demonstrated by using several real-world networks from different domains. Amirreza Salamat, Xiao Luo 0002 |
IJCNN | 2 |
| 2019 | Integrated Education of Data Analytics and Information Security through Cross-Curricular ActivitiesabstractThe National Research Council's report states that cross-sectional studies of multiple courses within a discipline, or all courses in a major, would enhance the understanding of how people learn the concepts, practices, and ways of thinking of science and engineering and the nature and development of expertise in a discipline. In science and engineering, ever-evolving technology and information make integrative abilities necessary and especially valuable. In this study, we investigated cross-curricular pedagogy, by engaging undergraduate students of two disciplines in collaboration on a common, context-connected project, so that students are better prepared for solving interdisciplinary problems in career settings. We implemented cross-curricular pedagogy in a network security course and a big data analytics course. The era of big data enables data-driven malicious detection, and big data analytics techniques have been applied to analyzing network logs to reinforce information security and predict abnormal behaviors, so these domains overlap. We investigated two forms of cross-curricular activities: one was integrated instructional units, and the other was cross-curricular knowledge integration projects. The results show significant improvements in students confidence in solving cross-disciplinary problems and a much better understanding of data analytics and information security, as well as the connections between them. This project is the first to study the loose integration of two context-connected courses that are taught in parallel. Xiao Luo 0002, Connie Justice, Brandon Herald Sorge |
FIE | 1 |
| 2019 | Incorporate Cross-Course Knowledge Integration into Computing EducationabstractThis research to practice full paper describes our cross-course knowledge integration approach, which uses a project-based learning environment. The structure and instructional pedagogies of some of our data management concentration courses in our department have not changed in a long time. Most of our data management courses cover theoretical knowledge without showing their practical applications. As a result, students don't find those courses interesting enough. In addition, none of our data management courses require cross-course interactions, which is the current instructional pedagogical direction. With the goal of addressing these issues and improving student learning, faculty employed a cross-course knowledge integration model to teach the advanced database design course CIT 44400. We redesigned this advanced database course as part of our data management curriculum enhancement also to align with the trends of the IT industry and to enable students to correlate knowledge learned in various courses. In this redesign process, we integrated a project that aligns with industrial projects in this higher-level course curriculum, so that students can integrate and apply theoretical knowledge gained in lower-level courses through participating in a project-based learning approach, using current technology of the field. Evaluation results show crosscourse knowledge integration-based pedagogy gives students comprehensive knowledge that improves students' performance and helps them link and apply knowledge learned in various courses. Shamima Mithun, Xiao Luo 0002 |
FIE | 2 |
| 2018 | Incorporating Syntactic Dependencies into Semantic Word Vector Model for Medical Text Processing
Maia Iyer, Christopher Zou, Xiao Luo 0002 |
BIBM | 3 |
| 2017 | Extracting Modifiable Risk Factors from Narrative Preventive Healthcare Guidelines for EHR IntegrationabstractGeneral criteria of preventive healthcare based on the preventive care guidelines have been integrated with Electronic Health Record (EHR) systems through decision support systems and led to improved performance in healthcare delivery. Advanced integration which considers factors such as ethnicity, social history, medical history, family history need to be investigated. Integrating the preventive healthcare guidelines with the EHR based on above factors requires the extraction of the relevant information from these guidelines using text mining and natural language processing techniques. In this research, we propose a framework to extract information according to the EHR modules. Our results show that the proposed framework successfully extracts terms and concepts, and adequately maps them to the proposed data interchange structure that is based on the EHR functional modules. The extracted information and the populated data interchange structures eases the integration of the modifiable risk factors with the patient's records in the EHR. The proposed framework can be extended to other clinical healthcare guidelines where modifiable risk factors are critical. Setu Shah, Xiao Luo 0002 |
BIBE | 2 |
| 2017 | Exploring diseases based biomedical document clustering and visualization using self-organizing mapsabstractDocument clustering is a text mining technique used to provide better document search and browsing in digital libraries or online corpora. In this research, a vector representation of concepts of diseases and similarity measurement between concepts are proposed. They identify the closest concepts of diseases in the context of a corpus. Each document is represented by using the vector space model. A weight scheme is proposed to consider both local content and associations between concepts. Self-Organizing Maps (SOM) are often used as document clustering algorithm. The vector projection and visualization features of SOM enable visualization and analysis of the cluster distribution and relationships on the two dimensional space. The Davies-Bouldin index is used to validate the clusters based on the visualized cluster distributions. The results show that the proposed document clustering framework generates meaningful clusters and can facilitate clustering visualization and information retrieval based on the concepts of diseases. Setu Shah, Xiao Luo 0002 |
Healthcom | 2 |
| 2017 | Exploring a service-based normal behaviour profiling system for botnet detectionabstractEffective detection of botnet traffic becomes difficult as the attackers use encrypted payload and dynamically changing port numbers (protocols) to bypass signature based detection and deep packet inspection. In this paper, we build a normal profiling-based botnet detection system using three unsupervised learning algorithms on service-based flow-based data, including self-organizing map, local outlier, and k-NN outlier factors. Evaluations on publicly available botnet data sets show that the proposed system could reach up to 91% detection rate with a false alarm rate of 5%. Weikeng Chen, Xiao Luo 0002, Nur Zincir-Heywood |
IM | 2 |
| 2017 | Investigation of malicious portable executable file detection on the network using supervised learning techniquesabstractMalware continues to be a critical concern for everyone from home users to enterprises. Today, most devices are connected through networks to the Internet. Therefore, malicious code can easily and rapidly spread. The objective of this paper is to examine how malicious portable executable (PE) files can be detected on the network by utilizing machine learning algorithms. The efficiency and effectiveness of the network detection rely on the number of features and the learning algorithms. In this work, we examined 28 features extracted from metadata, packing, imported DLLs and functions of four different types of PE files for malware detection. The returned results showed that the proposed system can achieve 98.7% detection rates, 1.8% false positive rate, and with an average scanning speed of 0.5 seconds per file in our testing environment. Rushabh Vyas, Xiao Luo 0002, Nichole McFarland, Connie Justice |
IM | 2 |
| 2017 | Are Recent Terrorism Trends Reflected in Social Media?abstractSocial media plays an important role in shaping the beliefs and sentiments of an audience regarding an event. A comparison between public data sets that have holistic features and social media data set that include more user features would give insight into the spread of misinformation and aspects of events that are reflected in user behavior. In this research, we compare the trends identified in the public data set - Global Terrorism Database (GTD) with the trends reflected through the social media data obtained using the Twitter API. The unsupervised learning algorithm Self-Organizing Map (SOM) is used to identify the features and trends summarized by the clusters. The results show discrepancies in the features and related trends of terrorism events in the GTD data set and obtained Twitter data set to suggest some media bias and public perception on terrorism. Ivana Terziyska, Setu Shah, Xiao Luo 0002 |
MASS | 3 |
| 2015 | Predictive Analysis on Tracking Emails for Targeted Marketing
Xiao Luo 0002, Revanth Nadanasabapathy, Nur Zincir-Heywood, Keith Gallant, Janith Peduruge |
Discovery Science | 1 |
| 2006 | Evolving Recurrent Linear-GP for Document Classification and Word TrackingabstractIn this paper, we propose a novel document classification system where the recurrent linear Genetic Programming is employed to classify the documents that are represented in encoded word sequences. During this process, word sequences of documents are tracked, frequent patterns are detected and document is classified. We describe the word encoding model and the recurrent linear Genetic Programming based classification mechanism. The performance results on benchmark data set Reuters 21578 show that this system can analyze the temporal sequence patterns of a document and get competitive performance on classification. We expect that it can be easily applied to other application areas, where the temporal sequences are very significant. Xiao Luo 0002, Nur Zincir-Heywood |
IEEE Congress on Evolutionary Computation | 1 |
| 2005 | Evolving recurrent models using linear GPabstractTuring complete Genetic Programming (GP) models introduce the concept of internal state, and therefore have the capacity for identifying interesting temporal properties. Surprisingly, there is little evidence of the application of such models to problems for prediction. An empirical evaluation is made of a simple recurrent linear GP model over standard prediction problems. Xiao Luo 0002, Malcolm I. Heywood, Nur Zincir-Heywood |
GECCO | 1 |
| 2005 | Comparison of a SOM based sequence analysis system and naive Bayesian classifier for spam filteringabstractThe problem introduced by the unsolicited bulk emails, also known as "spam" generates a need for reliable anti-spam filters. In this paper, we design and compare the performance of a newly designed SOM based sequence analysis (SBSA) system for the spam filtering task. The system is based on a SOM based sequential data representation combined with a kNN classifier designed to make use of word sequence information. We compare this system with the traditional baseline method naive Bayesian filter. Three different cost scenarios and suitable cost-sensitive measurements are employed. The results show that the SBSA system is superior to the naive Bayesian filter, particularly when the misclassification cost for non-spam message is high. Xiao Luo 0002, Nur Zincir-Heywood |
IJCNN | 1 |
| 2005 | Evaluation of Two Systems on Multi-class Multi-label Document Classification
Xiao Luo 0002, Nur Zincir-Heywood |
ISMIS | 1 |
| 2004 | Analyzing the Temporal Sequences for Text Categorization
Xiao Luo 0002, Nur Zincir-Heywood |
KES | 1 |
| 2003 | A comparison of SOM based document categorization systemsabstractThis paper describes the development and evaluation of two unsupervised learning mechanisms for solving the automatic document categorization problem. Both mechanisms are based on a hierarchical structure of self-organizing feature maps. Specifically, one architecture is based on the vector space model whereas the other one is based on a code-books model. Results show that the latter architecture performs better than the first one which is based on the quality of the returned clusters. Xiao Luo 0002, Nur Zincir-Heywood |
IJCNN | 1 |