EDBT 2026 Demo / reviewers in the wild / expert
Dan Chia-Tien Lo
dblp:65/6293 · also C. Dan Lo, Chia-Tien Dan Lo, Dan C. Lo, Dan Lo 0001
· DBLP profile ↗
16ranked-venue papers in the field
5as first author
8since 2021 · last 2024
0000-0003-4396-0370ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 16 (5 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Enhancing Contextual Understanding in Knowledge Graphs: Integration of Quantum Natural Language Processing with Neo4j LLM Knowledge GraphabstractTraditional Knowledge Graphs (KGs), such as Neo4j, face challenges in managing high-dimensional relationships and capturing semantic nuances due to their deterministic nature. Quantum Natural Language Processing (QNLP) introduces probabilistic reasoning into the KG context. This integration leverages quantum principles, such as superposition, which allows relationships to exist in multiple states simultaneously, and entanglement, where the state of one entity dynamically influences the state of another. This quantum-based probabilistic reasoning provides a richer, more flexible representation of connections, moving beyond binary relationships to model the nuances and variability of real-world interactions. Our research demonstrates that QNLP enhances Neo4j’s ability to analyze context-rich data, improving tasks like entity extraction and knowledge inference. By modeling relationship states probabilistically, QNLP addresses limitations in traditional methods, providing nuanced insights and enabling more advanced, context-aware NLP applications. Suman Bharti, Dan Chia-Tien Lo, Yong Shi 0002 |
IEEE Big Data | 2 |
| 2024 | Practical Considerations of Fully Homomorphic Encryption in Privacy-Preserving Machine LearningabstractMachine learning has been successfully applied to big data analytics across various disciplines. However, as data is collected from diverse sectors, much of it is private and confidential. At the same time, one of the major challenges in machine learning is the slow training speed of large models, which often requires high-performance servers or cloud services. To protect data privacy while still allowing model training on such servers, privacy-preserving machine learning using Fully Homomorphic Encryption (FHE) has gained significant attention. However, its widespread adoption is hindered by performance degradation. This paper presents our experiments on training models over encrypted data using FHE. The results show that while FHE ensures privacy, it can significantly degrade performance, requiring complex tuning to optimize. Dan Chia-Tien Lo, Yong Shi 0002, Hossain Shahriar, Bobin Deng, Xinyue Zhang 0001, Mei-Lan Chen |
IEEE Big Data | 1 |
| 2024 | Towards More Robust and Scalable Deep Learning Systems for Medical Image AnalysisabstractDeep learning (DL) has attracted interest in healthcare for disease diagnosis systems in medical imaging analysis (MedIA) and is especially applicable in Big Data environments like federated learning (FL) and edge computing. However, there is little research into mitigating the vulnerabilities and robustness of such systems against adversarial attacks, which can force DL models to misclassify, leading to concerns about diagnosis accuracy. This paper aims to evaluate the robustness and scalability of DL models for MedIA applications against adversarial attacks while ensuring their applicability in FL settings with Big Data. We fine-tune three state-of-the-art transfer learning models, DenseNet121, MobileNet-V2, and ResNet50, on several MedIA datasets of varying sizes and show that they are effective at disease diagnosis. We then apply the Fast Gradient Sign Method (FGSM) to attack the models and utilize adversarial training (AT) and knowledge distillation to defend them. We provide a performance comparison of the original transfer learning models and the defended models on the clean and perturbed data. The experimental results show that the defensive techniques can improve the robustness of the models to the FGSM attack and be scaled for Big Data as well as utilized for edge computing environments. Akshaj Yenumala, Xinyue Zhang 0001, Dan Chia-Tien Lo |
IEEE Big Data | 3 |
| 2023 | Deep Machine Learning on Segmenting and Classifying Crop Images Taken by Unmanned Aerial VehicleabstractIn the realm of precision agriculture, a crucial element involves the precise quantification or estimation of seedlings, fruits, and other agricultural produce on expansive multi-acre farms at various stages of cultivation. With the advent of unmanned aerial vehicles (UAVs), capturing images of watermelon fields has become a straightforward task. These images can be subsequently processed, segmented, and categorized to determine the total count of watermelons. Currently, conventional methods are employed to address this challenge, but they have their limitations. The field has benefited from the evolution of machine learning, which has the potential to streamline the process. Nevertheless, the training phase is intricate, and achieving a valuable model can be demanding. This research delves into an examination and presentation of the existing pre-trained models for image processing in this context. Dan Chia-Tien Lo, Bobin Deng, Yong Shi 0002 |
IEEE Big Data | 1 |
| 2023 | Design and Implementation of an ERC-20 Smart Contract on the Ethereum BlockchainabstractAs technology continues to develop, there is a growing need to find sustainable solutions in all industries, including cryptocurrency. Due to the high energy consumption that cryptocurrencies are known for, there have been efforts to reduce waste consumption and in turn minimize the carbon footprint. We support the trends for creating an environment-friendly crypto token using the ERC-20 standard on the Ethereum blockchain. We outline the various aspects that make a token more sustainable and highlight the potential benefits of such tokens. Our proposal involves the design of a smart contract that incorporates eco-friendly features such as lower energy consumption, carbon offsetting, and more efficient methods or algorithms. We also discuss the importance of transparency and accountability in the design and implementation of such tokens. This paper discusses not only the practical tools and steps necessary in creating a crypto token but also highlights the challenges associated with creating a more sustainable token. Joshua Priest, Cameron Cooper, Savvy Lovell, Yong Shi 0002, Dan Chia-Tien Lo |
IEEE Big Data | 5 |
| 2022 | Authentic Learning of Machine Learning in Cybersecurity with Portable Hands-on Labware: Neural Network Algorithms for Network Denial of Service (DOS) DetectionabstractThe primary goal of the authentic learning approach is to engage and motivate students in a learning environment that encourages all students in learning. This approach provides students with hands-on experiences in solving real-world security problems. We designed and developed ten learning modules based on 10 cybersecurity cases with different ML solutions. Each learning module consists of pre-lab, lab, and post-lab (Pre/Lab/Post) activities. All portable labs are made available on Google CoLab for ML to cybersecurity so that students can access and practice these hands-on labs anywhere and anytime without time tedious installation and configuration which will engage students in learning concepts and getting more experience for hands-on problem-solving skills. In this paper, we adopt Neural Network Algorithms for Network Denial of Service (DOS) Detection where we apply the KDDCup 1999 datasets contain a standard set of data to be audited, which includes a wide variety of intrusions simulated in a military network environment. Our primary goal of this lab is to show whether a link is a malicious or safe connection. Our demonstration shows an achieved accuracy of 99.89%. Md. Jobair Hossain Faruk, Hossain Shahriar, Dan Chia-Tien Lo, Michael E. Whitman, Alfredo Cuzzocrea, Fan Wu 0013, Victor Clincy |
IEEE Big Data | 4 |
| 2021 | Malware Detection and Prevention using Artificial Intelligence TechniquesabstractWith the rapid technological advancement, security has become a major issue due to the increase in malware activity that poses a serious threat to the security and safety of both computer systems and stakeholders. To maintain stakeholder’s, particularly, end user’s security, protecting the data from fraudulent efforts is one of the most pressing concerns. A set of malicious programming code, scripts, active content, or intrusive software that is designed to destroy intended computer systems and programs or mobile and web applications is referred to as malware. According to a study, naive users are unable to distinguish between malicious and benign applications. Thus, computer systems and mobile applications should be designed to detect malicious activities towards protecting the stakeholders. A number of algorithms are available to detect malware activities by utilizing novel concepts including Artificial Intelligence, Machine Learning, and Deep Learning. In this study, we emphasize Artificial Intelligence (AI) based techniques for detecting and preventing malware activity. We present a detailed review of current malware detection technologies, their shortcomings, and ways to improve efficiency. Our study shows that adopting futuristic approaches for the development of malware detection applications shall provide significant advantages. The comprehension of this synthesis shall help researchers for further research on malware detection and prevention using AI. Md. Jobair Hossain Faruk, Hossain Shahriar, Maria Valero, Farhat Lamia Barsha, Shahriar Sobhan, Md Abdullah Khan, Michael E. Whitman, Alfredo Cuzzocrea, Dan Chia-Tien Lo, Akond Ashfaque Ur Rahman, Fan Wu 0013 |
IEEE BigData | 9 |
| 2021 | Colab Cloud Based Portable and Shareable Hands-on Labware for Machine Learning to CybersecurityabstractMachine Learning (ML) analyze, and process data and develop patterns. In the case of cybersecurity, it helps to better analyze previous cyber attacks and develop proactive strategy to detect, prevent the security threats. Both ML and cybersecurity are important subjects in computing curriculum but ML for security is not well presented there. We design and develop case-study based portable labware on Google CoLab for ML to cybersecurity so that students can access, share, collaborate, and practice these hands-on labs anywhere and anytime without time tedious installation and configuration which will help students more focus on learning of concepts and getting more experience for hands-on problem solving skills. Dan Chia-Tien Lo, Hossain Shahriar, Michael E. Whitman, Fan Wu 0013 |
IEEE BigData | 1 |
| 2019 | Subject-Oriented Data Retrieval and Analysis on Sina WeiboabstractThe goal of this project is to timely identify and characterize hot topics discussed on the Chinese microblogging website, Sina Weibo. First, we develop a subject-oriented crawler to retrieve Weibo posts. Second, a probabilistic topic model is applied to the dataset to identify prominent topics and keywords. A comparison on hot topics, Word Cloud, and Latent Dirichlet Allocation results are discussed, followed by future research directions. Dan Chia-Tien Lo, Charles Garnder, Pascal Paschos, Chung Ng |
IEEE BigData | 1 |
| 2019 | Experiential Learning: Case Study-Based Portable Hands-on Regression Labware for Cyber Fraud PredictionabstractMachine Learning (ML) analyzes, and processes data and discover patterns. In cybersecurity, it effectively analyzes big data from existing cybersecurity attacks and develop proactive strategies to detect current and future cybersecurity attacks. Both ML and cybersecurity are important subjects in computing curriculum, but using ML for cybersecurity is not commonly explored. This paper designs and presents a case study-based portable labware experience built on Google's CoLaboratory (CoLab) for a ML cybersecurity application to provide students with hands-on labs accessing from anywhere and anytime, reducing or eliminating tedious installations and configurations. This approach allows students to focus on learning essential concepts and gaining valuable experience through hands-on problem solving skills. Our preliminary results and student evaluations are reported for a case-based hands-on regression labware in cyber fraud prediction using credit card fraud as an example. Hossain Shahriar, Michael E. Whitman, Dan Chia-Tien Lo, Fan Wu 0013, Cassandra Thomas, Alfredo Cuzzocrea |
IEEE BigData | 3 |
| 2019 | Big Data Analysis on Social NetworkingabstractSocial networking media, such as Twitters, Facebook, and Chinese Weibo, has become a major means for people to deliberately express their ideas, thoughts, and views about everything. A huge amount of posts in these social media are issued and viewed by the general public daily. Evidently the social networking media can directly affect people's perceptions on a specific topic. Those data can be used to valuable information that will help organizations to understand what thetrends or sentiments are. Most of the research efforts in social networking data analytics are conducted in English-based social networking media. Research on Chinese social networking media receives relatively little attention. In this paper, we examine the key problems in this field, focus particularly on the characteristics of new vocabulary, emotion expressions, and hierarchical structures in Chinese “Weibo”. Associated theoretical and technological methods to address these problems are reviewed and discussed. Zhengwu Sun, Dan Chia-Tien Lo, Yong Shi 0002 |
IEEE BigData | 2 |
| 2017 | Automated microsoft office macro malware detection using machine learningabstractMacro malware in Microsoft (MS) Office files has long persisted as a cybersecurity threat. Though it ebbed after its initial rampages around the turn of the century, it has reemerged as threat. Attackers are taking a persuasive approach and using document engineering, aided by improved data mining methods, to make MS Office file malware appear legitimate. Recent attacks have targeted specific corporations with malicious documents containing unusually relevant information. This development undermines the ability of users to distinguish between malicious and legitimate MS Office files and intensifies the need for automating macro malware detection. This study proposes a method of classifying MS Office files containing macros as malicious or benign using the K-Nearest Neighbors machine learning algorithm, feature selection, and TFIDF where p-code opcode n-grams (translated VBA macro code) compose the file features. This study achieves a 96.3% file classification accuracy on a sample set of 40 malicious and 118 benign MS Office files containing macros, and it demonstrates the effectiveness of this approach as a potential defense against macro malware. Finally, it discusses the challenges automated macro malware detection faces and possible solutions. Ruth Bearden, Dan Chia-Tien Lo |
IEEE BigData | 2 |
| 2017 | Training on the poles for review sentiment polarity classificationabstractSupport Vector Machines (SVMs) are well-known tools for the task of big data text classification. This research studies the effects of omitting non-polar training samples on the performance of a SVM-based binary text classifier. The classifier operates on a large corpus of Amazon product review text bodies, sampled from various product categories and predicts the polarity of reviews. Our results show that training on a smaller, more concise training set of only 1 and 5-star reviews offers similar performance to training on a much larger dataset that includes 1, 2, 4, and 5-star reviews, without sacrificing much in terms of precision or recall. These results hold true using both count-based and TF-IDF vectorization methodologies. The technique of training on the poles is demonstrated to offer efficient means to build a high-performing SVM-based review sentiment polarity classifier, especially in cases where labeled data may not be readily available and training time is constrained. Michael Kranzlein, Dan Chia-Tien Lo |
IEEE BigData | 2 |
| 2017 | Binary malware image classification using machine learning with local binary patternabstractMalware classification is a critical part in the cyber-security. Traditional methodologies for the malware classification typically use static analysis and dynamic analysis to identify malware. In this paper, a malware classification methodology based on its binary image and extracting local binary pattern (LBP) features is proposed. First, malware images are reorganized into 3 by 3 grids which is mainly used to extract LBP feature. Second, the LBP is implemented on the malware images to extract features in that it is useful in pattern or texture classification. Finally, Tensorflow, a library for machine learning, is applied to classify malware images with the LBP feature. Performance comparison results among different classifiers with different image descriptors such as GIST, a spatial envelop, and the LBP demonstrate that our proposed approach outperforms others. Jhu-Sin Luo, Dan Chia-Tien Lo |
IEEE BigData | 2 |
| 2016 | Towards an effective and efficient malware detection systemabstractThe ubiquitous advance of technology used on the Internet, computers, smart phones and tablets has been conducive to the creation and proliferation of cyber threats resulting in attacks that have grown exponentially. Consequently, anti-virus companies and researchers have developed new approaches for dealing with discovering and classifying malware. Among these, machine learning and big data technologies have been used for feature extraction, detection, and clustering of cyber threats. In this paper a dataset of malware and clean files (goodware) was created and analyzed from the static and dynamic features provided by the online framework VirusTotal. The purpose is to select the smallest number of features that keep classification accuracy as high as possible in order to decrease the use of resources for monitoring as well as extracting features and the time for detection. In this research, it was found that “9” features are enough to distinguish malware from “goodware” files with an accuracy of 99.60%. Selecting the most representative features for malware detection relies on the possibility of creating an embedded program that monitors the processes executed by the operating system (OS) and looks for the characteristics that match malware behavior. In addition, classification algorithms such as Random Forest (RF), Support Vector Machine (SVM) and Neural Networks (NN) were used in a novel combination that not only showed an increase in accuracy, but also in the training speed from hours to just minutes. Finally, the trained model (which was trained with a dataset of malware samples seen before September 2015) was tested on a new dataset of malware samples seen by first time between October 2015 and June 2016 and showed that the model is still effective for detection of unseen malware files. Dan Chia-Tien Lo, Pablo Ordóñez, Carlos Cepeda Mora |
IEEE BigData | 1 |
| 2007 | Compression for Low Power Consumption in Battery-powered HandsetsabstractJava has been introduced to the mobile/wireless handsets and Java enabled handsets are now prevalent in the market, showing their successful launch. A variety of attractive services have been developed and promise a lot of fun to our daily lives, which are however targeted to the high end products that are very expensive. We focus on economic types with similar performance to the high end products. Memory compression is a key technique to achieve the goal and is integrated into a Java runtime environment. Mayumi Kato, Dan Chia-Tien Lo |
DCC | 2 |