EDBT 2026 Demo / reviewers in the wild / expert
Keith S. Jones
dblp:09/1258
· DBLP profile ↗
6ranked-venue papers in the field
0as first author
3since 2021 · last 2025
0000-0002-3463-0401ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Emotion Detection in Imbalanced Conversational Data Using Transformer-Based Language Models
Bipsa Paka, Akbar Siami Namin, Faranak Abri, Keith S. Jones |
IEEE Big Data | 4 |
| 2023 | Detecting Phishing URLs using the BERT Transformer ModelabstractPhishing websites many a times look-alike to benign websites with the objective being to lure unsuspecting users to visit them. The visits at times may be driven through links in phishing emails, links from web pages as well as web search results. Although the precise motivations behind phishing websites may differ the common denominator lies in the fact that unsuspecting users are mostly required to take some action e.g., clicking on a desired Uniform Resource Locator (URL). To accurately identify phishing websites, the cybersecurity community has relied on a variety of approaches including blacklisting, heuristic techniques as well as content-based approaches among others. The identification techniques are every so often enhanced using an array of methods i.e., honeypots, features recognitions, manual reporting, web-crawlers among others. Nevertheless, a number of phishing websites still escape detection either because they are not blacklisted, are too recent or were incorrectly evaluated. It is therefore imperative to enhance solutions that could mitigate phishing websites threats. In this study, the effectiveness of the Bidirectional Encoder Representations from Transformers (BERT) is investigated as a possible tool for detecting phishing URLs. The experimental results detail that the BERT transformer model achieves acceptable prediction results without requiring advanced URLs feature selection techniques or the involvement of a domain specialist. Denish Omondi Otieno, Faranak Abri, Akbar Siami Namin, Keith S. Jones |
IEEE Big Data | 4 |
| 2022 | Using Transformers for Identification of Persuasion Principles in Phishing EmailsabstractIt is important to learn about attackers and their attacking strategies so that better and more effective defense systems can be built. During the reconnaissance stage, attackers intend to probe potential targets through various techniques including social engineering attacks. Phishing through email is a well-known, cheap, easy, and surprisingly effective technique for obtaining the needed information. This type of attack targets individuals and thus utilizes weaknesses that might exist in each person. Given the uniqueness of each individual’s personality, attackers make sure the right persuasion principle technique is employed for each targeted individual. This paper describes efforts to build machine-learning transformers, the emerging technique in language modeling, with the goal of building classifiers that take into account different types of persuasion principles. More specifically, the paper describes efforts to build machine-learning transformers based on BERT, RoBERTa, and DistilBERT and captures their classification results. The results show that these transformers are accurate enough to build a classification of phishing emails with respect to persuasion techniques. Furthermore, we report that the RoBERTa model is able to train faster than BERT and DistilBERT models. Bimal Karki, Faranak Abri, Akbar Siami Namin, Keith S. Jones |
IEEE Big Data | 4 |
| 2020 | Predicting Emotions Perceived from SoundsabstractSonification is the science of communication of data and events to users through sounds. Auditory icons, earcons, and speech are the common auditory display schemes utilized in sonification, or more specifically in the use of audio to convey information. Once the captured data are perceived, their meanings, and more importantly, intentions can be interpreted more easily and thus can be employed as a complement to visualization techniques. Through auditory perception it is possible to convey information related to temporal, spatial, or some other context-oriented information. An important research question is whether the emotions perceived from these auditory icons or earcons are predictable in order to build an automated sonification platform. This paper conducts an experiment through which several mainstream and conventional machine learning algorithms are developed to study the prediction of emotions perceived from sounds. To do so, the key features of sounds are captured and then are modeled using machine learning algorithms using feature reduction techniques. We observe that it is possible to predict perceived emotions with high accuracy. In particular, the regression based on Random Forest demonstrated its superiority compared to other machine learning algorithms. Faranak Abri, Luis Felipe Gutiérrez, Akbar Siami Namin, David R. W. Sears, Keith S. Jones |
IEEE BigData | 5 |
| 2020 | Predicting Consequences of Cyber-AttacksabstractCyber-physical systems posit a complex number of security challenges due to interconnection of heterogeneous devices having limited processing, communication, and power capabilities. Additionally, the conglomeration of both physical and cyber-space further makes it difficult to devise a single security plan spanning both these spaces. Cyber-security researchers are often overloaded with a variety of cyber-alerts on a daily basis many of which turn out to be false positives. In this paper, we use machine learning and natural language processing techniques to predict the consequences of cyberattacks. The idea is to enable security researchers to have tools at their disposal that makes it easier to communicate the attack consequences with various stakeholders who may have little to no cybersecurity expertise. Additionally, with the proposed approach researchers' cognitive load can be reduced by automatically predicting the consequences of attacks in case new attacks are discovered. We compare the performance through various machine learning models employing word vectors obtained using both tf-idf and Doc2Vec models. In our experiments, an accuracy of 60% was obtained using tf-idf features and 57% using Doc2Vec method for models based on LinearSVC model. Prerit Datta, Natalie R. Lodinger, Akbar Siami Namin, Keith S. Jones |
IEEE BigData | 4 |
| 2020 | Email Embeddings for Phishing DetectionabstractThe problem of detecting phishing emails through machine learning techniques has been discussed extensively in the literature. Conventional and state-of-the-art machine learning algorithms have demonstrated the possibility of building classifiers with high accuracy. The existing research studies treat phishing and genuine emails through general indicators and thus it is not exactly clear what phishing features are contributing to variations of the classifiers. In this paper, we crafted a set of phishing and legitimate emails with similar indicators in order to investigate whether these cues are captured or disregarded by email embeddings, i.e., vectorizations. We then fed machine learning classifiers with the carefully crafted emails to find out about the performance of email embeddings developed. Our results show that using these indicators, email embeddings techniques is effective for classifying emails as phishing or legitimate. Luis Felipe Gutiérrez, Faranak Abri, Miriam Armstrong, Akbar Siami Namin, Keith S. Jones |
IEEE BigData | 5 |