Sicong Shao

dblp:207/3541 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
19since 2021 · last 2027
0000-0003-1205-3890ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2027 A Green Hybrid Transformer-based System Log Anomaly Detection Framework for System Security
Qinxuan Shi, Sicong Shao, Zhanglong Yang, Brian Terry
Future Gener. Comput. Syst.2
2026 LLM-MC-Affect: LLM-Based Monte Carlo Modeling of Affective Trajectories and Latent Ambiguity for Interpersonal Dynamic Insight
abstract
Emotional coordination is a core property of human interaction that shapes how relational meaning is constructed in real time. While text-based affect inference has become increasingly feasible, prior approaches often treat sentiment as a deterministic point estimate for individual speakers, failing to capture the inherent subjectivity, latent ambiguity, and sequential coupling found in mutual exchanges. We introduce LLM-MC-Affect, a probabilistic framework that characterizes emotion not as a static label, but as a continuous latent probability distribution defined over an affective space. By leveraging stochastic LLM decoding and Monte Carlo estimation, the methodology approximates these distributions to derive high-fidelity sentiment trajectories that explicitly quantify both central affective tendencies and perceptual ambiguity. These trajectories enable a structured analysis of interpersonal coupling through sequential cross-correlation and slope-based indicators, identifying leading or lagging influences between interlocutors. To validate the interpretive capacity of this approach, we utilize teacher-student instructional dialogues as a representative case study, where our quantitative indicators successfully distill high-level interaction insights such as effective scaffolding. This work establishes a scalable and deployable pathway for understanding interpersonal dynamics, offering a generalizable solution that extends beyond education to broader social and behavioral research.
Yu-Zheng Lin, Bono Po-Jen Shih, John Paul Martin Encinas, Elizabeth Victoria Abraham Achom, Karan Himanshu Patel, Jesus Horacio Pacheco, Sicong Shao, Jyotikrishna Dass, Soheil Salehi, Pratik Satam
ACL (1)7
2026 Robust Ensemble Log Anomaly Detection for High-performance Parallel and Distributed Computing Systems
abstract
As high-performance parallel and distributed computing (HPDC) systems grow in scale, heterogeneity, and operational complexity, maintaining their security and reliability demands fast, accurate log anomaly detection (LAD) across massive volumes of system logs. Existing LAD methods often struggle with log volume, diversity, manual threshold tuning, and rare or context-dependent anomalies. To address these limitations, this research presents RE-LAD, a streaming LAD ensemble framework for HPDC systems. RE-LAD combines scalable log ingestion via Apache Kafka with a robust ensemble anomaly detector whose base detectors extend Transformer-based anomaly detection and automatically calibrate detection thresholds. RE-LAD features BAT (Bagging-style ensemble of enhanced Anomaly Transformers) as its core detection component to improve detection robustness. The experimental results on Hadoop Distributed File System (HDFS) and Blue Gene/L supercomputer (BGL) logs show that BAT outperforms leading methods.
Qinxuan Shi, Zhanglong Yang, Sicong Shao
HPDC3
2026 Efficient System Log Analysis via Quantized On-Device Anomaly Detection and Response
Qinxuan Shi, Zhanglong Yang, Sicong Shao
PAM3
2026 Robust Ensemble-based Log Anomaly Detection for Securing Mobile Systems
Qinxuan Shi, Zhanglong Yang, Sicong Shao
SECON3
2026 TransEdge: a lightweight transformer-based anomaly detection framework for edge devices
abstract
The rapid growth of the Internet of Things (IoT) and the increasing interconnectivity of industrial systems are driving a critical need for anomaly detection to be performed locally on edge devices. However, a significant gap exists between the computational power required by advanced deep learning models and the limited resources of edge hardware, often forcing a compromise between detection capability and on-device feasibility. To bridge this gap, we first introduce the EM-AT and its variants (EM-AT-bin), which enhance the Anomaly Transformer (AT) by integrating the Expectation–Maximization (EM) algorithm to enable fully automated, data-driven threshold determination. The EM-AT model with a Bayesian information criterion (BIC) estimator achieves the highest detection performance across four public datasets (i.e., SWaT, WADI, HDFS, and OpenStack), with $$F_{1}$$ -scores of 96.32%, 92.47%, 98.90%, and 99.61%, respectively. Building on the EM-AT and its variants, we present Q-EM-AT (Quantized-EM-AT) and its variants (Q-EM-AT-bin), edge-optimized and quantized variants that leverage mixed-precision quantization to substantially reduce computational and memory overhead while preserving detection accuracy. Finally, we propose TransEdge, a lightweight edge anomaly detection framework that adopts Q-EM-AT as its core detector to balance detection performance and computational resource consumption. Comprehensive experiments show that TransEdge significantly reduces resource consumption while maintaining competitive detection performance, achieving $$F_{1}$$ -scores of 96.18% on SWaT, 92.36% on WADI, 98.65% on HDFS, and 99.43% on OpenStack.
Qinxuan Shi, Zhanglong Yang, Sicong Shao
J. Intell. Inf. Syst.3
2025 A Moss Growth Optimization Approach to Robot Path Planning
abstract
Safe and efficient path planning remains a key challenge for mobile robots, especially in cluttered and complex environments. Many existing methods struggle to balance global exploration and local optimization, often leading to suboptimal paths or high computational costs. This paper presents a bio-inspired Moss Growth Optimization (MGO) algorithm integrated with a graph-based approach for robotic navigation. The environment is efficiently represented using MAKLINK Graph Theory (MGT), enabling rapid initial path generation. MGO enhances global exploration by thoroughly searching the solution space, while an implicit memory mechanism supports local refinement of promising paths. To further improve search efficiency, a Line-of-Sight (LoS) Reduction technique minimizes unnecessary exploration. Simulation results and comparative evaluations show that the proposed approach achieves faster convergence and shorter paths compared to state-of-the-art methods in high-dimensional and cluttered spaces. With its nature-inspired design and memory-based updates, the MGO framework offers a robust, efficient, and scalable solution for optimal robot path planning in real-world autonomous navigation scenarios.
Samuel Steen, Tingjun Lei, Chaomin Luo, Sicong Shao, Lin Gong
CEC4
2025 HPCLOG: A Transformer-Based Log Anomaly Detection Framework for Distributed Systems
abstract
Distributed systems, prevalent in safety-critical domains, produce vast logs containing essential runtime data. Log Anomaly Detection (LAD) is vital for maintaining the security and reliability of these systems, yet traditional methods falter amid growing data complexity and scale. While modern deep learning approaches such as LSTM excel at pattern recognition, they inadequately model the long-term dependencies found in log sequences. To address these challenges, we introduce HPCLog (High-throughput Processing Collection Log), a transformerbased LAD framework tailored for distributed systems. HPCLog is built upon EM-AT, a novel unsupervised anomaly detection model that extends the Anomaly Transformer. By integrating an Expectation-Maximization (EM) algorithm, EM-AT enables fully automated threshold determination. HPCLog seamlessly integrates this EM-AT model with Apache Kafka, creating a highthroughput pipeline for log collection and processing that enables accurate anomaly detection with low latency. Experimental results on two public datasets demonstrate that HPCLog surpasses leading methods, achieving$F_{1}$-scores of$\mathbf{9 8. 9 0 \%}$on HDFS and 99.61% on OpenStack.
Qinxuan Shi, Zhanglong Yang, Sicong Shao
ICPADS3
2025 DualATLog: Dual-mode Anomaly Transformer-based Log Anomaly Detection Framework for Securing Software Systems (S)
abstract
Modern software systems generate massive numbers of log messages, driving the need for automatic log analysis to detect anomalies.Besides, the growing adoption of edge computing highlights the importance of performing log anomaly detection (LAD) on edge devices.However, traditional methods struggle to handle complex and largescale log data effectively.While recent deep learning methods, although effective, are hindered by high computational demands, they are undesirable for power-constrained edge deployment.To solve these problems, we introduce Du-alATLog, a Dual-mode Anomaly Transformer-based framework for software system log anomaly detection across varying environments.DualATLog operates in two distinct modes-a standard mode optimized for environments with sufficient computational resources, and a power-saving mode designed explicitly for low-power edge devices.Experiments on two popular public datasets demonstrate the superior performance of the standard mode with an F 1score of 98.78% on HDFS and 99.56% on OpenStack, while the power-saving mode exhibits only a slight performance reduction.Power consumption measurements further validate our framework's ability to be seamlessly deployed in different environments.
Qinxuan Shi, Zhanglong Yang, Sicong Shao
SEKE3
2025 Energy-efficient Anomaly Detection for Securing Water Treatment and Distribution Systems
abstract
Cyber-Physical Systems (CPS) in water treatment and distribution are critical to smart city infrastructure.While CPS integration enhances operational efficiency, it also expands the attack surface, making robust anomaly detection crucial.Traditional deep learning approaches, despite their effectiveness in anomaly detection, are often too computationally intensive for resource-constrained edge devices prevalent in smart city deployments.To address this challenge, we introduce EATW (a novel Energy-efficient Anomaly Transformer-based anomaly detection architecture for Water treatment and distribution systems).EATW offers two modes: a performance mode for optimal detection capability where computational resources are ample, and an energy-efficient mode for low-resource edge devices, leveraging torchao technology and ExecuTorch framework to achieve crucial balance between accurate detection and operational efficiency.Evaluations using SWaT and WADI water system facility datasets demonstrate EATW's efficacy.The performance mode achieved superior results on the WADI and equaled the best results on SWaT compared with leading methods, while the energy-efficient mode curtails resource consumption with only a marginal 1-2% decrease in F 1 score.This highlights EATW's capability to deliver efficient and reliable anomaly detection, enhancing the water system facilities' security within smart city application ecosystems.
Zhanglong Yang, Qinxuan Shi, Sicong Shao
SEKE3
2023 Anomaly Behavior Analysis of Smart Water Treatment Facility Service: Design, Analysis, and Evaluation
abstract
The current trends toward the design and deployment of smart city services, including water services, improve quality, reliability and reduce operational costs. These advancements have led to the proliferation of ubiquitous connectivity to critical infrastructures. However, although smart sensors and Industrial Internet of Things (IIoTs) expedites rigorous monitoring and control, they exponentially increase vulnerabilities that can be exploited by cyberattacks. Therefore, development of advanced cybersecurity tools and resilience methods for smart city services are critically important because compromising these services can lead to disasters, accidents or even loss of life. To address the cybersecurity challenges facing smart city services, researchers need realistic testbeds to perform experiments, collect real-time data, and evaluate different security algorithms to protect smart critical infrastructure services. This paper presents a Water Treatment Facility Testbed (WTFT), a Cyber-Physical System (CPS) developed to enable experimentation with cybersecurity and resilient algorithms to deliver smart water services that can tolerate cyberattacks. Furthermore, an anomaly-based detection unit for water quality is implemented and our experimental results show a 96.8% F1-score, and a 98.3% accuracy with an attack detection latency under two seconds.
Ibrahim Almazyad, Sicong Shao, Salim Hariri, Hisham A. Kholidy
AICCSA2
2023 An Explainable Outlier Detection-based Data Cleaning Approach for Intrusion Detection
abstract
The effectiveness of machine learning (ML)-based intrusion detection systems (IDSs) for detecting widespread cyberattacks on critical infrastructure and government systems has been demonstrated in recent years. Nevertheless, with ML models becoming more complex, people can hardly understand their decisions. Further, most works on model explanations focus on analyzing the ML model itself. However, data cleaning is also vital in influencing the model’s detection behavior. On the other hand, data cleaning for ML-based IDSs is challenging because modern IDS datasets may contain outliers that affect the training stage. In this work, we propose an explainable data cleaning approach for intrusion detection, which can effectively perform explainable isolation forest-based outlier detection in the data preprocessing stage for intrusion detection. Through experiments on real-world network intrusion datasets, we evaluate the effectiveness of our approach. Experiment results demonstrate that eliminating outliers improves intrusion detection and that data cleaning using outlier detection is explainable.
Theodore Ha, Sicong Shao, Salim Hariri
AICCSA2
2023 Resilient Machine Learning (rML) Against Adversarial Attacks on Industrial Control Systems
abstract
Machine learning (ML) algorithms have been widely used in many critical automated systems, including as a technique in Dynamic Data Driven Applications Systems (DDDAS)-based methods and areas such as financial trading, autonomous vehicles, and intrusion detection systems. However, malicious adversaries have strong interests in manipulating the operations of machine learning algorithms to achieve their objectives of gaining financial, social, or political influence. Adversarial ML (AML) users can be classified based on their capabilities and goals into three types: Adversary who has full knowledge of the ML models and parameters (white-box scenario), partial knowledge of ML models (gray-box scenario), and one who does not have any knowledge and uses guessing techniques to figure out the ML model and its parameters (black-box scenario). In these scenarios, the adversaries attempt to maliciously manipulate the model/data either during training or testing. Defending against these AML attacks can be successful by following methods such as making the ML model robust, validating and verifying inputs and outputs, and changing the ML architecture. This paper presents a resilient machine learning (rML) against adversarial attacks by dynamically conducting feature space anonymization and model randomization in ML services such that the adversaries lack knowledge about the feature space and model used and consequently prevent them from maliciously manipulating the ML operations during the runtime. In our approach, the rML utilizes autoencoders as an anonymization technique for encoding feature space to minimize the effect of adversarial samples. The rML method is evaluated using the benchmarking Industrial Control Systems (ICS) data and the corresponding adversarial data generated using the Jacobian-based Saliency Map Attack (JSMA) method. The experiment demonstrated that the proposed approach can detect attacks targeting ICS and prevent adversarial attacks compromising ML models used to secure ICS.
Likai Yao, Sicong Shao, Salim Hariri
AICCSA2
2023 Machine Learning for Intrusion Detection: Stream Classification Guided by Clustering for Sustainable Security in IoT
abstract
The Internet of Things (IoT) has brought about unprecedented connectivity and convenience in our daily lives, but with this newfound interconnectedness comes the threat of cyber-attacks. With ever-increasing IoT devices being connected to the internet, securing IoT devices is becoming increasingly urgent. Machine learning (ML) is among the most popular techniques used by intrusion detection systems (IDS) to enhance their detection performance when securing IoT. However, a key obstacle of ML-based IDS for IoT is learning from nonstationary streaming data, also known as concept drift. One of the most challenging learning scenarios under concept drift is extreme verification latency (EVL), which occurs when only unlabeled nonstationary streaming data is available after a small set of initial labeled data. Stream Classification Algorithm Guided by Clustering (SCARGC) is an algorithm that can effectively deal with the nonstationary data streams in EVL scenarios. Applying an EVL implementation provides the capability of adapting to nonstationary environments within the IoT domain. The SCARGC model, as an integrated IoT intrusion detection system, allows for sustainable security as new threats are identified in this non-stationary environment. Hence, in this project, we develop an innovative IoT intrusion detection approach by natively integrating SCARGC and intrusion detection to address the EVL challenges to provide sustainable security as the model adapts to nonstationary environments. We evaluated the proposed approach on real-world IoT cybersecurity datasets. The results demonstrate the feasibility of the proposed approach, which can lead to the development of sophisticated intrusion detection systems for IoT.
Martin Manuel Lopez, Sicong Shao, Salim Hariri, Soheil Salehi
ACM Great Lakes Symposium on VLSI2
2023 Quantized Transformer Language Model Implementations on Edge Devices
abstract
Large-scale transformer-based models like the Bidi-rectional Encoder Representations from Transformers (BERT) are widely used for Natural Language Processing (NLP) applications, wherein these models are initially pre-trained with a large corpus with millions of parameters and then fine-tuned for a downstream NLP task. One of the major limitations of these large-scale models is that they cannot be deployed on resource- constrained devices due to their large model size and increased inference latency. In order to overcome these limitations, such large-scale models can be converted to an optimized FlatBuffer format, tailored for deployment on resource-constrained edge devices. Herein, we evaluate the performance of such FlatBuffer transformed MobileBERT models on three different edge devices, fine-tuned for Reputation analysis of English language tweets in the Rep Lab 2013 dataset. In addition, this study encompassed an evaluation of the deployed models, wherein their latency, performance, and resource efficiency were meticulously assessed. Our experiment results show that, compared to the original BERT large model, the converted and quantized MobileBERT models have 160x smaller footprints for a 4.1 % drop in accuracy while analyzing at least one tweet per second on edge devices. Furthermore, our study highlights the privacy-preserving aspect of TinyML systems as all data is processed locally within a serverless environment.
Mohammad Wali Ur Rahman, Murad Mehrab Abrar, Hunter Gibbons Copening, Salim Hariri, Sicong Shao, Pratik Satam, Soheil Salehi
ICMLA5
2022 Blockchain Based Methodology for Zero Trust Modeling and Quantification for 5G Networks
abstract
The 5th generation mobile network (5G) is designed with a new core architecture that makes it quite extensible. The components of the 5G core architecture are no longer physical standalone devices, but rather software processes run on commercial off-the-shelf (COTS) servers. The backbone of 5G is software-defined networking (SDN) and network function virtualization (NFV), and they both bring unprecedented flexibility to network and resource management. In this context, 5G logical networks can be created by partitioning a shared physical infrastructure, and each network can be customized and optimized for specific entity. This concept is known as 5G network slicing. Despite the tremendous benefits of network slicing, it also brings many unprecedented security challenges because of the dynamism and diversity of slice's structure. Therefore, establishing trust in the 5G ecosystem is a cornerstone for global adaptation and tackling security and privacy risks. In this paper, we focus on the trust aspect between the network slice stakeholders (i.e slice owners, users, slice resource providers, and service providers), and we propose a blockchain-based zero trust model that addresses threat models that are based on the lack of trust between the entities in a network slice. Our approach for zero trust modeling and quantification is based on direct evidence and indirect evidence and the use of smart contracts with blockchain to maintain the required trust values at runtime. We provide details on how to model and quantify the trust of all the stakeholders of a given network slice and how the blockchain smart contract can enforce the zero-trust requirements for all network slice stakeholders.
Safwan Elmadani, Salim Hariri, Sicong Shao
AICCSA3
2022 A BERT-based Deep Learning Approach for Reputation Analysis in Social Media
abstract
Social media has become an essential part of the modern lifestyle, with its usage being highly prevalent. This has resulted in unprecedented amounts of data generated from users in social media, such as users' attitudes, opinions, interests, purchases, and activities across various aspects of their lives. Therefore, in a world of social media, where its power has shifted to users, actions taken by companies and public figures are subject to constantly being under scrutiny by influential global audiences. As a result, reputation management in social media has become essential as companies and public figures need to maintain their reputation to preserve their reputational capital. However, domain experts still face the challenge of lacking appropriate solutions to automate reliable online reputation analysis. To tackle this challenge, we proposed a novel reputation analysis approach based on the popular language model BERT (Bidirectional Encoder Representations from Transformers). The proposed approach was evaluated on the reputational polarity task using RepLab 2013 dataset. Compared to previous works, we achieved 5.8% improvement in accuracy, 26.9% improvement in balanced accuracy, and 21.8% improvement in terms of F-score.
Mohammad Wali Ur Rahman, Sicong Shao, Pratik Satam, Salim Hariri, Chris Padilla, Zoe Taylor, Carlos Nevarez
AICCSA2
2022 AI-based Arabic Language and Speech Tutor
abstract
In the past decade, we have observed a growing interest in using technologies such as artificial intelligence (AI), machine learning, and chatbots to provide assistance to language learners, especially in second language learning. By using AI and natural language processing (NLP) and chatbots, we can create an intelligent self-learning environment that goes beyond multiple-choice questions and/or fill in the blank exercises. In addition, NLP allows for learning to be adaptive in that it offers more than an indication that an error has occurred. It also provides a description of the error, uses linguistic analysis to isolate the source of the error, and then suggests additional drills to achieve optimal individualized learning outcomes. In this paper, we present our approach for developing an Artificial Intelligence-based Arabic Language and Speech Tutor (AI-ALST) for teaching the Moroccan Arabic dialect. The AI-ALST system is an intelligent tutor that provides analysis and assessment of students learning the Moroccan dialect at University of Arizona (UA). The AI-ALST provides a self-learned environment to practice each lesson for pronunciation training. In this paper, we present our initial experimental evaluation of the AI-ALST that is based on MFCC (Mel frequency cepstrum coefficient) feature extraction, bidirectional LSTM (Long Short-Term Memory), attention mechanism, and a cost-based strategy for dealing with class-imbalance learning. We evaluated our tutor on the word pronunciation of lesson 1 of the Moroccan Arabic dialect class. The experimental results show that the AI-ALST can effectively and successfully detect pronunciation errors and evaluate its performance by using$\boldsymbol{F}_{\mathbf{1}}$- score, accuracy, precision, and recall.
Sicong Shao, Saleem Alharir, Salim Hariri, Pratik Satam, Sonia Shiri, Abdessamad Mbarki
AICCSA1
2021 Multi-Layer Mapping of Cyberspace for Intrusion Detection
abstract
The ubiquity and vulnerability of computer applications make them ideal places for intrusion attacks that increase in intensity and complexity. Computer applications have a relationship with various networks, physical components, host devices, and users with different roles and requirements. Therefore, securing computer applications in such a complex and dynamic cyberspace is urgent and challenging. This paper attempts to tackle the challenges by proposing a Multi-Layer Abnormal Behaviors Analysis (MLABA) framework for intrusion detection associated with three layers (i.e., system, process, and network layers) in cyberspace for characterizing their normal operations and detect any abnormal behavior that might be triggered by malicious activities. The proposed technique was evaluated on several popular applications (i.e., Firefox, Opera, Chrome, and Ruby). The experimental results demonstrate the feasibility of MLABA framework that can detect the intrusion and abuse for applications.
Sicong Shao, Pratik Satam, Shalaka Satam, Khalid Al-Awady, Gregory Ditzler, Salim Hariri, Cihan Tunc
AICCSA1
2020 Video Anomaly Detection using Pre-Trained Deep Convolutional Neural Nets and Context Mining
abstract
Anomaly detection is critically important for intelligent surveillance systems to detect in a timely manner any malicious activities. Many video anomaly detection approaches using deep learning methods focus on a single camera video stream with a fixed scenario. These deep learning methods use large-scale training data with large complexity. As a solution, in this paper, we show how to use pre-trained convolutional neural net models to perform feature extraction and context mining, and then use denoising autoencoder with relatively low model complexity to provide efficient and accurate surveillance anomaly detection, which can be useful for the resource-constrained devices such as edge devices of the Internet of Things (IoT). Our anomaly detection model makes decisions based on the high-level features derived from the selected embedded computer vision models such as object classification and object detection. Additionally, we derive contextual properties from the high-level features to further improve the performance of our video anomaly detection method. We use two UCSD datasets to demonstrate that our approach with relatively low model complexity can achieve comparable performance compared to the state-of-the-art approaches.
Chongke Wu, Sicong Shao, Cihan Tunc, Salim Hariri
AICCSA2
2019 One-Class Classification with Deep Autoencoder Neural Networks for Author Verification in Internet Relay Chat
abstract
Social networks are highly preferred to express opinions, share information, and communicate with others on arbitrary topics. However, the downside is that many cybercriminals are leveraging social networks for cyber-crime. Internet Relay Chat (IRC) is the important social networks which can grant the anonymity to users by allowing them to connect channels without sign-up process. Therefore, IRC has been the playground of hackers and anonymous users for various operations such as hacking, cracking, and carding. Hence, it is urgent to study effective methods which can identify the authors behind the IRC messages. In this paper, we design an autonomic IRC monitoring system, performing recursive deep learning for classifying threat levels of messages and develop a novel author verification approach with one-class classification with deep autoencoder neural networks. The experimental results show that our approach can successfully perform effective author verification for IRC users.
Sicong Shao, Cihan Tunc, Amany Al-Shawi, Salim Hariri
AICCSA1
2019 Automated Twitter Author Clustering with Unsupervised Learning for Social Media Forensics
abstract
Twitter is one of the key social media platforms, which is also used for cyber-crimes. Hence, monitoring and detecting the malicious activities of Twitter users is critically important for cybersecurity concerns around the globe since cybercriminals are heavily using Twitter for illegal purpose. It is increasingly common for cybercriminals signing up many accounts while masquerading different users for malicious behaviors. This fact has brought forward the issue of identifying the authors of Twitter accounts. In this paper, we propose a novel approach through a combination of feature extraction methods and then convert high dimensional data to kernel matrix for Twitter author clustering. The experimental results show that our approach can be used to effectively identify the groups among more than one hundred Twitter aliases even without knowing the number of authors.
Sicong Shao, Cihan Tunc, Amany Al-Shawi, Salim Hariri
AICCSA1
2018 Autonomic Author Identification in Internet Relay Chat (IRC)
abstract
With the advances in Internet technologies and services, the social media has been gaining excessive popularity, especially because these technologies provide anonymity where they use nicknames to post their messages. Unfortunately, the anonymity feature has been exploited by the cyber-criminals to hide their identities and their operations. Hence, there is a growing interest in cybersecurity research domain to identify the authors of malicious messages and activities. Internet Relay Chat (IRC) channels are widely used to exchange messages and information among malicious users involved in cybercrimes. In this paper, we present an autonomic author identification technique based on personality profile and analysis of IRC messages. We first monitor the IRC channels using our autonomic bots and then create a personality profile for each targeted author. We demonstrate that personality analysis for author detection/identification is an efficient approach and has high detection rates.
Sicong Shao, Cihan Tunc, Amany Al-Shawi, Salim Hariri
AICCSA1