Murat Simsek

dblp:29/10350 · DBLP profile ↗
← Back
40ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0003-3156-5760ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 30 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Revisiting Table Detection Datasets for Visually Rich Documents
Bin Xiao 0008, Murat Simsek, Burak Kantarci, Ala Abu Alkheir
Int. J. Document Anal. Recognit.2
2025 Rethinking detection based table structure recognition for visually rich document images
Bin Xiao 0008, Murat Simsek, Burak Kantarci, Ala Abu Alkheir
Expert Syst. Appl.2
2025 On-Dyn-CDA: A Real-Time Cost-Driven Task Offloading Algorithm for Vehicular Networks With Reduced Latency and Task Loss
abstract
Real-time task processing is a critical challenge in vehicular networks, where achieving low latency and minimizing dropped task ratio depend on efficient task execution. Our primary objective is to maximize the number of completed tasks while minimizing overall latency, with a particular focus on reducing number of dropped tasks. To this end, we investigate both static and dynamic versions of an optimization algorithm. The static version assumes full task availability, while the dynamic version manages tasks as they arrive. We also distinguish between online and offline cases: the online version incorporates execution time into the offloading decision process, whereas the offline version excludes it, serving as a theoretical benchmark for optimal performance. We evaluate our proposed Online Dynamic Cost-Driven Algorithm (On-Dyn-CDA) against these baselines. Notably, the static Particle Swarm Optimization (PSO) baseline assumes all tasks are transferred to the RSU and processed by the MEC, and its offline version disregards execution time, making it infeasible for real-time applications despite its optimal performance in theory. Our novel On-Dyn-CDA completes execution in just 0.05 seconds under the most complex scenario, compared to 1330.05 seconds required by Dynamic PSO. It also outperforms Dynamic PSO by 3.42% in task loss and achieves a 29.22% reduction in average latency in complex scenarios. Furthermore, it requires neither a dataset nor a training phase, and its low computational complexity ensures efficiency and scalability in dynamic environments.
Mahsa Paknejad, Parisa Fard Moshiri, Murat Simsek, Burak Kantarci, Hussein T. Mouftah
IEEE Internet Things J.3
2025 Joint Optimization of Completion Ratio and Latency of Offloaded Tasks With Multiple Priority Levels in 5G Edge
abstract
Multi-Access Edge Computing (MEC) is widely recognized as an essential enabler for applications that necessitate minimal latency. However, the dropped task ratio metric has not been studied thoroughly in literature. Neglecting this metric can potentially reduce the system’s capability to effectively manage tasks, leading to an increase in the number of eliminated or unprocessed tasks. This paper presents a 5G-MEC task offloading scenario with a focus on minimizing the dropped task ratio, computational latency, and communication latency. We employ Mixed Integer Linear Programming (MILP), Particle Swarm Optimization (PSO), and Genetic Algorithm (GA) to optimize the latency and dropped task ratio. We conduct an analysis on how the quantity of tasks and User Equipment (UE) impacts the ratio of dropped tasks and the latency. The tasks that are generated by UEs are classified into two categories: urgent tasks and non-urgent tasks. The UEs with urgent tasks are prioritized in processing to ensure a zero-dropped task ratio. Our proposed method improves the performance of the baseline methods, First Come First Serve (FCFS) and Shortest Task First (STF), in the context of 5G-MEC task offloading. Under the MILP-based approach, the latency is reduced by approximately 55% compared to GA and 35% compared to PSO. The dropped task ratio under the MILP-based approach is reduced by approximately 70% compared to GA and by 40% compared to PSO.
Parisa Fard Moshiri, Murat Simsek, Burak Kantarci
IEEE Trans. Netw. Serv. Manag.2
2024 All Predict Cost Efficient Decides: A New Cost-Centric Ensemble Learning Method for Network Intrusions Detection
abstract
Machine Learning (ML) techniques have gained extensive attention for network intrusion detection. However, integrating ML approaches faces two primary challenges due to the presence of multi-class attacks and their varying impact levels on the network: the one-size-fits-all dilemma and the consideration of intrusion costs. Since ML models exhibit differing detection performances for each attack class, a single ML model may not suffice for predicting all attacks. Additionally, intrusion cost, a crucial concern for users and network service providers, is often overlooked in intrusion detection scheme development. To address these challenges, we propose a novel ensemble-learning framework called All Predict Cost Efficient Decides (APCED). APCED integrates multiple ML models, selecting an expert ML model for each attack class to minimize intrusion costs. In APCED, both damage cost and response cost determine the cost-efficient base estimators for ensemble learning, with an aggregation strategy employed for final decisions. We evaluate the performance of APCED using the NSL-KDD dataset. Numerical results demonstrate that APCED enhances the overall weighted F1 score by 81.13% compared to Adaboost and achieves an overall cost reduction of 54.7% and 87% compared to XGBoost and Adaboost, respectively.
Murat Simsek, Poonam Lohan, Burak Kantarci, Petar Djukic
GLOBECOM2
2024 Towards Sustainable Edge Computing: Efficient Task Offloading for Energy Efficiency and Latency Reduction
abstract
In networks with limited resources, the concept of offloading computation to Mobile Edge Computing (MEC) has emerged as a promising research direction with the advent of new services in fifth-generation (5G) networks. However, poorly designed offloading strategies can lead to excessive energy consumption and unpredictable latency, while the number of dropped tasks significantly impacts system efficiency. This paper presents a 5G-MEC task offloading scenario aimed at minimizing computation and communication latency, energy consumption, and the rate of dropped tasks. To achieve this, we employ Mixed Integer non-Linear Programming (MINLP) and Mixed Integer Linear Programming (MILP), comparing their performance with Particle Swarm Optimization (PSO) and Genetic Algorithm (GA). Our analysis considers the impact of the quantity of tasks and User Equipment (UE) on network parameters, distinguishing between urgent and non-urgent tasks. We ensure a zero-dropped task rate for urgent tasks. The proposed approach outperforms baseline techniques such as First Come First Serve (FCFS), Shortest Deadline First (SDF), and Urgent Tasks First (UTF) in the context of 5G-MEC task offloading. Specifically, compared to MILP, PSO, and GA, the MINLP-based approach reduces total latency by 12%, 34%, and 44%, respectively. Moreover, it decreases energy consumption by 8%, 30%, and 47% compared to MILP, PSO, and GA, respectively. The dropped task ratio is also reduced by 17%, 42%, and 65% under the MINLP-based approach compared to MILP, PSO, and GA, respectively.
Parisa Fard Moshiri, Murat Simsek, Burak Kantarci
GLOBECOM2
2024 On the Interplay Between Network Metrics and Performance of Mobile Edge Offloading
abstract
Multi-Access Edge Computing (MEC) emerged as a viable computing allocation method that facilitates offloading tasks to edge servers for efficient processing. The integration of MEC with 5G, referred to as 5G-MEC, provides real-time processing and data-driven decision-making in close proximity to the user. The 5G- MEC has gained significant recognition in task offloading as an essential tool for applications that require low delay. Nevertheless, few studies consider the dropped task ratio metric. Disregarding this metric might possibly undermine system efficiency. In this paper, the dropped task ratio and delay has been minimized in a realistic 5G- MEC task offloading scenario implemented in NS3. We utilize Mixed Integer Linear Programming (MILP) and Genetic Algorithm (GA) to optimize delay and dropped task ratio. We examined the effect of the number of tasks and users on the dropped task ratio and delay. Compared to two traditional offloading schemes, First Come First Serve (FCFS) and Shortest Task First (STF), our proposed method effectively works in 5G-MEC task offloading scenario. For MILP, the dropped task ratio and delay has been minimized by 20% and 2ms compared to GA.
Parisa Fard Moshiri, Murat Simsek, Burak Kantarci
ICC2
2024 Real-Time Binary Cell Phone Usage Detection and Classification on Vehicular Edge Devices
abstract
IoT binary classification tasks can benefit significantly from edge computing because it allows for real-time processing and decision-making. By gathering and processing data locally, edge devices can lower latency and enable quicker response times, which is beneficial for applications whose main aims are safety and security. Cell phone usage while driving is one of the worst scenarios that decreases traffic safety and causes accidents. A wide range of new applications and services could be possible with the convergence of IoT and cell phone detection in the car while in driving mode. Machine learning methods, which include the ability to track people and objects in realtime, increase public safety by identifying and preventing potential security threats and improve transportation system efficiency by streamlining traffic and easing congestion. Although object detection is the most common approach for cell phone detection, the binary classification approach has been proposed because of its fast processing ability and easy deployment on edge devices. The device, in consideration, incorporates an inside camera to gather driver image data to perform binary classification. After collecting images from edge devices, these data are prepared in detail in an IID (independently and identically distributed) manner for better training for deep learning models. After training, test results are obtained by interpolation and extrapolation analyses. Results show that interpolation accuracy increases by $1.1 \%$ and extrapolation accuracy increases by $16.5 \%$.
Murat Arda Onsu, Pankti Shah, Murat Simsek, Mark Fobert, Burak Kantarci
IWCMC3
2024 Machine learning-enabled hybrid intrusion detection system with host data transformation and an advanced two-stage classifier
abstract
Network Intrusion Detection Systems (NIDS) have been extensively investigated by monitoring real network traffic and analyzing suspicious activities. However, there are limitations in detecting specific types of attacks with NIDS, such as Advanced Persistent Threats (APT). Additionally, NIDS is restricted in observing complete traffic information due to encrypted traffic or a lack of authority. To address these limitations, a Host-based Intrusion Detection system (HIDS) evaluates resources in the host, including logs, files, and folders, to identify APT attacks that routinely inject malicious files into victimized nodes. In this study, a hybrid network intrusion detection system that combines NIDS and HIDS is proposed to improve intrusion detection performance. The host data undergoes a Language Processing (NLP)-based Bidirectional Encoder Representations from Transformers (BERT) model from textual representation to a numerical one in order to process host data in a similar way to the network flow data through machine learning models. The feature flattening technique is applied to flatten two-dimensional host-based features that is provided by BERT into one-dimensional vectors so that host-based and network flow-based features can be processed by advanced Machine Learning (ML) models. In order to enhance HIDS effectiveness, a two-stage collaborative classifier is utilized, which applies two tiers of machine learning algorithms, binary and multi-class classifiers, to detect network intrusions. Once a binary classifier is used to detect benign samples to reduce the complexity of the original problem, the attack data are classified by a multi-class supervised learner to identify attack types. Hence, the overall performance of the two-stage collaborative model outperforms the baseline classifier, XGBoost. The proposed method is shown to generalize across two well-known datasets, CICIDS 2018 and NDSec-1. The performance of XGBoost, which represents conventional ML, is evaluated. Combining host and network features enhances attack detection performance (macro average F1 score) by 8.1% under the CICIDS 2018 dataset and 3.7% under the NDSec-1 dataset. Meanwhile, the two-stage collaborative classifier improves detection performance for most single classes, especially for DoS-LOIC-UDP and DoS-SlowHTTPTest, with improvements of 30.7% and 84.3%, respectively, when compared with the traditional ML models.
Murat Simsek, Burak Kantarci, Mehran Bagheri, Petar Djukic
Comput. Networks2
2024 TableStrRec: framework for table structure recognition in data sheet images
Johan Fernandes, Bin Xiao 0008, Murat Simsek, Burak Kantarci, Shahzad Khan 0002, Ala Abu Alkheir
Int. J. Document Anal. Recognit.3
2023 Multidomain transformer-based deep learning for early detection of network intrusion
abstract
Timely response of Network Intrusion Detection Systems (NIDS) is constrained by the flow generation process which requires accumulation of network packets. This paper introduces Multivariate Time Series (MTS) early detection into NIDS to identify malicious flows prior to their arrival at target systems. With this in mind, we first propose a novel feature extractor, Time Series Network Flow Meter (TS-NFM), that represents network flow as MTS with explainable features, and a new benchmark dataset is created using TS-NFM and the meta-data of CICIDS2017, called SCVIC-TS-2022. Additionally, a new deep learning-based early detection model called Multi-Domain Transformer (MDT) is proposed, which incorporates the frequency domain into Transformer. This work further proposes a Multi-Domain Multi-Head Attention (MD-MHA) mechanism to improve the ability of MDT to extract better features. Based on the experimental results, the proposed methodology improves the earliness of the conventional NIDS (i.e., percentage of packets that are used for classification) by 5 ×104times and duration-based earliness (i.e., percentage of duration of the classified packets of a flow) by a factor of 60, resulting in a 84.1% macro F1 score (31% higher than Transformer) on SCVIC-TS-2022. Additionally, the proposed MDT outperforms the state-of-the-art early detection methods by 5% and 6% on ECG and Wafer datasets, respectively.
Jinxin Liu 0001, Murat Simsek, Michele Nogueira Lima, Burak Kantarci
GLOBECOM2
2023 Anomalous Behaviour Detection via Event-Based Metric with Sequential Tracking in a V2X Environment
abstract
Massive amount of data transmission in vehicle-to-everything (V2X) settings lead to heavy utilization of communication channels. Furthermore, reliable connectivity and fast data transmission can be achieved by either re-engineering the network architecture or using efficient methods that alter data attributes, such as volume. This paper proposes a new method called Sequential Tracking along with an event-based distracted driving detection algorithm. Existing studies using machine learning models and object detection aim to detect distracted drivers and send their detection outcomes to the cloud or edge units. However, minimizing the data transmission or exchange overhead remains understudied. Therefore, the proposed Sequential Tracking method with an event-based algorithm is applied to distracted driving detection models to reduce the data transmission overhead due to false predictions, i.e., false positives or false negatives. Furthermore, the proposed method considers the camera's inference time and storage capacity since AI models are deployed to edge units for these kinds of tasks. Numerical results, with the inclusion of parameter tuning, confirm that the overall accuracy performance of the model can be improved from 87% to 91%. Moreover, following upon parameter-tuning, false predictions in the test dataset are eliminated, and the number of data points is reduced to less than one-tenth leading to significant traffic reduction between the edge unit and the cloud.
Murat Arda Onsu, Murat Simsek, Burak Kantarci
GLOBECOM2
2023 Knowledge-Based Zero-Touch Security under Host and Network Flow Features Merger
abstract
Incorporating machine learning algorithms with Intrusion Detection System (IDS) can detect network intrusions without human intervention and aims for Zero Touch Networks (ZTN). In this research, an automatic network-based features and host-based features integrated intrusion detection scheme is presented to improve the performance of network attack detection under the SCVIC-CIDS-2021 dataset which is derived from the integration of network packets and host logs of the CSE-CIC-IDS2018 dataset. Auto-encoder (AE) and Gated Recurrent Unit (GRU) are utilized for feature derivation to overcome the dimensionality mismatch between network-based and host-based features. The knowledge-based Prior Knowledge Input (PKI) model is used to combine unsupervised extra knowledge with a pre-trained supervised model for the final classification results. The results of the experiment reveal that the integration of network-based and host-based features is effective and the PKI model improves the performance of the original ML classification algorithm as well. Under the test set, the maximum achievable macro average F1-score reaches up to 97.08% which points out approximately 9% improvement compared to the best baseline performance.
Yu Shen 0001, Murat Simsek, Burak Kantarci, Hussein T. Mouftah, Mehran Bagheri, Petar Djukic
ICC2
2023 Multi-Modal OCR System for the ICT Global Supply Chain
abstract
Optical Character Recognition (OCR) tools have been widely used to extract text content from images in many applications including Information and Communications Technology (ICT) supply chains. Due to the characteristics of datasheets in global ICT supply chains, models trained with popular public datasets often suffer from domain adaptation problems. First, popular open source text recognition datasets do not contain all the characters and symbols that appear in the ICT documents, meaning that models trained with these datasets cannot recognize these special characters and symbols. Second, these datasets also do not contain the samples with multiple words and multiple lines assuming that there is an Text Detection model that can extract text areas perfectly, which is not practical for ICT documents. Besides, as far as we know, there is no open-source dataset specifically designed that can be used to evaluate the OCR tools in the ICT domain. Therefore, in this study, we first build a benchmark dataset for the text recognition problem in the ICT domain, which includes the special characters and symbols in the ICT domain, and samples with multiple lines and multiple words. Then we propose a novel multi-modal sequence-to-sequence model, which not only take images as input but also their corresponding their texts generated by a pre-trained model. We conducted extensive experiments to evaluate the proposed multi-modal method on the proposed dataset, and the empirical results show that the proposed method can recognize special characters and symbols, samples with multiple lines and multiple words, outperform benchmark models consistently.
Bin Xiao 0008, Yakup Akkaya, Murat Simsek, Burak Kantarci, Ala Abu Alkheir
ICC3
2023 AI-enabled cluster head selection through modified density based clustering in Aeronautical Ad Hoc Networks
Mohsen Shahbazi, Murat Simsek, Burak Kantarci
Ad Hoc Networks2
2023 Table detection for visually rich document images
Bin Xiao 0008, Murat Simsek, Burak Kantarci, Ala Abu Alkheir
Knowl. Based Syst.2
2023 Utility-Aware Legitimacy Detection of Mobile Crowdsensing Tasks via Knowledge-Based Self Organizing Feature Map
abstract
In Mobile Crowdsensing (MCS), fake tasks can drain significant amount of resources. This paper proposes a new methodology to determine a proper time window for the training dataset and the impact of the accuracy of task legitimacy detection on the MCS campaign performance. To reach the desired performance, the task legitimacy detection is utilized in such a way that while legitimate tasks are kept, the fake tasks are eliminated as much as possible in the MCS platform through machine learning (ML) prediction. The proposed methodology is evaluated for legitimacy detection under multiple ML methods. Moreover, a knowledge-based fake task detection technique with effective feature selection is formulated to ensure fake tasks are filtered at the MCS servers. Detection accuracy is improved by using shorter time frame in training and longer time frame in prediction. The overall performance improvement based on profit, cost, legitimate tasks loss ratio, and fake tasks elimination ratio has been achieved under three different sizes of training datasets to verify the efficiency of the proposed methodology. Moreover, Prior Knowledge Input with Self-Organizing Feature Map outperforms the conventional legitimacy detection by 5.48%, 12.11% and 58.05% in terms of test accuracy, profit and cost under the small dataset, respectively.
Murat Simsek, Burak Kantarci, Azzedine Boukerche
IEEE Trans. Mob. Comput.1
2022 Collaborative Feature Maps of Networks and Hosts for AI-driven Intrusion Detection
abstract
Intrusion Detection Systems (IDS) are critical secu-rity mechanisms that protect against a wide variety of network threats and malicious behaviors on networks or hosts. As both Network-based IDS (NIDS) or Host-based IDS (HIDS) have been widely investigated, this paper aims to present a Combined Intrusion Detection System (CIDS) that integrates network and host data in order to improve IDS performance. Due to the scarcity of datasets that include both network packet and host data, we present a novel CIDS dataset formation framework that can handle log files from a variety of operating systems and align log entities with network flows. A new CIDS dataset named SCVIC-CIDS-2021 is derived from the meta-data from the well-known benchmark dataset, CIC-IDS-2018 by utilizing the proposed framework. Furthermore, a transformer-based deep learning model named CIDS-Net is proposed that can take network flow and host features as inputs and outperform baseline models that rely on network flow features only. Experimental results to evaluate the proposed CIDS-Net under the SCVIC-CIDS-2021 dataset support the hypothesis for the benefits of combining host and flow features as the proposed CIDS- N et can improve the macro F1 score of baseline solutions by 6.36 % (up to 99.89%).
Jinxin Liu 0001, Murat Simsek, Burak Kantarci, Mehran Bagheri, Petar Djukic
GLOBECOM2
2022 Prior Knowledge based Advanced Persistent Threats Detection for IoT in a Realistic Benchmark
abstract
The number of Internet of Things (IoT) devices being deployed into networks is growing at a phenomenal pace, which makes IoT networks more vulnerable in the wireless medium. Advanced Persistent Threat (APT) is malicious to most of the network facilities and the available attack data for training the machine learning-based Intrusion Detection System (IDS) is limited when compared to the normal traffic. Therefore, it is quite challenging to enhance the detection performance in order to mitigate the influence of APT. Therefore, Prior Knowledge Input (PKI) models are proposed and tested using the SCVIC-APT-2021 dataset. To obtain prior knowledge, the proposed PKI model pre-classifies the original dataset with unsupervised clustering method. Then, the obtained prior knowledge is incorporated into the supervised model to decrease training complexity and assist the supervised model in determining the optimal mapping between the raw data and true labels. The experimental findings indicate that the PKI model outperforms the supervised baseline, with the best macro average F1-score of 81.37%, which is 10.47% higher than the baseline.
Yu Shen 0001, Murat Simsek, Burak Kantarci, Hussein T. Mouftah, Mehran Bagheri, Petar Djukic
GLOBECOM2
2022 Efficient Information Sharing in ICT Supply Chain Social Network via Table Structure Recognition
abstract
The global Information and Communications Technology (ICT) supply chain is a complex network consisting of all types of participants. It is often formulated as a Social Network to discuss the supply chain network's relations, properties, and development in supply chain management. Information sharing plays a crucial role in improving the efficiency of the supply chain, and datasheets are the most common data format to describe e-component commodities in the ICT supply chain because of human readability. However, with the surging number of electronic documents, it has been far beyond the capacity of human readers, and it is also challenging to process tabular data automatically because of the complex table structures and heterogeneous layouts. Table Structure Recognition (TSR) aims to represent tables with complex structures in a machine-interpretable format so that the tabular data can be processed automatically. In this paper, we formulate TSR as an object detection problem and propose to generate an intuitive representation of a complex table structure to enable structuring of the tabular data related to the commodities. To cope with border-less and small layouts, we propose a cost-sensitive loss function by considering the detection difficulty of each class. Besides, we propose a novel anchor generation method using the character of tables that columns in a table should share an identical height, and rows in a table should share the same width. We implement our proposed method based on Faster-RCNN and achieve 94.79% on mean Average Precision (AP), and consistently improve more than 1.5% AP for different benchmark models.
Bin Xiao 0008, Yakup Akkaya, Murat Simsek, Burak Kantarci, Ala Abu Alkheir
GLOBECOM3
2022 Handling big tabular data of ICT supply chains: a multi-task, machine-interpretable approach
abstract
Due to the characteristics of Information and Communications Technology (ICT) products, the critical information of ICT devices is often summarized in big tabular data shared across supply chains. Therefore, it is critical to automatically interpret tabular structures with the surging amount of electronic assets. To transform the tabular data in electronic documents into a machine-interpretable format and provide layout and semantic information for information extraction and interpretation, we define a Table Structure Recognition (TSR) task and a Table Cell Type Classification (CTC) task. We use a graph to represent complex table structures for the TSR task. Meanwhile, table cells are categorized into three groups based on their functional roles for the CTC task, namely Header, Attribute, and Data. Subsequently, we propose a multi-task model to solve the defined two tasks simultaneously by using the text modal and image modal features. Our experimental results show that our proposed method can outperform state-of-the-art methods on ICDAR2013 and UNLV datasets.
Bin Xiao 0008, Murat Simsek, Burak Kantarci, Ala Abu Alkheir
GLOBECOM2
2022 Collaborative Self Organizing Map with DeepNNs for Fake Task Prevention in Mobile Crowdsensing
abstract
Mobile Crowdsensing (MCS) is a sensing paradigm that has transformed the way that various service providers collect, process, and analyze data. MCS offers novel processes where data is sensed and shared through mobile devices of the users to support various applications and services for cutting-edge technologies. However, various threats, such as data poisoning, clogging task attacks and fake sensing tasks adversely affect the performance of MCS systems, especially their sensing, and computational capacities. Since fake sensing task submissions aim at the successful completion of the legitimate tasks and mobile device resources, they also drain MCS platform resources. In this work, Self Organizing Feature Map (SOFM), an artificial neural network that is trained in an unsupervised manner, is utilized to pre-cluster the legitimate data in the dataset, thus fake tasks can be detected more effectively through less imbalanced data where legitimate/fake tasks ratio is lower in the new dataset. After pre-clustered legitimate tasks are separated from the original dataset, the remaining dataset is used to train a Deep Neural Network (DeepNN) to reach the ultimate performance goal. Pre-clustered legitimate tasks are appended to the positive prediction outputs of DeepNN to boost the performance of the proposed technique, which we refer to as pre-clustered DeepNN (PrecDeepNN). The results prove that the initial average accuracy to discriminate the legitimate and fake tasks obtained from DeepNN with the selected set of features can be improved up to an average accuracy of 0.9812 obtained from the proposed machine learning technique.
Murat Simsek, Burak Kantarci, Azzedine Boukerche
ICC1
2022 On cropped versus uncropped training sets in tabular structure detection
Yakup Akkaya, Murat Simsek, Burak Kantarci, Shahzad Khan 0002
Neurocomputing2
2022 TableDet: An end-to-end deep learning approach for table detection and table image classification in data sheet images
Johan Fernandes, Murat Simsek, Burak Kantarci, Shahzad Khan 0002
Neurocomputing2
2021 All Predict Wisest Decides: A Novel Ensemble Method to Detect Intrusive Traffic in IoT Networks
abstract
Internet of things (IoT) networks confront vari-ous network intrusion threats due to massively interconnected nodes that form an extensive attack surface for adversaries. Machine learning (ML)-based approaches are widely investigated to address network intrusions. It becomes further challenging to achieve promising performance for multi-class classification so to identify each attack type rather than detection of the presence of intrusion, which involves binary classification. ML models perform divergent detection performance in each class, so it is challenging to select one ML model applicable to all classes prediction. With this in mind, we propose an innovative ensemble learning framework, namely All Predict Wisest Decides (APWD) that builds on training of multiple ML models and testing them independently so to obtain prediction performance for all classes. For each attack category, an expert (i.e., wisest) model that performs the best F1 score, accuracy, lowest false detection rate is determined according to individual model results. The aggregation module makes decisions relying upon the wisest model determined for each class. APWD is a generic framework, and the types of MLs and the number of MLs can be customized in APWD. Experiments under a popular public dataset, NSL-KDD verify the proposed approach APWD by demonstrating that APWD boosts overall accuracy to 0.797, comparing 0.772 by XGBoost, 0.758 by RF, and 0.584 by Adaboost. Moreover, in certain attack types R2L, APWD increases F1 score by a factor of 18, from 0.022 by RF to 0.421.
Murat Simsek, Burak Kantarci, Petar Djukic
GLOBECOM2
2021 Federated Learning-Based Risk-Aware Decision to Mitigate Fake Task Impacts on Crowdsensing Platforms
abstract
Mobile crowdsensing (MCS) leverages distributed and non-dedicated sensing concepts by utilizing sensors embedded in a large number of mobile smart devices. However, the openness and distributed nature of MCS leads to various vulnerabilities and consequent challenges to address. A malicious user submitting fake sensing tasks to an MCS platform may be attempting to consume resources from any number of participants’ devices; as well as attempting to clog the MCS server. In this paper, a novel approach that is based on horizontal federated learning is proposed to identify fake tasks that contain a number of independent detection devices and an aggregation entity. Detection devices are deployed to operate in parallel with each device equipped with a machine learning (ML) module, and an associated training dataset. Furthermore, the aggregation module collects the prediction results from individual devices and determines the final decision with the objective of minimizing the prediction loss. Loss measurement considers the lost task values with respect to misclassification, where the final decision utilizes a risk-aware approach where the risk is formulated as a function of the utility loss. Experimental results demonstrate that using federated learning-driven illegitimate task detection with a risk aware aggregation function improves the detection performance of the traditional centralized framework. Furthermore, the higher performance of detection and lower loss of utility can be achieved by the proposed framework. This scheme can even achieve 100% detection accuracy using small training datasets distributed across devices, while achieving slightly over an 8% increase in detection improvement over traditional approaches.
Murat Simsek, Burak Kantarci
ICC2
2021 Prior Knowledge Input to Improve LSTM Auto-encoder-based Characterization of Vehicular Sensing Data
abstract
Precision in event characterization in connected vehicles has become increasingly important with the responsive connectivity that is available to modern vehicles. Event characterization via vehicular sensors is utilized in safety and autonomous driving applications in vehicles. While characterization systems are capable of predicting risky driving patterns, the precision of such systems remains an open issue. The major issues against the driving event characterization systems need to be addressed in connected vehicle settings, which are the heavy imbalance and the event infrequency of the driving data and the existence of the time-series detection systems that are optimized for vehicular settings. To overcome the problems, we introduce the application of the prior-knowledge input method to the characterization systems. Furthermore, we propose a recurrent-based denoising auto-encoder network to populate the existing data for a more robust training process. The results of the conducted experiments show that the introduction of knowledge-based modeling enables the existing systems to reach significantly higher accuracy and F1-score levels. Ultimately, the combination of the two methods enables the proposed model to attain a 14.7% accuracy boost over the baseline by achieving an accuracy of 0.96.
Nima Taherifard, Murat Simsek, Charles Lascelles, Burak Kantarci
ICC2
2021 TabCellNet: Deep learning-based tabular cell structure detection
JiChu Jiang, Murat Simsek, Burak Kantarci, Shahzad Khan 0002
Neurocomputing2
2021 Empowering Self-Organized Feature Maps for AI-Enabled Modeling of Fake Task Submissions to Mobile Crowdsensing Platforms
abstract
Mobile crowdsensing (MCS) has emerged as a ubiquitous solution for data collection from embedded sensors of smart devices to improve the sensing capacity and reduce sensing costs in large regions. Due to the ubiquitous nature of MCS services, smart devices require awareness of against misbehaving users that are becoming smarter to clog the resources in such a nondedicated sensing environment. In an MCS setting, the primary goal of a fake sensing task submission is to keep participant devices occupied, such as the battery, sensing, storage, and computing. Since the development of robust sensing campaigns highly depends on the existence of a realistic model of misbehaving users, this article leverages artificial intelligence and introduces a region-based self-organizing feature map (SOFM)-based model on user movement patterns so as to place the fake sensing tasks with the objective of maximum impacted participants and recruits. Uniformly and randomly initialized neurons are designed with fixed and adaptive quantities that are determined based upon the affected area on the covered terrain. Through numerical studies, we show that the impact of the SOFM structures can affect up to 46% of the participants and up to 37% of the recruits under various SOFM topologies. Furthermore, SOFM-based task submission models can increase the energy consumption in recruited devices by up to 39% due to the illegitimate task submission.
Yueqian Zhang, Murat Simsek, Burak Kantarci
IEEE Internet Things J.2
2021 AI-driven autonomous vehicles as COVID-19 assessment centers: A novel crowdsensing-enabled strategy
Murat Simsek, Azzedine Boukerche, Burak Kantarci, Shahzad Khan 0002
Pervasive Mob. Comput.1
2020 Region-Aware Bagging and Deep Learning-Based Fake Task Detection in Mobile Crowdsensing Platforms
abstract
Mobile crowdsensing (MCS) is a distributed sensing concept that enables ubiquitous sensing services via various builtin sensors in smart devices. However, MCS systems are vulnerable because of being non-dedicated. Especially, submission of fake tasks with the aim of clogging participants device resources as well as MCS servers is a crucial threat to MCS platforms. In this paper, we propose an ensemble learning-based solution for MCS platforms to mitigate illegitimate tasks. Furthermore, we also integrate k-means-based classification with the proposed method to extract region-specific features as input to the machine learningbased fake task detection. Through simulations, we compare the ensemble method to a previously proposed Deep Belief Network (DBN)-based fake task detection, which is also shown to improve performance in terms of accuracy, F1 score, recall, precision and geometric mean score (G-mean) with the integration of regionawareness. Our validation results show that the ensemble machine learning-based detection can eliminate majority of the fake tasks, with up to 0.995 precision, 0.997 recall, 0.996 F1, 0.993 accuracy and 0.982 G-Mean. Furthermore, the proposed solution introduces savings up to 12.18% battery of mobile devices while reducing the impacted recruits to 0.25% and protecting up to 10.59% participants against malicious sensing tasks.
Murat Simsek, Burak Kantarci
GLOBECOM2
2020 Self Organizing Feature Map-Integrated Knowledge-Based Deep Network Against Fake Crowdsensing Tasks
abstract
Mobile Crowdsensing (MCS) builds on the Sensing as a Service model, and is considered to be an integral component of the Internet of Things systems. Since MCS does not build on a thoroughly assessed and established trust mechanism between all parties various threats including data poisoning, fake sensing tasks and clogging task attacks remain challenges. Fake task submissions are the least investigated although they have the potential to drain significant amount of resources (e.g. battery, sensors, processing, storage) and clog the MCS servers. This paper proposes a knowledge-based technique alongside sequential feature selection methodology to detect fake sensing tasks submitted to the MCS servers so that the tasks do not get assigned to the participants but filtered at the MCS servers. The proposed methodology is compared to fake task detection under Knowledge Based Deep Neural Network which is also enhanced by feature selection, and the simulation results show that the proposed methodology, by utilizing Deep Prior Knowledge Input with Self-Organizing Feature Map can outperform the deep neural network-based detection by 9.7% in terms of accuracy.
Murat Simsek, Burak Kantarci, Azzedine Boukerche
GLOBECOM1
2020 Deep Belief Network-based Fake Task Mitigation for Mobile Crowdsensing under Data Scarcity
abstract
Mobile crowdsensing (MCS) is a ubiquitous sensing paradigm that emerged in the form of”sensed data as a service” model in the Internet of Things Era. Distributed nature of MCS results in vulnerabilities at the MCS platforms as well as participating devices that provide sensory data services. Submission of fake tasks with the aim of clogging sensing server resources and draining participating device batteries is a crucial threat that has not been investigated well. In this paper, we provide a detailed analysis by modeling a deep belief network (DBN) when the available sensory data is scarce for analysis. With oversampling to cope with the class imbalance challenge, a Principal Component Analysis (PCA) module is implemented prior to the DBN and weights of various features of sensing tasks are analyzed under varying inputs. The experimental results show that the presented DBN-driven fake task mitigation detection of fake sensing tasks can ensure up to 0.92 accuracy, 0.943 precision and up to 0.928 F1 score outperforming prior work on MCS data with deep learning networks.
Yueqian Zhang, Murat Simsek, Burak Kantarci
ICC3
2020 Ensemble Learning Against Adversarial AI-driven Fake Task Submission in Mobile Crowdsensing
abstract
Non-dedicated nature of mobile crowdsensing (MCS) systems introduces vulnerabilities for MCS platforms in terms of sensing, computing, storage, and battery resources. The advent of adversarial artificial intelligence (AI) leads to high impact malicious behavior when adversaries aim to clog the resources of such a non-dedicated and ubiquitous system. This paper proposes an ensemble learning-based methodology for MCS platforms in order to mitigate the impacts of adversarial AI-driven fake task submission attacks, which are intelligently designed so to clog resources such as batteries, sensing, or memory resources. We validate our proposal through realistic simulations to generate crowdsensing data under two different cities, and intelligent fake task submissions under adversarial self-organizing maps. The experimental results show that when the submitted tasks undergo a Gradient Boosting-based classifier prior to being assigned to participants, the proposed solution can introduce battery savings at the participant devices up to 23%, and the impacted recruit population can be reduced from 24% to 6% whereas the defense mechanism can achieve an overall accuracy level above 98% concerning the legitimacy of the submitted tasks.
Yueqian Zhang, Murat Simsek, Burak Kantarci
ICC2
2020 High Precision Deep Learning-Based Tabular Position Detection
abstract
Documents are constantly being processed within supply chains in various industries throughout the globe. Within those documents, often times the most important content is stored in tabular format. Therefore an automated technique for supply chain document processing is highly desired. Deep learning approaches show promise to deliver an end-to-end extraction model. However, it has been shown that tabular detection accuracy is not always correlated to tabular localization accuracy. Portions of the desired tabular information can easily be cropped out due to a lack of localization accuracy. In this paper, we propose a two stage convolutional neural network-based deep learning framework to improve tabular localization accuracy. We use pre-trained backbone network ResNet-50 and then apply transfer learning to fit our application. One of our main contributions is the introduction of the KL loss function into Faster-RCNN. Once the bounding box variances are acquired from the KL loss function, we introduce a voting procedure with soft-non-maximum suppression (Soft-NMS) to improve localization performance. The proposed framework is trained and evaluated on public and private datasets that span from scientific documents to various electronic components. Our test results show that the precision of tabular detection can be improved by 1.2% while achieving the same recall as other state-of-the-art models on the public ICDAR2013 dataset. Furthermore, a large improvement in precision has been achieved at extremely high intersection over union (IoU) thresholds (i.e. 95%). Thus, 10.9% higher precision is achieved at 95% IoU for ICDAR2013. For another public dataset, namely ICDAR2017, 8.4% higher precision is achieved at 95% IoU .
JiChu Jiang, Murat Simsek, Burak Kantarci, Shahzad Khan 0002
ISCC2
2020 Knowledge-Based Machine Learning Boosting for Adversarial Task Detection in Mobile Crowdsensing
abstract
Mobile Crowdsensing (MCS) leverages Sensing as a Service paradigm to contribute to the Internet of Things ecosystems through non-dedicated sensing capabilities of smart mobile devices. Distributed and non-trusted nature of MCS systems are vulnerable against various threats for the devices, MCS platforms, as well as the participating devices that provide sensory data services. Out of the many threats, submission of fake tasks may lead to drained resources at the participating devices, and clogged sensing server resources at MCS platforms. In this paper, classical machine learning (ML) performance is boosted by knowledge-based methods and sequential feature selection which is proposed for the first time against fake tasks submission to MCS platforms. Prior Knowledge Input and Prior Knowledge Input with Difference exploit AdaBoost and Decision Tree methods as initial accuracy to improve the accuracy of learning the legitimacy of submitted tasks to MCS platforms. Moreover, Sequential Feature Selection is implemented to investigate further improvements for the detection of task legitimacy in MCS campaigns. Intelligently selected 5 features amongst 10 possible features and implementation of knowledge-based methods boost the accuracy of machine learning performance from 93.67% to 97.37% for AdaBoost, and from 92.28% to 97.58% for Decision Trees.
Murat Simsek, Burak Kantarci, Azzedine Boukerche
ISCC1
2020 Holistic design for deep learning-based discovery of tabular structures in datasheet images
Ertugrul Kara, Mark Traquair, Murat Simsek, Burak Kantarci, Shahzad Khan 0002
Eng. Appl. Artif. Intell.3
2019 Sensory Data-Driven Modeling of Adversaries in Mobile Crowdsensing Platforms
abstract
The advent of Mobile crowdsensing (MCS) facilitates the adoption of ubiquitous sensing solutions in smart environments. Despite its benefits, MCS calls for proper security and trust solutions. Various threatening attacks, such as injection attacks, can compromise both the veracity and integrity of crowdsensed data. This work leverages adversarial machine learning to introduce a smart injection attacker model (SINAM) that may be used in the design of security solutions against injection attacks in MCS. SINAM has been validated during an authentic MCS campaign. Unlike most random data injection models, SINAM monitors data traffic in an online- learning manner, successfully injecting malicious data across multiple victims with near-perfect accuracy rates of 99%. SINAM uses accomplices within the sensing campaign to predict accurate injections based on both behavioral analysis and context similarities.
Kyle Quintal, Ertugrul Kara, Murat Simsek, Burak Kantarci, Herna L. Viktor
GLOBECOM3
2019 Self Organizing Feature Map for Fake Task Attack Modelling in Mobile Crowdsensing
abstract
Clogging attacks in mobile crowdsensing (MCS) denote injection of fake sensing tasks into MCS campaigns in order to interfere with service ability and user participation in sensing campaigns. Due to the lack of a realistic location-based and energy-oriented clogging attack model in MCS, this type of attacks have not been well investigated. To this end, for the first time, we introduce a self organizing feature map (SOFM)-based clogging attack model that aims at maximizing the number of affected participants according to the location of attack zones. These zones are identified by clustering 2-D coordinates of all potential participants and finding out aggregation areas of their mobile devices. We evaluate and verify the introduced attack model via simulations by comparing it to an attack model that relies on random mobility of illegitimate tasks over the attack zones. Our simulation results demonstrate that SOFM-based modeling of clogging attacks in MCS results in a significant impact with almost 50% affected participant population, 24% affected recruitment decisions, and up to 28% energy overhead introduced by illegitimate tasks injected to the MCS campaigns.
Yueqian Zhang, Murat Simsek, Burak Kantarci
GLOBECOM2
2019 Trustworthiness and Comfort-Aware Participant Recruitment for Mobile Crowd-Sensing in Smart Environments
abstract
Mobile crowd-sensing (MCS) has gained significant momentum in recent years for sensory data acquisition through non-dedicated sensors. Even though most of the participants are assumed to be willing to participate in the sensing campaigns, some smartphone users may be reluctant to grant access to particular sensors in their devices. Besides this, the presence of adversaries in the participant pool makes the participant selection problem further challenging. With these in mind, we introduce Reputation and Comfort Level Aware Participant Selection (RACLAPS)which allows participants to modify their available set of sensors that can be accessed by central platform/task publisher i.e. participants can turn-off sensors according to their comfort. To demonstrate the effect of RA-CLAPS, we utilize Selective and Reputation-aware Recruitment (SRR) in which participants are selective in choosing their task to sense solely to improve income. To enrich the discussion, Non-Selective and Reputation-aware Recruitment (NSR) is also considered in which participants are given no choice regarding selectiveness, sensor configuration. Simulation results show that RA-CLAPS improves average user utility by 7.6% compared to its predecessor, SRR. We also mention that average discomfort per participant is reduced by 6.4% under RA-CLAPS when compared to non-selective and selective recruitment approaches.
Venkat Surya Dasari, Burak Kantarci, Murat Simsek
ISCC3