VLDB 2026 Research / reviewers in the wild / expert
Khaled B. Shaban
dblp:20/835 · also Khaled Bashir Shaban
· DBLP profile ↗
60ranked-venue papers
2as first author
25since 2021 · last 2026
0000-0002-5688-7515ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 14 since 2021Computer networks · 12 · 1 since 2021Security and privacy · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashDetR: A deep learning pipeline for early detection and time estimation of flashover in high-voltage insulators using infrared videosabstractFlashover in high-voltage insulators poses a significant risk to power system reliability, potentially leading to outages and safety hazards. This study introduces an innovative deep learning-based approach for early prediction of flashover events and time-to-flashover estimation by analyzing infrared videos of dry band arcing, a known precursor to flashover. In this work, we propose a pipeline named Flashover Detector and Time Estimator , which integrates a transformer-based model to accurately predict flashover occurrences, while a Three Dimensional Convolutional Neural Network-based model estimates the time to flashover. Flashover Detector and Time Estimator progressively samples video frames at multiple scales, enhancing prediction accuracy. Experimental results demonstrate that the models achieve up to 88.73% accuracy in predicting flashover events and a mean absolute error of 3.41 in time-to-flashover estimation. These findings substantially improve the ability to implement preventive measures. Flashover Detector and Time Estimator thus represents a significant advancement in proactively managing power system reliability, with demonstrated effectiveness and real-time application potential. • End-to-end DL model for early flashover prediction and precise time-to-flashover. • IR video dataset captured in controlled conditions, showing full DBA progression. • High-accuracy models for early flashover detection and low MAE for time-to-flashover prediction. Najmath Ottakath, Abdulla Lutfi, Ali Hamdi, Khaled B. Shaban, Ayman H. El-Hag |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | LLM-Ops and Ensemble Intelligence for Robust LLM Performance: Integrating Fine-Tuning and Majority VotingabstractThis paper presents a novel approach that combines LLM-Ops with ensemble intelligence to enhance document processing accuracy. We introduce a multi-OCR pipeline that leverages four distinct OCR engines and four fine-tuned lightweight LLMs in a two-tier majority voting framework. Through automated fine-tuning after every 500 processed records, our system demonstrates that lightweight (7B parameter) models can achieve performance comparable to much larger (27B parameter) alternatives. Experimental results show field accuracy improvements from $85.6 \%$ to $94.5 \%$ after three fine-tuning cycles, with processing speeds twice as fast as larger models. The continuous improvement loop enabled by our LLM-Ops framework ensures the system evolves with minimal human intervention, making advanced document intelligence more accessible and deployable for real-world applications. Osama Hosam Abdellatif, Ahmed Ayman, Abdelrahman Nader, Ali Hamdi, Khaled B. Shaban |
AICCSA | 5 |
| 2025 | Scaling Arabic Medical Chatbots Using Synthetic Data: Enhancing Generative AI with Synthetic Patient RecordsabstractThe development of medical chatbots in Arabic is significantly constrained by the scarcity of large-scale, highquality annotated datasets. While prior efforts compiled a dataset of 20,000 Arabic patient-doctor interactions from social media to fine-tune large language models (LLMs), model scalability and generalization remained limited. In this study, we propose a scalable synthetic data augmentation strategy to expand the training corpus to 100,000 records. Using advanced generative AI systems-ChatGPT-4o and Gemini 2.5 Pro-we generated 80,000 contextually relevant and medically coherent synthetic question-answer pairs grounded in the structure of the original dataset. These synthetic samples were semantically filtered, manually validated, and integrated into the training pipeline. We fine-tuned five LLMs, including Mistral-7B and AraGPT2, and evaluated their performance using BERTScore metrics and expert-driven qualitative assessments. To further analyze the effectiveness of synthetic sources, we conducted an ablation study comparing ChatGPT-4o and Gemini-generated data independently. The results showed that ChatGPT-4o data consistently led to higher F1-scores and fewer hallucinations across all models. Overall, our findings demonstrate the viability of synthetic augmentation as a practical solution for enhancing domain-specific language models in low-resource medical NLP, paving the way for more inclusive, scalable, and accurate Arabic healthcare chatbot systems. Abdulrahman Allam, Seif Ahmed, Ali Hamdi, Khaled B. Shaban |
AICCSA | 4 |
| 2025 | MHA-DQN: Personalized Route Planning for Asthma Patients Using Multi-Head Attention and Deep Reinforcement LearningabstractIndividuals with asthma face significant health risks in urban environments with poor air quality, yet traditional navigation systems fail to account for environmental factors that exacerbate respiratory conditions. To address this, we propose a novel hybrid deep learning framework integrating Multi-Head Attention (MHA) with Deep Q-Network (DQN) reinforcement learning for personalized, health-aware pedestrian route optimization. Our approach uniquely combines real-time environmental data—such as air quality index (AQI), temperature, and humidity—with individual mobility patterns and health profiles to prioritize respiratory safety. The model achieves $87.5 \%$ prediction accuracy and an $84 \%$ F1-score, with spatial errors of 1.43 MAE and 2.24 RMSE for route precision, outperforming conventional routing algorithms and standalone deep learning models. This work offers a scalable, adaptive solution for smart cities, enhancing mobility while safeguarding the health of vulnerable populations. Nada Ayman, Shaimaa Alaa Esmail, Ali Hamdi, Khaled B. Shaban, Hozaifa Kassab |
AICCSA | 4 |
| 2025 | Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer ExtractionabstractQuranic Question Answering presents unique challenges due to the linguistic complexity of Classical Arabic and the semantic richness of religious texts. In this paper, we propose a novel two-stage framework that addresses both passage retrieval and answer extraction. For passage retrieval, we ensemble finetuned Arabic language models to achieve superior ranking performance. For answer extraction, we employ instruction-tuned large language models with few-shot prompting to overcome the limitations of fine-tuning on small datasets. Our approach achieves state-of-the-art results on the Quran QA 2023 Shared Task, with a MAP@10 of 0.3128 and MRR@10 of 0.5763 for retrieval, and a pAP@10 of 0.669 for extraction, substantially outperforming previous methods. These results demonstrate that combining model ensembling and instruction-tuned language models effectively addresses the challenges of low-resource question answering in specialized domains. Mohamed Basem, Islam Oshallah, Ali Hamdi, Khaled B. Shaban, Hozaifa Kassab |
AICCSA | 4 |
| 2025 | EECG: An Efficient and Scalable Blockchain Solution for Securing Two-Way Cryptographic Communications in Smart GridsabstractIn the smart grid, data communication between smart meters and utility servers should be authentic, private, have integrity while being accessible. To mitigate the risks of potential attacks, securing these two-way communications is crucial. Equally important is maintaining near real-time communication and avoiding significant delays when extra security levels are involved. Existing research on smart grids has not simultaneously tackled the issues of security, communication speed, and network scalability. In this work, we propose a novel delay-optimized blockchain solution for securing cryptographic communication between consumers and the utility in a smart grid. Our solution, based on EOS smart contracts, Edge computing, asymmetric Cryptographic functions, and Group signatures ($E E C G$), treats data communication as transactions that are asymmetrically encrypted and signed in groups before being stored on the EOS blockchain, ensuring confidentiality, privacy, availability, and low cost. The use of edge computing reduces the computational burden of smart meters, increases transaction speed, enhances data privacy, and improves scalability. Furthermore, an optimization problem for associating smart meters with edge nodes is formulated to minimize data exchange and processing delays over the blockchain, facilitating near real-time secure data access. Ahmad El-Hajj, Alaa Awad, Mohammed Al-Husseini, Wassim El-Hajj, Hazem M. Hajj, Khaled B. Shaban, Rabih A. Jabr |
AICCSA | 6 |
| 2025 | Balancing Factual Consistency and Diversity in Abstractive Summarization via Model-Agnostic Composite RerankingabstractAbstractive text summarization has achieved remarkable progress with transformer-based models, yet these systems often produce fluent outputs that suffer from factual inconsistency and redundancy, limiting their reliability in realworld use. Existing solutions to improve factual accuracy typically rely on additional training, specialized architectures, or large-scale annotations, which are computationally costly and difficult to deploy in resource-constrained environments. This paper introduces a lightweight, training-free framework for enhancing abstractive summarization by combining decoding diversity with multi-metric reranking. Our method generates candidate summaries using both deterministic beam search and stochastic top-k sampling, then applies a composite scoring function that integrates ROUGE-1, METEOR, and BERTScore to select the most accurate and informative summary. Experiments on the XSum dataset demonstrate consistent improvements, achieving ROUGE-1 =51.7, ROUGE-2 =27.5, and ROUGE-L =42.7, outperforming competitive reranking baselines. A preference-based evaluation further showed that reranked outputs were favored in 71% of cases, confirming that automatic metric gains align with human-perceived quality. The results highlight a practical and resource-efficient solution for improving summarization quality without retraining. Mariam Elewa, Ali Hamdi, Hozaifa Kassab, Khaled B. Shaban |
AICCSA | 4 |
| 2025 | An Ensemble Classification Approach in A Multi-Layered Large Language Model Framework for Disease PredictionabstractSocial telehealth has made remarkable progress in healthcare by allowing patients to post symptoms and participate in medical consultations remotely. Users frequently post symptoms on social media and online health platforms, creating a huge repository of medical data that can be leveraged for disease classification. Large language models (LLMs) such as LLAMA3 and GPT-3.5, along with transformer-based models like BERT, have demonstrated strong capabilities in processing complex medical text. In this study, we evaluate three Arabic medical text preprocessing methods such as summarization, refinement, and Named Entity Recognition (NER) before applying finetuned Arabic transformer models (CAMeLBERT, AraBERT, and AsafayaBERT). To enhance robustness, we adopt a majority voting ensemble that combines predictions from original and preprocessed text representations. This approach achieved the best classification accuracy of 80.56%, thus showing its effectiveness in leveraging various text representations and model predictions to improve the understanding of medical texts. To the best of our knowledge, this is the first work that integrates LLM-based preprocessing with fine-tuned Arabic transformer models and ensemble learning for disease classification in Arabic social telehealth data. Ali Hamdi, Malak Mohamed, Rokaia Emad, Khaled B. Shaban |
AICCSA | 4 |
| 2025 | CAKD: A Confidence-Aware Knowledge Distillation Approach for Building Compact and Efficient LLMsabstractHigh-quality models across various natural language processing tasks, such as summarization and chatbots, often rely on large architectures, making them computationally intensive and challenging to deploy in resource-constrained environments. While knowledge distillation enables smaller student models to approximate the performance of larger teacher models, existing methods frequently encounter significant trade-offs between accuracy and efficiency. Additionally, uncertain predictions from teacher models can negatively impact the student’s learning process. In this paper, we introduce CAKD, a novel approach that optimizes the training of student models by selectively emphasizing the teacher model’s most reliable predictions using confidence scores. By integrating entropybased confidence weighting into the distillation loss, CAKD effectively prioritizes high-confidence samples, resulting in improved performance and efficiency. Our experiments on text summarization (using a BART-based model on the CNN/DM dataset) and chatbot tasks (using Llamabased model on the DailyDialog and PersonaChat datasets) demonstrate that CAKD achieves significant performance gains over larger teacher models, with improvements of 10.53, 2.1 and 0.38 ROUGE-L points respectively. Mohammad Basheer Kotit, Omama Hamad, Khaled B. Shaban, Ali Hamdi |
AICCSA | 3 |
| 2025 | Attentional Trajectory Modeling for Text-to-3D Generation with Gaussian Multi-View Diffusion and SDS++abstractThe advancement in converting text into 3D scenes has driven significant improvements in generating realistic and adaptable 3D models. However, existing methods face persistent challenges, including inconsistent multi-view generation, limited scene complexity, and an inability to handle real-world datasets with varying camera trajectories. To address these limitations, we introduce a novel approach utilizing a four-part system: the Cinematographer (Trajectory Diffusion Transformer - Traj-DiT), Decorator (Gaussian-driven Multi-view Latent Diffusion Model - GM-LDM, and Detailer (SDS++ loss). Our model enhances 3D scene generation by aligning 3D Gaussians with pixel data, refining 3D structures, and applying realistic surface properties while ensuring view-to-view consistency and accommodating complex scenes. Our research methodology integrates dense-view trajectories processed through BERT, employing multi-head selfattention to handle intricate, real-world camera movements. We conducted extensive experimental comparisons with state-of-theart models, including DreamFusion, Magic3D, LatentNeRF, SJC, Fantasia3D, ProlificDreamer, and Director3D, using BRISQUE, NIQE, and CLIP-Score metrics. Our approach achieved a BRISQUE score of 23.3, NIQE score of 4.34, and CLIP-Score score of 86.1, significantly outperforming all competing methods. These results demonstrate our model’s superior visual clarity, multi-view consistency, geometric accuracy, and photo-realistic rendering. This work represents a substantial advancement in text-to-3D generation, with promising applications in gaming, simulation, and virtual reality. Marena Anis Labib, Ali Hamdi, Khaled B. Shaban |
AICCSA | 3 |
| 2025 | MultiFuzz: A Dense Retrieval-based Multi-Agent System for Network Protocol FuzzingabstractTraditional protocol fuzzing techniques, such as those employed by AFL-based systems, often lack effectiveness due to a limited semantic understanding of complex protocol grammars and rigid seed mutation strategies. Recent works, such as ChatAFL, have integrated Large Language Models (LLMs) to guide protocol fuzzing and address these limitations, pushing protocol fuzzers to wider exploration of the protocol state space. But ChatAFL still faces issues like unreliable output, LLM hallucinations, and assumptions of LLM knowledge about protocol specifications. This paper introduces MultiFuzz, a novel dense retrieval-based multi-agent system designed to overcome these limitations by integrating semantic-aware context retrieval, specialized agents, and structured tool-assisted reasoning. MultiFuzz utilizes agentic chunks of protocol documentation (RFC Documents) to build embeddings in a vector database for a retrieval-augmented generation (RAG) pipeline, enabling agents to generate more reliable and structured outputs, enhancing the fuzzer in mutating protocol messages with enhanced state coverage and adherence to syntactic constraints. The framework decomposes the fuzzing process into modular groups of agents that collaborate through chain-of-thought reasoning to dynamically adapt fuzzing strategies based on the retrieved contextual knowledge. Experimental evaluations on the Real-Time Streaming Protocol (RTSP) demonstrate that MultiFuzz significantly improves branch coverage and explores deeper protocol states and transitions over state-of-the-art (SOTA) fuzzers such as NSFuzz, AFLNet, and ChatAFL. By combining dense retrieval, agentic coordination, and language model reasoning, MultiFuzz establishes a new paradigm in autonomous protocol fuzzing, offering a scalable and extensible foundation for future research in intelligent agentic-based fuzzing systems. Youssef Maklad, Fares Wael, Ali Hamdi, Wael Elsersy, Khaled B. Shaban |
AICCSA | 5 |
| 2025 | Weather-Aware Transformer for Real-Time Route Optimization in Drone-as-a-Service OperationsabstractThis paper presents a novel framework to accelerate route prediction in Drone-as-a-Service operations through weather-aware deep learning models. While classical pathplanning algorithms, such as $\mathrm{A}^{*}$ and Dijkstra, provide optimal solutions, their computational complexity limits real-time applicability in dynamic environments. We address this limitation by training machine learning and deep learning models on synthetic datasets generated from classical algorithm simulations. Our approach incorporates transformer-based and attention-based architectures that utilize weather heuristics to predict optimal nextnode selections while accounting for meteorological conditions affecting drone operations. The attention mechanisms dynamically weight environmental factors including wind patterns, wind bearing, and temperature to enhance routing decisions under adverse weather conditions. Experimental results demonstrate that our weather-aware models achieve significant computational speedup over traditional algorithms while maintaining route optimization performance, with transformer-based architectures showing superior adaptation to dynamic environmental constraints. The proposed framework enables real-time, weatherresponsive route optimization for large-scale DaaS operations, representing a substantial advancement in the efficiency and safety of autonomous drone systems. Kamal Mohamed, Lillian Wassim, Ali Hamdi, Khaled B. Shaban |
AICCSA | 4 |
| 2025 | Efficient Segmentation of Solar Panel Defects Using Knowledge DistillationabstractEfficient identification of anomalies in solar panels is essential for ensuring optimal energy generation and long-term system reliability. This work introduces a two-stage automated framework leveraging drone-captured infrared imagery to detect and segment common defects such as hotspots, dirt accumulation, and shadowing. In the first stage, multiple YOLO-based detection models were evaluated to localize defective regions. YOLOv10 emerged as the best-performing model, achieving a mAP50 of $85 \%$. The detected regions were then used to guide the segmentation stage. A combination of advanced data augmentation techniques (e.g., flipping, lighting variations, and rotation) and a knowledge distillation strategy - where YOLOv11-Seg (student) learned from YOLOv8-Seg (teacher) - significantly improved segmentation accuracy.The final results of the evaluation avaluation of the test set demonstrated strong segmentation performance. The model achieved a class-averaged mask mAP50 of 0.891, with individual class results as follows: 0.975 for Serious Hot Spot, 0.931 for Slight Hot Spot, and 0.766 for Dirt. These results highlight the effectiveness of combining detection-guided segmentation with targeted training enhancements for real-world solar panel inspection. Shahd Tarek, Ali Hamdi, Khaled B. Shaban |
AICCSA | 3 |
| 2025 | MSLEF: Multi-Segment LLM Ensemble Finetuning in RecruitmentabstractThis paper presents MSLEF, a multi-segment ensemble framework that employs LLM fine-tuning to enhance resume parsing in recruitment automation. It integrates finetuned Large Language Models (LLMs) using weighted voting, with each model specializing in a specific resume segment to boost accuracy. Building on MLAR [1], MSLEF introduces a segmentaware architecture that leverages field-specific weighting tailored to each resume part, effectively overcoming the limitations of single-model systems by adapting to diverse formats and structures. The framework incorporates Gemini-2.5-Flash LLM as a high-level aggregator for complex sections and utilizes Gemma 9B, LLaMA 3.1 8B, and Phi-4 14B. MSLEF achieves significant improvements in Exact Match (EM), F1 score, BLEU, ROUGE, and Recruitment Similarity (RS) metrics, outperforming the best single model by up to +7 % in RS. Its segment-aware design enhances generalization across varied resume layouts, making it highly adaptable to real-world hiring scenarios while ensuring precise and reliable candidate representation. Omar Walid, Mohamed T. Younes, Khaled B. Shaban, Mai Hassan, Ali Hamdi |
AICCSA | 3 |
| 2025 | Augmented Fine-Tuned LLMs for Enhanced Recruitment AutomationabstractThis paper presents a novel approach to recruitment automation. Large Language Models (LLMs) were fine-tuned to improve accuracy and efficiency. Building upon our previous work on the Multilayer Large Language Model-Based Robotic Process Automation Applicant Tracking (MLAR) system [1]. This work introduces a novel methodology. Training fine-tuned LLMs specifically tuned for recruitment tasks. The proposed framework addresses the limitations of generic LLMs by creating a synthetic dataset that uses a standardized JSON format. This helps ensure consistency and scalability. In addition to the synthetic data set, the resumes were parsed using DeepSeek, a high-parameter LLM. The resumes were parsed into the same structured JSON format and placed in the training set. This will help improve data diversity and realism. Through experimentation, we demonstrate significant improvements in performance metrics, such as exact match, F1 score, BLEU score, ROUGE score, and overall similarity compared to base models and other state-of-the-art LLMs. In particular, the fine-tuned Phi-4 model achieved the highest F1 score of $90.62 \%$, indicating exceptional precision and recall in recruitment tasks. This study highlights the potential of fine-tuned LLMs. Furthermore, it will revolutionize recruitment workflows by providing more accurate candidate-job matching. Mohamed T. Younes, Omar Walid, Khaled B. Shaban, Ali Hamdi, Mai Hassan |
AICCSA | 3 |
| 2025 | LexiSem: A re-ranker balancing lexical and semantic quality for enhanced abstractive summarizationabstractSequence-to-sequence neural networks have recently achieved significant success in abstractive summarization, especially through fine-tuning large pre-trained language models on downstream datasets. However, these models frequently suffer from exposure bias, which can impair their performance. To address this, re-ranking systems have been introduced, but their potential remains underexplored despite some demonstrated performance gains. Most prior work relies on ROUGE scores and aligned candidate summaries for ranking, exposing a substantial gap between semantic similarity and lexical overlap metrics. In this study, we demonstrate that a second-stage model can be trained to re-rank a set of summary candidates, significantly enhancing performance. Our novel approach leverages a re-ranker that balance lexical and semantic quality. Additionally, we introduce a new strategy for defining negative samples in ranking models. Through experiments on the CNN/DailyMail, XSum and Reddit TIFU datasets, we show that our method effectively estimates the semantic content of summaries without compromising lexical quality. In particular, our method sets a new performance benchmark on the CNN/DailyMail dataset (48.18 R1, 24.46 R2, 45.05 RL) and on Reddit TIFU (30.37 R1,RL 23.87). Eman Aloraini, Hozaifa Kassab, Ali Hamdi, Khaled B. Shaban |
Neurocomputing | 4 |
| 2024 | ASEM: Enhancing Empathy in Chatbot through Attention-based Sentiment and Emotion ModelingabstractEffective feature representations play a critical role in enhancing the performance of text generation models that rely on deep neural networks. However, current approaches suffer from several drawbacks, such as the inability to capture the deep semantics of language and sensitivity to minor input variations, resulting in significant changes in the generated text. In this paper, we present a novel solution to these challenges by employing a mixture of experts, multiple encoders, to offer distinct perspectives on the emotional state of the user’s utterance while simultaneously enhancing performance. We propose an end-to-end model architecture called ASEM that performs emotion analysis on top of sentiment analysis for open-domain chatbots, enabling the generation of empathetic responses that are fluent and relevant. In contrast to traditional attention mechanisms, the proposed model employs a specialized attention strategy that uniquely zeroes in on sentiment and emotion nuances within the user’s utterance. This ensures the generation of context-rich representations tailored to the underlying emotional tone and sentiment intricacies of the text. Our approach outperforms existing methods for generating empathetic embeddings, providing empathetic and diverse responses. The performance of our proposed model significantly exceeds that of existing models, enhancing emotion detection accuracy by 6.2% and lexical diversity by 1.4%. ASEM code is released at https://github.com/MIRAH-Official/Empathetic-Chatbot-ASEM.git Omama Hamad, Khaled B. Shaban, Ali Hamdi |
LREC/COLING | 2 |
| 2024 | Optimal operation of reverse osmosis desalination process with deep reinforcement learning methodsabstractAbstract The reverse osmosis (RO) process is a well-established desalination technology, wherein energy-efficient techniques and advanced process control methods significantly reduce production costs. This study proposes an optimal real-time management method to minimize the total daily operation cost of an RO desalination plant, integrating a storage tank system to meet varying daily freshwater demand. Utilizing the dynamic model of the RO process, a cascade structure with two reinforcement learning (RL) agents, namely the deep deterministic policy gradient (DDPG) and deep Q-Network (DQN), is developed to optimize the operation of the RO plant. The DDPG agent, manipulating the high-pressure pump, controls the permeate flow rate to track a reference setpoint value. Simultaneously, the DQN agent selects the optimal setpoint value and communicates it to the DDPG controller to minimize the plant’s operation cost. Monitoring storage tanks, permeate flow rates, and water demand enables the DQN agent to determine the required amount of permeate water, optimizing water quality and energy consumption. Additionally, the DQN agent monitors the storage tank’s water level to prevent overflow or underflow of permeate water. Simulation results demonstrate the effectiveness and practicality of the designed RL agents. Arash Golabi, Abdelkarim Erradi, Hazim Qiblawey, Ashraf Tantawy, Ahmed Ben Said, Khaled B. Shaban |
Appl. Intell. | 6 |
| 2023 | ECC: Enhancing Smart Grid Communication with Ethereum Blockchain, Asymmetric Cryptography, and Cloud ServicesabstractSmart grids are suscceptible to security vulnerabilities of cyber-physical systems due to the heterogeneity of their interconnected components. There are high risks associated with potential attacks targeting the two-way communication between the smart meters and the utility servers. It is vital to ensure that data communicated between consumers and the utility is not tampered with and is authentic, private, and available. Conventional security measures in traditional communication and network systems fail to secure the data communication aspect in the complex network that composes the advanced metering infrastructure (AMI). In this work, we propose ECC: a novel prevention approach based on Ethereum smart contracts, asymmetric cryptographic functions, and cloud services for securing the two-way communication between smart meters and utility servers. Ethereum blockchain is utilized as a building block where communicated data is treated as transactions encrypted and stored in a distributed fashion to ensure data availability, confidentiality, and privacy. We also augment the Ethereum architecture with cloud services to extend the number of allowable transactions, ensure the availability of the electricity data, and reduce the cost associated with Ethereum transactions. The conducted experiments illustrate the efficacy of ECC in terms of the achieved security properties. This paper shows that the Ethereum Blockchain coupled with Cloud services can improve the efficiency of a system solely based on the Ethereum Blockchain. Raphaelle Akhras, Wassim El-Hajj, Hazem M. Hajj, Khaled B. Shaban, Rabih Jaber |
DSAA | 4 |
| 2023 | Reliable Federated Learning for Age Sensitive Mobile Edge Computing SystemsabstractThe conventional approach for Federated Learning (FL) is to train a global model by averaging local models trained on local data sets. However, given the limited computing resources at the mobile-edge nodes, unreliable models may be received from the Edge Nodes (ENs), which can lead to a significant performance degradation in the FL. Thus, this paper proposes a reliable and age sensitive FL framework that captures the dynamic nature of the local data and computing resources at each participating EN. Specifically, we formulate two optimization problems to select the optimal subset of ENs that can upload their local models in each round of the global model training, given a limited learning cost budget. The first problem aims at selecting the most reliable ENs that should cooperate to complete the FL process, while considering stationary data distributions at different ENs. The second problem aims at minimizing the average age of information experienced by each EN while selecting the most reliable ENs, given fast changing data distributions. Efficient solutions are proposed for the two problems with a worst-case linear complexity. Our Results, leveraging a real-world dataset, depict the efficiency of our solutions in obtaining a better performance compared to conventional FL approach. Alaa Awad, Mhd Saria Allahham, Noor Khial, Amr Mohamed 0001, Aiman Erbad, Khaled B. Shaban |
ICC | 6 |
| 2023 | LIME: Long-Term Forecasting Model for Desalination Membrane Fouling to Estimate the Remaining Useful Life of Membrane
Sohaila Eltanbouly, Abdelkarim Erradi, Ashraf Tantawy, Ahmed Ben Said, Khaled B. Shaban, Hazim Qiblawey |
IEA/AIE (2) | 5 |
| 2023 | Open-Domain Response Generation in Low-Resource Settings using Self-Supervised Pre-Training of Warm-Started TransformersabstractLearning response generation models constitute the main component of building open-domain dialogue systems. However, training open-domain response generation models requires large amounts of labeled data and pre-trained language generation models that are often nonexistent for low-resource languages. In this article, we propose a framework for training open-domain response generation models in low-resource settings. We consider Dialectal Arabic (DA) as a working example. The framework starts by warm-starting a transformer-based encoder-decoder with pre-trained language model parameters. Next, the resultant encoder-decoder model is adapted to DA by employing self-supervised pre-training on large-scale unlabeled data in the desired dialect. Finally, the model is fine-tuned on a very small labeled dataset for open-domain response generation. The results show significant performance improvements on three spoken Arabic dialects after adopting the framework’s three stages, highlighted by higher BLEU and lower Perplexity scores compared with multiple baseline models. Specifically, our models are capable of generating fluent responses in multiple dialects with an average human-evaluated fluency score above 4. Our data is made publicly available. Tarek Naous, Zahraa Bassyouni, Basel Mousi, Hazem M. Hajj, Wassim El-Hajj, Khaled B. Shaban |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2023 | Metadial: A Meta-learning Approach for Arabic Dialogue GenerationabstractDialogue generation is the automatic generation of a text response, given a user’s input. Dialogue generation for low-resource languages has been a challenging tasks for researchers. However, the advancements in deep learning models have made developing conversational agents that perform the tasks of dialogue generation not only possible, but also effective and helpful in many applications spanning a variety of domains. Nevertheless, work on conversational bots for low-resource languages such as the Arabic language is still limited due to various challenges, including the language structure, vocabulary, and the scarcity of its data resources. Meta-learning has been introduced before in the natural language processing (NLP) realm and showed significant improvements in many tasks; however, it has rarely been used in natural language generation (NLG) tasks and never in Arabic NLG. In this work, we propose a meta-learning approach for Arabic dialogue generation for fast adaptation on low-resource domains, namely, Arabic. We start by using existing pre-trained models; we then meta-learn the initial parameters on high-resource dataset before finetuning the parameters on the target tasks. We prove that the proposed model that employs meta-learning techniques improves generalization and enables fast adaptation of the transformer model on low-resource NLG tasks. We report gains in the BLEU-4 and improvements in Semantic textual Similarity (STS) metrics when compared to the existing state-of-the-art approach. We also do a further study on the effectiveness of the meta-learning algorithms on the response generation of the models. Mohsen Shamas, Wassim El-Hajj, Hazem M. Hajj, Khaled B. Shaban |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | Automated Generation of Human-readable Natural Arabic Text from RDF DataabstractWith the advances in Natural Language Processing (NLP), the industry has been moving towards human-directed artificial intelligence (AI) solutions. Recently, chatbots and automated news generation have captured a lot of attention. The goal is to automatically generate readable text from tabular data or web data commonly represented in Resource Description Framework (RDF) format. The problem can then be formulated as Data-to-text (D2T) generation from structured non-linguistic data into human-readable natural language. Despite the significant work done for the English language, no efforts are being directed towards low-resource languages like the Arabic language. This work promotes the development of the first RDF data-to-text (D2T) generation system for the Arabic language while trying to address the low-resource limitation. We develop several models for the Arabic D2T task using transfer learning from large language models (LLM) such as AraBERT, AraGPT2, and mT5. These models include a baseline Bi-LSTM Sequence-to-Sequence (Seq2Seq) model, as well as encoder-decoder transformers like BERT2BERT, BERT2GPT, and T5. We then provide a detailed comparative study highlighting the strengths and limitations of these methods setting the stage for further advancement in the field. We also introduce a new Arabic dataset (AraWebNLG) that can be used for new model development in the field. To ensure a comprehensive evaluation, general-purpose automated metrics (BLEU and Perplexity scores) are used as well as task-specific human evaluation metrics related to the accuracy of the content selection and fluency of the generated text. The results highlight the importance of pre-training on a large corpus of Arabic data and show that transfer learning from AraBERT gives the best performance. Text-to-text pre-training using mT5 achieves second best performance results even with multilingual weights. Roudy Touma, Hazem M. Hajj, Wassim El-Hajj, Khaled B. Shaban |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2021 | MARL: Multimodal Attentional Representation Learning for Disease Prediction
Ali Hamdi, Amr Aboeleneen, Khaled B. Shaban |
ICVS | 3 |
| 2020 | Securing Smart Grid Communication using Ethereum Smart ContractsabstractSmart grids are being continually adopted as a replacement of the traditional power grid systems to ensure safe, efficient, and cost-effective power distribution. The smart grid is a heterogeneous communication network made up of various devices such as smart meters, automation, and emerging technologies interacting with each other. As a result, the smart grid inherits most of the security vulnerabilities of cyber systems, putting the smart grid at risk of cyber-attacks. To secure the communication between smart grid entities, namely the smart meters and the utility, we propose in this paper a communication infrastructure built on top of a blockchain network, specifically Ethereum. All two-way communication between the smart meters and the utility is assumed to be transactions governed by smart contracts. Smart contracts are designed in such a way to ensure that each smart meter is authentic and each smart meter reading is reported securely and privately. We present a simulation of a sample smart grid and report all the costs incurred from building such a grid. The simulations illustrate the feasibility and security of the proposed architecture. They also point to weaknesses that must be addressed, such as scalability and cost. Raphaelle Akhras, Wassim El-Hajj, Michel Majdalani, Hazem M. Hajj, Rabih A. Jabr, Khaled B. Shaban |
IWCMC | 6 |
| 2020 | Model-based risk assessment for cyber physical systems security
Ashraf Tantawy, Sherif Abdelwahed, Abdelkarim Erradi, Khaled B. Shaban |
Comput. Secur. | 4 |
| 2019 | Data-driven Curation, Learning and Analysis for Inferring Evolving IoT Botnets in the WildabstractThe insecurity of the Internet-of-Things (IoT) paradigm continues to wreak havoc in consumer and critical infrastructure realms. Several challenges impede addressing IoT security at large, including, the lack of IoT-centric data that can be collected, analyzed and correlated, due to the highly heterogeneous nature of such devices and their widespread deployments in Internet-wide environments. To this end, this paper explores macroscopic, passive empirical data to shed light on this evolving threat phenomena. This not only aims at classifying and inferring Internet-scale compromised IoT devices by solely observing such one-way network traffic, but also endeavors to uncover, track and report on orchestrated "in the wild" IoT botnets. Initially, to prepare the effective utilization of such data, a novel probabilistic model is designed and developed to cleanse such traffic from noise samples (i.e., misconfiguration traffic). Subsequently, several shallow and deep learning models are evaluated to ultimately design and develop a multi-window convolution neural network trained on active and passive measurements to accurately identify compromised IoT devices. Consequently, to infer orchestrated and unsolicited activities that have been generated by well-coordinated IoT botnets, hierarchical agglomerative clustering is deployed by scrutinizing a set of innovative and efficient network feature sets. By analyzing 3.6 TB of recent darknet traffic, the proposed approach uncovers a momentous 440,000 compromised IoT devices and generates evidence-based artifacts related to 350 IoT botnets. While some of these detected botnets refer to previously documented campaigns such as the Hide and Seek, Hajime and Fbot, other events illustrate evolving threats such as those with cryptojacking capabilities and those that are targeting industrial control system communication and control services. Morteza Safaei Pour, Antonio Mangino, Kurt Friday, Matthias Rathbun, Elias Bou-Harb, Farkhund Iqbal, Khaled B. Shaban, Abdelkarim Erradi |
ARES | 7 |
| 2019 | A Survey of Opinion Mining in Arabic: A Comprehensive System Perspective Covering Challenges and Advances in Tools, Resources, Models, Applications, and VisualizationsabstractOpinion-mining or sentiment analysis continues to gain interest in industry and academics. While there has been significant progress in developing models for sentiment analysis, the field remains an active area of research for many languages across the world, and in particular for the Arabic language, which is the fifth most-spoken language and has become the fourth most-used language on the Internet. With the flurry of research activity in Arabic opinion mining, several researchers have provided surveys to capture advances in the field. While these surveys capture a wealth of important progress in the field, the fast pace of advances in machine learning and natural language processing (NLP) necessitates a continuous need for a more up-to-date literature survey. The aim of this article is to provide a comprehensive literature survey for state-of-the-art advances in Arabic opinion mining. The survey goes beyond surveying previous works that were primarily focused on classification models. Instead, this article provides a comprehensive system perspective by covering advances in different aspects of an opinion-mining system, including advances in NLP software tools, lexical sentiment and corpora resources, classification models, and applications of opinion mining. It also presents future directions for opinion mining in Arabic. The survey also covers latest advances in the field, including deep learning advances in Arabic Opinion Mining. The article provides state-of-the-art information to help new or established researchers in the field as well as industry developers who aim to deploy an operational complete opinion-mining system. Key insights are captured at the end of each section for particular aspects of the opinion-mining system giving the reader a choice of focusing on particular aspects of interest. Gilbert Badaro, Ramy Baly, Hazem M. Hajj, Wassim El-Hajj, Khaled B. Shaban, Nizar Habash, Ahmad A. Al Sallab, Ali Hamdi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2018 | Passive inference of attacks on CPS communication protocols
Elias Bou-Harb, Nasir Ghani, Abdelkarim Erradi, Khaled B. Shaban |
J. Inf. Secur. Appl. | 4 |
| 2018 | CLASENTI: A Class-Specific Sentiment Analysis FrameworkabstractArabic text sentiment analysis suffers from low accuracy due to Arabic-specific challenges (e.g., limited resources, morphological complexity, and dialects) and general linguistic issues (e.g., fuzziness, implicit sentiment, sarcasm, and spam). The limited resources problem requires efforts to build new and improved Arabic corpora and lexica. We propose a class-specific sentiment analysis (CLASENTI) framework. The framework includes a new annotation approach to build multi-faceted Arabic corpus and lexicon allowing for simultaneous annotation of different facets, including domains, dialects, linguistic issues, and polarity strengths. Each of these facets has multiple classes (e.g., the nine classes representing dialects found in the Arab world). The new corpus and lexicon annotations facilitate the development of new class-specific classification models and polarity strength calculation. For the new sentiment classification models, we propose a hybrid model combining corpus-based and lexicon-based models. The corpus-based model has two interrelated phases to build; (1) full-corpus classification models for all facets; and (2) class-specific models trained on filtered subsets of the corpus according to the performances of the full-corpus models. To calculate polarity strengths, the lexicon-based model filters the annotated lexicon based on the specific classes of the domain and dialect. As a case study, we collect and annotate 15274 reviews from various sources, including surveys, Facebook comments, and Twitter posts, pertaining to governmental services. In addition, we develop a new web-based application to apply the proposed framework on the case study. CLASENTI framework reaches up to 95% accuracy and 93% F1-Score surpassing the best-known sentiment classifiers implemented in Scikit-learn library that achieve 82% accuracy and 81% F1-Score for Arabic when tested on the same dataset. Ali Hamdi, Khaled B. Shaban, Anazida Zainal |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2017 | Modelling and performance analysis of TCP variants for data collection in smart power grids
Tarek Khalifa, Khaled B. Shaban, Atef Abdrabou, Abdoulmenim Bilh, Sagar Naik |
Comput. Commun. | 2 |
| 2017 | Post-failure repair for cloud-based infrastructure services after disasters
Mahsa Pourvali, Cicek Cavdar, Khaled B. Shaban, Jorge Crichigno, Nasir Ghani |
Comput. Commun. | 3 |
| 2017 | Restoration methods for cloud multicast virtual networks
Sara Ayoubi, Chadi Assi, Yiheng Chen, Tarek Khalifa, Khaled B. Shaban |
J. Netw. Comput. Appl. | 5 |
| 2017 | A Sentiment Treebank and Morphologically Enriched Recursive Deep Models for Effective Sentiment Analysis in ArabicabstractAccurate sentiment analysis models encode the sentiment of words and their combinations to predict the overall sentiment of a sentence. This task becomes challenging when applied to morphologically rich languages (MRL). In this article, we evaluate the use of deep learning advances, namely the Recursive Neural Tensor Networks (RNTN), for sentiment analysis in Arabic as a case study of MRLs. While Arabic may not be considered the only representative of all MRLs, the challenges faced and proposed solutions in Arabic are common to many other MRLs. We identify, illustrate, and address MRL-related challenges and show how RNTN is affected by the morphological richness and orthographic ambiguity of the Arabic language. To address the challenges with sentiment extraction from text in MRL, we propose to explore different orthographic features as well as different morphological features at multiple levels of abstraction ranging from raw words to roots. A key requirement for RNTN is the availability of a sentiment treebank; a collection of syntactic parse trees annotated for sentiment at all levels of constituency and that currently only exists in English. Therefore, our contribution also includes the creation of the first Arabic Sentiment Treebank (A r S en TB) that is morphologically and orthographically enriched. Experimental results show that, compared to the basic RNTN proposed for English, our solution achieves significant improvements up to 8% absolute at the phrase level and 10.8% absolute at the sentence level, measured by average F1 score. It also outperforms well-known classifiers including Support Vector Machines, Recursive Auto Encoders, and Long Short-Term Memory by 7.6%, 3.2%, and 1.6% absolute respectively, all models being trained with similar morphological considerations. Ramy Baly, Hazem M. Hajj, Nizar Habash, Khaled B. Shaban, Wassim El-Hajj |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2017 | AROMA: A Recursive Deep Learning Model for Opinion Mining in Arabic as a Low Resource LanguageabstractWhile research on English opinion mining has already achieved significant progress and success, work on Arabic opinion mining is still lagging. This is mainly due to the relative recency of research efforts in developing natural language processing (NLP) methods for Arabic, handling its morphological complexity, and the lack of large-scale opinion resources for Arabic. To close this gap, we examine the class of models used for English and that do not require extensive use of NLP or opinion resources. In particular, we consider the Recursive Auto Encoder (RAE). However, RAE models are not as successful in Arabic as they are in English, due to their limitations in handling the morphological complexity of Arabic, providing a more complete and comprehensive input features for the auto encoder, and performing semantic composition following the natural way constituents are combined to express the overall meaning. In this article, we propose A R ecursive Deep Learning Model for O pinion M ining in A rabic (AROMA) that addresses these limitations. AROMA was evaluated on three Arabic corpora representing different genres and writing styles. Results show that AROMA achieved significant performance improvements compared to the baseline RAE. It also outperformed several well-known approaches in the literature. Ahmad A. Al Sallab, Ramy Baly, Hazem M. Hajj, Khaled B. Shaban, Wassim El-Hajj, Gilbert Badaro |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2017 | Delay-Aware Flow Scheduling In Low Latency Enterprise Datacenter Networks: Modeling and Performance AnalysisabstractReal-time interactive application workloads (e.g., Web search, social networking, and so on) appear in the form of a large number of mini requests and responses flowing over the datacenters' networks. They end up being sewed all together to constitute a user-requested task or computation (e.g., display a complete Facebook timeline). Applications as such strictly impose low latency flow completion, since the service's quality is decreed by quick aggregation of responses to the largest possible fraction of requests and their delivery back to the user. This paper presents a deadline-aware flow scheduling (DAFS). In addition to reducing the average flow completion time (FCT), DAFS aims at decreasing the deadline mismatch and blocking probabilities, hence improving the average application throughput. An analytical queuing model is formulated herein to capture the datacenter's network dynamics and evaluate its performance when operating under DAFS. The model is validated through extensive simulations whose results also show that DAFS outperforms existing multi-queue-based priority mechanisms by 52% in terms of the average FCT and a range of 7%-29% in terms of the average throughput. Maurice Khabbaz, Khaled B. Shaban, Chadi Assi |
IEEE Trans. Commun. | 2 |
| 2017 | A Reliability-Aware Network Service Chain Provisioning With Delay Guarantees in NFV-Enabled Enterprise Datacenter NetworksabstractTraditionally, service-specific network functions (NFs) (e.g., Firewall, intrusion detection system, etc.) are executed by installation-and maintenance-costly hardware middleboxes that are deployed within a datacenter network following a strictly ordered chain. NF virtualization (NFV) virtualizes these NFs and transforms them into instances of plain software referred to as virtual NFs (VNFs) and executed by virtual machines, which, in turn, are hosted over one or multiple industry-standard physical machines. The failure (e.g., hardware or software) of any one of a service chain's VNFs leads to breaking down the entire chain and causing significant data losses, delays, and resource wastage. This paper establishes a reliability-aware and delay-constrained (READ) routing optimization framework for NFV-enabled datacenter networks. READ encloses the formulation of a complex mixed integer linear program (MILP) whose resolution yields an optimal network service VNF placement and traffic routing policy that jointly maximizes the achieved respective reliabilities of supported network services and minimizes these services' respective end-to-end delays. A heuristic algorithm dubbed Greedy-k-shortest paths (GSP) is proposed for the purpose of overcoming the MILP's complexity and develop an efficient routing scheme whose results are comparable to those of READ's optimal counterparts. Thorough numerical analyses are conducted to evaluate the network's performance under GSP, and hence, gauge its merit; particularly, when compared to existing schemes, GSP exhibits an improvement of 18.5% in terms of the average end-to-end delay as well as 7.4% to 14.8% in terms of reliability. Long Qu, Chadi Assi, Khaled B. Shaban, Maurice Khabbaz |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2017 | Optimizing Cloud-Service Performance: Efficient Resource Provisioning via Optimal Workload AllocationabstractCloud computing is being widely accepted and utilized in the business world. From the perspective of businesses utilizing the cloud, it is critical to meet their customers' requirements by achieving service-level-objectives. Hence, the ability to accurately characterize and optimize cloud-service performance is of great importance. In this paper a stochastic multi-tenant framework is proposed to model the service of customer requests in a cloud infrastructure composed of heterogeneous virtual machines. Two cloud-service performance metrics are mathematically characterized, namely the percentile and the mean of the stochastic response time of a customer request, in closed form. Based upon the proposed multi-tenant framework, a workload allocation algorithm, termed maxmin-cloud algorithm, is then devised to optimize the performance of the cloud service. A rigorous optimality proof of the max-min-cloud algorithm is also given. Furthermore, the resource-provisioning problem in the cloud is also studied in light of the max-min-cloud algorithm. In particular, an efficient resource-provisioning strategy is proposed for serving dynamically arriving customer requests. These findings can be used by businesses to build a better understanding of how much virtual resource in the cloud they may need to meet customers' expectations subject to cost constraints. Zhuoyao Wang 0001, Majeed M. Hayat, Nasir Ghani, Khaled B. Shaban |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Enhancing the performance of post-failure restoration schemes in multi-tenant networksabstractFailure in the physical network can cause a temporal or permanent unavailability of some resources, which can lead to a quality of service (QoS) degradation and loss of revenue. While much work has been dedicated to the survival of delay-constrained applications, little attention has been paid to enhancing the performance of the post-failure recovery scheme by maintaining the performance for the affected applications. In this paper, we introduce a QoS provisioning framework that overcomes the limitations of post-failure recovery techniques. The framework aims at not only fixing the failed applications, but also providing sufficient QoS guarantees for the hosted applications in case of link/node failure while maximizing the resources utilization. Simulation results of a data center hosting multicast applications prove that the proposed method boosts the recovery scheme to achieves better restoration ratio in a considerably fast execution time, and increases the revenue. Abdulaziz M. Ghaleb, Tarek Khalifa, Khaled B. Shaban |
CNSM | 3 |
| 2016 | Reliability-aware service provisioning in NFV-enabled enterprise datacenter networksabstractNetwork Function Visualization (NFV) enables the complete decoupling of Network Functions (NFs) (e.g., firewall, intrusion detection, routing, etc.) from physical middleboxes used to implement service-specific and strictly ordered chains of these NFs. Precisely, NFV allows for dispatching NFs as plain software instances called Virtual Network Functions (VNFs) running on virtual machines hosted by one or more industry standard physical machines. This, however, introduces vulnerabilities (e.g., hard-/soft-ware failures, etc) causing the break down of the entire VNF chain. The functionality of NFV-enabled networks impose higher reliability requirements than traditional networks. This paper encloses an in-depth investigation of a reliability-aware joint VNF placement and flow routing optimization problem. This problem is formulated as a complex Integer Linear Program (ILP). A heuristic is proposed in order to overcome this ILP's complexity. Thorough numerical analysis are conducted to verify and assert the correctness and effectiveness of the proposed heuristic. Long Qu, Chadi Assi, Khaled B. Shaban, Maurice Khabbaz |
CNSM | 3 |
| 2016 | Directed graph-based wireless EEG sensor channel selection approach for cognitive task classificationabstractWireless electroencephalogram (EEG) sensors have been successfully applied in many medical and computer brain interface classifications. A common characteristic of wireless EEG sensors is that they are low powered devices, and hence an efficient usage of sensor energy resources is critical for any practical application. One way of minimizing energy consumption by the EEG sensors is by reducing the number of EEG channels participating in the classification process. For the purpose of classifying EEG signals, we propose a directed acyclic graph (DAG)-based channel selection algorithm. To achieve this objective, the EEG sensor channels are first realized in a complete undirected graph, where each channel is represented by a node. An edge between any two nodes indicates the collaboration between these nodes in identifying the system state; and the significance of this collaboration is quantified by a weight assigned to the edge. The complete graph is then reduced into a directed acyclic graph that encodes the knowledge of the non-increasing order of the channel ranking for each cognitive task. The channel selection algorithm utilizes this directed graph to find a maximum path such that the total weight of this path satisfies a predefined threshold. It has been demonstrated experimentally that channel utilization has been reduced by 50% in the worst case scenario for a three-state system and an EEG sensor with 14 channels; and the best classification accuracy obtained is 81%. Abduljalil Mohamed, Khaled B. Shaban, Amr Mohamed 0001 |
IWCMC | 2 |
| 2016 | Arabic Corpora for Credibility Analysis
Ayman Al Zaatari, Rim El Ballouli, Shady Elbassuoni, Wassim El-Hajj, Hazem M. Hajj, Khaled B. Shaban, Nizar Habash, Emad Yahya |
LREC | 6 |
| 2016 | Surviving link failures in multicast VN embedded applicationsabstractVirtual network embedding (VNE) is defined as the allocation of network resources to multiple virtual networks (VNs) and is recognized to be a challenging task to perform efficiently. Virtual network survivability is a new term that describes the measures taken to provide a failure-proof VN against physical link and/or node failure. Indeed, a single link or node failure in a substrate network can bring down multiple hosted VNs, i.e., the ones that utilize that failed link or node. As such, virtual network survivability becomes an essential part of VNE. While much work has been dedicated to studying the impact of a variety of failure cases in a VN, little attention has been directed towards studying the link failure impact on multicast virtual network (MVN) applications, which principally restrict end-to-end delay and delay variation measures. In fact, most of the introduced survivability schemes adopt protection techniques by reserving backup resources prior to embedding, which inevitably leads to under-utilization of the network resources. In this paper, we first investigate the impact of physical link failure on MVNs. Then, we introduce a novel recovery approach to restore MVNs while considering their end-delay and delay variation requirements. Simulation experiments prove that our recovery technique achieves good restoration ratio in considerably fast execution time and low link mapping cost with little impact on the admittance ratio. Abdulaziz M. Ghaleb, Tarek Khalifa, Sara Ayoubi, Khaled B. Shaban |
NOMS | 4 |
| 2016 | Network function virtualization scheduling with transmission delay optimizationabstractTo accelerate the implementation of network functions/middle boxes and reduce the deployment cost, recently the concept of Network Function Virtualization (NFV) has emerged and became a topic of much interest attracting the attention of researchers from both industry and academia. Unlike the traditional implementation of network functions, a software-oriented approach for network functions create more flexible and dynamic network services to meet a more diversified demand. In this paper, we study the Virtual Network Function (VNF) chaining scheduling problem with limited network resources. We consider VNF transmission and processing delays, and formulate the VNFs chaining scheduling as a new Mixed Integer Linear Programming (MILP) problem. Our objective is to minimize the latency of the overall VNFs' schedule. Reducing the scheduling latency enables cloud operators to service (and admit) more customers, thereby increasing operators' revenues. Owing to the complexity of the problem, we develop a Genetic Algorithm (GA) based method for solving the problem efficiently. Finally, the effectiveness of our heuristic algorithm is verified through numerical results. Long Qu, Chadi Assi, Khaled B. Shaban |
NOMS | 3 |
| 2016 | Overlay network scheduling design
Khaled B. Shaban, Mahmoud A. Khodeir, Jorge Crichigno, Samee Ullah Khan, Nasir Ghani |
Comput. Commun. | 2 |
| 2016 | Delay-Aware Scheduling and Resource Optimization With Network Function VirtualizationabstractTo accelerate the implementation of network functions/middle boxes and reduce the deployment cost, recently, the concept of network function virtualization (NFV) has emerged and become a topic of much interest attracting the attention of researchers from both industry and academia. Unlike the traditional implementation of network functions, a software-oriented approach for virtual network functions (VNFs) creates more flexible and dynamic network services to meet a more diversified demand. Software-oriented network functions bring along a series of research challenges, such as VNF management and orchestration, service chaining, VNF scheduling for low latency and efficient virtual network resource allocation with NFV infrastructure, among others. In this paper, we study the VNF scheduling problem and the corresponding resource optimization solutions. Here, the VNF scheduling problem is defined as a series of scheduling decisions for network services on network functions and activating the various VNFs to process the arriving traffic. We consider VNF transmission and processing delays and formulate the joint problem of VNF scheduling and traffic steering as a mixed integer linear program. Our objective is to minimize the makespan/latency of the overall VNFs' schedule. Reducing the scheduling latency enables cloud operators to service (and admit) more customers, and cater to services with stringent delay requirements, thereby increasing operators' revenues. Owing to the complexity of the problem, we develop a genetic algorithm-based method for solving the problem efficiently. Finally, the effectiveness of our heuristic algorithm is verified through numerical evaluation. We show that dynamically adjusting the bandwidths on virtual links connecting virtual machines, hosting the network functions, reduces the schedule makespan by 15%-20% in the simulated scenarios. Long Qu, Chadi Assi, Khaled B. Shaban |
IEEE Trans. Commun. | 3 |
| 2016 | Surviving Multiple Failures in Multicast Virtual Networks With Virtual Machines MigrationabstractThis paper deals with the multiple link/node substrate failures that impact a multicast virtual network (MVN) in which link recovery is not feasible and node migration is mandatory. A novel restoration approach is introduced to repair the failed MVNs while maintaining their quality of service requirements (e.g., end-to-end delay and delay variations). This approach relies on reducing the search region and exploiting nodes ranking and filtering (NRF) techniques to speed up the recovery process of finding an alternative node to which to migrate. The performance is extensively evaluated against multiple failures, with and without NRF, compared with complete re-embedding technique, link failure algorithms for single link failure, and previous work for single node failure. Simulation results prove that our recovery technique achieves good restoration ratio in considerably fast execution time, low link mapping cost (gain) with a slight impact on the admission ratio. Abdulaziz M. Ghaleb, Tarek Khalifa, Sara Ayoubi, Khaled B. Shaban, Chadi Assi |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2016 | A Meta-Framework for Modeling the Human Reading Process in Sentiment AnalysisabstractThis article introduces a sentiment analysis approach that adopts the way humans read, interpret, and extract sentiment from text. Our motivation builds on the assumption that human interpretation should lead to the most accurate assessment of sentiment in text. We call this automated process Human Reading for Sentiment (HRS). Previous research in sentiment analysis has produced many frameworks that can fit one or more of the HRS aspects; however, none of these methods has addressed them all in one approach. HRS provides a meta-framework for developing new sentiment analysis methods or improving existing ones. The proposed framework provides a theoretical lens for zooming in and evaluating aspects of any sentiment analysis method to identify gaps for improvements towards matching the human reading process. Key steps in HRS include the automation of humans low-level and high-level cognitive text processing. This methodology paves the way towards the integration of psychology with computational linguistics and machine learning to employ models of pragmatics and discourse analysis for sentiment analysis. HRS is tested with two state-of-the-art methods; one is based on feature engineering, and the other is based on deep learning. HRS highlighted the gaps in both methods and showed improvements for both. Ramy Baly, Roula Hobeica, Hazem M. Hajj, Wassim El-Hajj, Khaled B. Shaban, Ahmad A. Al Sallab |
ACM Trans. Inf. Syst. | 5 |
| 2015 | Multicast Tree Repair and Maintenance in the CloudabstractNetwork virtualization enables the multi-tenancy concept where multiple tenants's services can cohabit the same substrate network and share its resources. With multi-tenancy, the problem of allocating resources to the various tenants emerges as a challenging problem. This former is commonly known as the virtual network embedding problem (VNE), which has attracted numerous effort from the research industry due to its NP-Hard nature. Yet, most of the existing work overlook the various modes of communication a virtual network (VN) can exhibit, assuming it is always a one-to-one communication between virtual machines (VMs). The recent technological advancements (such as Software Defined Networks (SDNs)) have paved the way for efficient multicast in data center networks, thereby leveraging the support of services and applications which multicast data in large volumes. While much work has been devoted for studying the problem of multicast virtual network (MVN) embedding in the cloud, little attention has been paid to investigating the impact of failure on this service class. In this paper, we study the impact of facility node failure on embedded MVNs, and introduce a novel post-failure restoration scheme to repair failed MVNs while maintaining their requested Quality of Service (QoS). Our numerical results prove that our suggested method achieves encouraging restoration ratio in considerably fast execution time. Sara Ayoubi, Yiheng Chen, Chadi Assi, Tarek Khalifa, Khaled B. Shaban |
CLOUD | 5 |
| 2015 | Survivable Cloud Network Mapping for Disaster Recovery SupportabstractNetwork virtualization is a key provision for improving the scalability and reliability of cloud computing services. In recent years, various mapping schemes have been developed to reserve VN resources over substrate networks. However, many cloud providers are very concerned about improving service reliability under catastrophic disaster conditions yielding multiple system failures. To address this challenge, this work presents a novel failure region-disjoint VN mapping scheme to improve VN mapping survivability. The problem is first formulated as a mixed integer linear programming problem and then two heuristic solutions are proposed to compute a pair of failure region-disjoint VN mappings. The solution also takes into account mapping costs and load balancing concerns to help improve resource efficiencies. The schemes are then analyzed in detail for a variety of networks and their overall performances compared to some existing survivable VN mapping schemes. Khaled B. Shaban, Nasir Ghani, Samee Ullah Khan, Mahshid Rahnamay-Naeini, Majeed M. Hayat, Chadi Assi |
IEEE Trans. Computers | 2 |
| 2015 | MINTED: Multicast VIrtual NeTwork Embedding in Cloud Data Centers With Delay ConstraintsabstractNetwork virtualization is regarded as the pillar of cloud computing, enabling the multi-tenancy concept where multiple Virtual Networks (VNs) can cohabit the same substrate network. With network virtualization, the problem of allocating resources to the various tenants, commonly known as the Virtual Network Embedding problem, emerges as a challenge. Its NP-Hard nature has drawn a lot of attention from the research community, many of which however overlooked the type of communication that a given VN may exhibit, assuming that they all exhibit a one-to-one (unicast) communication only. In this paper, we motivate the importance of characterizing the mode of communication in VN requests, and we focus our attention on the problem of embedding VNs with a one-to-many (multicast) communication mode. Throughout this paper, we highlight the unique properties of multicast VNs and its distinct Quality of Service (QoS) requirements, most notably the end-delay and delay-variation constraints for delay-sensitive multicast services. Further, we showcase the limitations of handling a multicast VN as unicast. To this extent, we formally define the VNE problem for Multicast VNs (MVNs) and prove its NP-Hard nature. We propose two novel approach to solve the Multicast VNE (MVNE) problem with end-delay and delay variation constraints: A 3-Step MVNE technique, and a Tabu-Search algorithm. We motivate the intuition behind our proposed embedding techniques, and provide a competitive analysis of our suggested approaches over multiple metrics and against other embedding heuristics. Sara Ayoubi, Chadi Assi, Khaled B. Shaban, Lata Narayanan |
IEEE Trans. Commun. | 3 |
| 2014 | Multicast Virtual Network Embedding in Cloud Data Centers with Delay ConstraintsabstractNetwork virtualization enables the multi-tenancy concept and paves the way towards more advancements and innovation in the underlying infrastructure. With network virtualization, allocating resources to Virtual Networks (VNs) that represent tenants' requests emerges as a challenging problem. This problem is commonly known as the Virtual Network Embedding (VNE) problem, and its NP-Hard nature has drawn a lot of attention from the research community. A common feature in the existing work is that the type of communication in the VN requests was never characterized, assuming that they exhibit unicast communication only. In this paper, we motivate the importance of characterizing the type of communication in VN requests. We present a formal definition of the VNE problem for VNs with multicast communication. To the best of our knowledge, the multicast VNE problem has not been addressed in the frame of cloud computing, where the location of all the virtual machines in a given multicast VN is unknown. We propose a novel 3-steps heuristic to solve the multicast VNE problem with end-delay and delay variation constraints. Our numerical results prove the efficiency of our suggested approach over multiple metrics and against numerous embedding heuristics. Sara Ayoubi, Khaled B. Shaban, Chadi Assi |
IEEE CLOUD | 2 |
| 2014 | Traffic engineering in cloud data centers: A column generation approachabstractWhile many have advocated for the use of Virtual Local Area Networks (VLANs) as a way to provide scalable traffic management, finding the optimal traffic split (mapping) among VLANs to achieve load balancing has turned out to be a very challenging and combinatorially complex problem to solve. This paper considers the traffic engineering problem in data center networks by studying the joint problem of finding spanning trees for VLANs and optimally selecting the most promising spanning trees to map the traffic flows onto. We mathematically model this problem using Integer Linear Program (ILP) techniques and follow a primal-dual decomposition approach, using column generation, to solve exactly a relaxed mapping version of the problem, as well we present approximate solutions to the original problem. We show through numerical evaluations an outstanding scalability of the decomposed version of the problem and we use our results to study the performance of traffic engineering protocols developed in recent literature for data center networks. Sara Ayoubi, Samir Sebbah, Khaled B. Shaban, Chadi Assi |
NOMS | 3 |
| 2014 | Classification ensemble to improve medical Named Entity RecognitionabstractAn accurate Named Entity Recognition (NER) is important for knowledge discovery in text mining. This paper proposes an ensemble machine learning approach to recognise Named Entities (NEs) from unstructured and informal medical text. Specifically, Conditional Random Field (CRF) and Maximum Entropy (ME) classifiers are applied individually to the test data set from the i2b2 2010 medication challenge. Each classifier is trained using a different set of features. The first set focuses on the contextual features of the data, while the second concentrates on the linguistic features of each word. The results of the two classifiers are then combined. The proposed approach achieves an f-score of 81.8%, showing a considerable improvement over the results from CRF and ME classifiers individually which achieve f-scores of 76% and 66.3% for the same data set, respectively. Sara Keretna, Chee Peng Lim, Douglas C. Creighton, Khaled B. Shaban |
SMC | 4 |
| 2014 | Transport layer performance analysis and optimization for smart metering infrastructure
Tarek Khalifa, Atef Abdrabou, Khaled B. Shaban, Maazen Alsabaan, Sagar Naik |
J. Netw. Comput. Appl. | 3 |
| 2014 | Towards Scalable Traffic Management in Cloud Data CentersabstractCloud Computing is becoming a mainstream paradigm, as organizations, large and small, begin to harness its benefits. This novel technology brings new challenges, mostly in the protocols that govern its underlying infrastructure. Traffic engineering in cloud data centers is one of these challenges that has attracted attention from the research community, particularly since the legacy protocols employed in data centers offer limited and unscalable traffic management. Many advocated for the use of VLANs as a way to provide scalable traffic management, however, finding the optimal traffic split between VLANs is the well known NP-Complete VLAN assignment problem. The size of the search space of the VLAN assignment problem is huge, even for small size networks. This paper introduce a novel decomposition approach to solve the VLAN mapping problem in cloud data centers through column generation. Column generation is an effective technique that is proven to reach optimality by exploring only a small subset of the search space. We introduce both an exact and a semi-heuristic decomposition with the objective to achieve load balancing by minimizing the maximum link load in the network. Our numerical results have shown that our approach explores less than 1% of the available search space, with an optimality gap of at most 4%. We have also compared and assessed the performance of our decomposition model and state of the art protocols in traffic engineering. This comparative analysis proves that our model attains encouraging gain over its peers. Chadi Assi, Sara Ayoubi, Samir Sebbah, Khaled B. Shaban |
IEEE Trans. Commun. | 4 |
| 2012 | Air Quality Monitoring and Prediction System Using Machine-to-Machine Platform
Abdullah Kadri, Khaled B. Shaban, Elias Yaacoub, Adnan A. Abu-Dayya |
ICONIP (4) | 2 |
| 2010 | Credit default swap pricing using artificial neural networksabstractThe credit derivatives market has experienced unprecedented growth over the past few years. As such, there is a growing interest in tools for pricing the most prominent credit derivative, the credit default swap. In this paper, we present several artificial neural networks that predict real-world credit default swap prices. In addition to the input parameters used by analytical pricing strategies, these networks explore the use of historic credit default swap prices and equity prices. It was found that the inclusion of historic parameters has increased the accuracy of the network's prediction of credit default swap prices.. Khaled B. Shaban, Abdunnaser Younes, R. Lam, M. Allison, S. Kathirgamanathan |
IJCNN | 1 |
| 2006 | Document Mining Based on Semantic Understanding of Text
Khaled B. Shaban, Otman A. Basir, Mohamed S. Kamel |
CIARP | 1 |