EDBT 2026 Demo / reviewers in the wild / expert
Ali Hamdi
dblp:203/3857
· DBLP profile ↗
23ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-2301-6588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashDetR: A deep learning pipeline for early detection and time estimation of flashover in high-voltage insulators using infrared videosabstractFlashover in high-voltage insulators poses a significant risk to power system reliability, potentially leading to outages and safety hazards. This study introduces an innovative deep learning-based approach for early prediction of flashover events and time-to-flashover estimation by analyzing infrared videos of dry band arcing, a known precursor to flashover. In this work, we propose a pipeline named Flashover Detector and Time Estimator , which integrates a transformer-based model to accurately predict flashover occurrences, while a Three Dimensional Convolutional Neural Network-based model estimates the time to flashover. Flashover Detector and Time Estimator progressively samples video frames at multiple scales, enhancing prediction accuracy. Experimental results demonstrate that the models achieve up to 88.73% accuracy in predicting flashover events and a mean absolute error of 3.41 in time-to-flashover estimation. These findings substantially improve the ability to implement preventive measures. Flashover Detector and Time Estimator thus represents a significant advancement in proactively managing power system reliability, with demonstrated effectiveness and real-time application potential. • End-to-end DL model for early flashover prediction and precise time-to-flashover. • IR video dataset captured in controlled conditions, showing full DBA progression. • High-accuracy models for early flashover detection and low MAE for time-to-flashover prediction. Najmath Ottakath, Abdulla Lutfi, Ali Hamdi, Khaled B. Shaban, Ayman H. El-Hag |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | LLM-Ops and Ensemble Intelligence for Robust LLM Performance: Integrating Fine-Tuning and Majority VotingabstractThis paper presents a novel approach that combines LLM-Ops with ensemble intelligence to enhance document processing accuracy. We introduce a multi-OCR pipeline that leverages four distinct OCR engines and four fine-tuned lightweight LLMs in a two-tier majority voting framework. Through automated fine-tuning after every 500 processed records, our system demonstrates that lightweight (7B parameter) models can achieve performance comparable to much larger (27B parameter) alternatives. Experimental results show field accuracy improvements from $85.6 \%$ to $94.5 \%$ after three fine-tuning cycles, with processing speeds twice as fast as larger models. The continuous improvement loop enabled by our LLM-Ops framework ensures the system evolves with minimal human intervention, making advanced document intelligence more accessible and deployable for real-world applications. Osama Hosam Abdellatif, Ahmed Ayman, Abdelrahman Nader, Ali Hamdi, Khaled B. Shaban |
AICCSA | 4 |
| 2025 | Scaling Arabic Medical Chatbots Using Synthetic Data: Enhancing Generative AI with Synthetic Patient RecordsabstractThe development of medical chatbots in Arabic is significantly constrained by the scarcity of large-scale, highquality annotated datasets. While prior efforts compiled a dataset of 20,000 Arabic patient-doctor interactions from social media to fine-tune large language models (LLMs), model scalability and generalization remained limited. In this study, we propose a scalable synthetic data augmentation strategy to expand the training corpus to 100,000 records. Using advanced generative AI systems-ChatGPT-4o and Gemini 2.5 Pro-we generated 80,000 contextually relevant and medically coherent synthetic question-answer pairs grounded in the structure of the original dataset. These synthetic samples were semantically filtered, manually validated, and integrated into the training pipeline. We fine-tuned five LLMs, including Mistral-7B and AraGPT2, and evaluated their performance using BERTScore metrics and expert-driven qualitative assessments. To further analyze the effectiveness of synthetic sources, we conducted an ablation study comparing ChatGPT-4o and Gemini-generated data independently. The results showed that ChatGPT-4o data consistently led to higher F1-scores and fewer hallucinations across all models. Overall, our findings demonstrate the viability of synthetic augmentation as a practical solution for enhancing domain-specific language models in low-resource medical NLP, paving the way for more inclusive, scalable, and accurate Arabic healthcare chatbot systems. Abdulrahman Allam, Seif Ahmed, Ali Hamdi, Khaled B. Shaban |
AICCSA | 3 |
| 2025 | MHA-DQN: Personalized Route Planning for Asthma Patients Using Multi-Head Attention and Deep Reinforcement LearningabstractIndividuals with asthma face significant health risks in urban environments with poor air quality, yet traditional navigation systems fail to account for environmental factors that exacerbate respiratory conditions. To address this, we propose a novel hybrid deep learning framework integrating Multi-Head Attention (MHA) with Deep Q-Network (DQN) reinforcement learning for personalized, health-aware pedestrian route optimization. Our approach uniquely combines real-time environmental data—such as air quality index (AQI), temperature, and humidity—with individual mobility patterns and health profiles to prioritize respiratory safety. The model achieves $87.5 \%$ prediction accuracy and an $84 \%$ F1-score, with spatial errors of 1.43 MAE and 2.24 RMSE for route precision, outperforming conventional routing algorithms and standalone deep learning models. This work offers a scalable, adaptive solution for smart cities, enhancing mobility while safeguarding the health of vulnerable populations. Nada Ayman, Shaimaa Alaa Esmail, Ali Hamdi, Khaled B. Shaban, Hozaifa Kassab |
AICCSA | 3 |
| 2025 | Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer ExtractionabstractQuranic Question Answering presents unique challenges due to the linguistic complexity of Classical Arabic and the semantic richness of religious texts. In this paper, we propose a novel two-stage framework that addresses both passage retrieval and answer extraction. For passage retrieval, we ensemble finetuned Arabic language models to achieve superior ranking performance. For answer extraction, we employ instruction-tuned large language models with few-shot prompting to overcome the limitations of fine-tuning on small datasets. Our approach achieves state-of-the-art results on the Quran QA 2023 Shared Task, with a MAP@10 of 0.3128 and MRR@10 of 0.5763 for retrieval, and a pAP@10 of 0.669 for extraction, substantially outperforming previous methods. These results demonstrate that combining model ensembling and instruction-tuned language models effectively addresses the challenges of low-resource question answering in specialized domains. Mohamed Basem, Islam Oshallah, Ali Hamdi, Khaled B. Shaban, Hozaifa Kassab |
AICCSA | 3 |
| 2025 | Balancing Factual Consistency and Diversity in Abstractive Summarization via Model-Agnostic Composite RerankingabstractAbstractive text summarization has achieved remarkable progress with transformer-based models, yet these systems often produce fluent outputs that suffer from factual inconsistency and redundancy, limiting their reliability in realworld use. Existing solutions to improve factual accuracy typically rely on additional training, specialized architectures, or large-scale annotations, which are computationally costly and difficult to deploy in resource-constrained environments. This paper introduces a lightweight, training-free framework for enhancing abstractive summarization by combining decoding diversity with multi-metric reranking. Our method generates candidate summaries using both deterministic beam search and stochastic top-k sampling, then applies a composite scoring function that integrates ROUGE-1, METEOR, and BERTScore to select the most accurate and informative summary. Experiments on the XSum dataset demonstrate consistent improvements, achieving ROUGE-1 =51.7, ROUGE-2 =27.5, and ROUGE-L =42.7, outperforming competitive reranking baselines. A preference-based evaluation further showed that reranked outputs were favored in 71% of cases, confirming that automatic metric gains align with human-perceived quality. The results highlight a practical and resource-efficient solution for improving summarization quality without retraining. Mariam Elewa, Ali Hamdi, Hozaifa Kassab, Khaled B. Shaban |
AICCSA | 2 |
| 2025 | An Ensemble Classification Approach in A Multi-Layered Large Language Model Framework for Disease PredictionabstractSocial telehealth has made remarkable progress in healthcare by allowing patients to post symptoms and participate in medical consultations remotely. Users frequently post symptoms on social media and online health platforms, creating a huge repository of medical data that can be leveraged for disease classification. Large language models (LLMs) such as LLAMA3 and GPT-3.5, along with transformer-based models like BERT, have demonstrated strong capabilities in processing complex medical text. In this study, we evaluate three Arabic medical text preprocessing methods such as summarization, refinement, and Named Entity Recognition (NER) before applying finetuned Arabic transformer models (CAMeLBERT, AraBERT, and AsafayaBERT). To enhance robustness, we adopt a majority voting ensemble that combines predictions from original and preprocessed text representations. This approach achieved the best classification accuracy of 80.56%, thus showing its effectiveness in leveraging various text representations and model predictions to improve the understanding of medical texts. To the best of our knowledge, this is the first work that integrates LLM-based preprocessing with fine-tuned Arabic transformer models and ensemble learning for disease classification in Arabic social telehealth data. Ali Hamdi, Malak Mohamed, Rokaia Emad, Khaled B. Shaban |
AICCSA | 1 |
| 2025 | CAKD: A Confidence-Aware Knowledge Distillation Approach for Building Compact and Efficient LLMsabstractHigh-quality models across various natural language processing tasks, such as summarization and chatbots, often rely on large architectures, making them computationally intensive and challenging to deploy in resource-constrained environments. While knowledge distillation enables smaller student models to approximate the performance of larger teacher models, existing methods frequently encounter significant trade-offs between accuracy and efficiency. Additionally, uncertain predictions from teacher models can negatively impact the student’s learning process. In this paper, we introduce CAKD, a novel approach that optimizes the training of student models by selectively emphasizing the teacher model’s most reliable predictions using confidence scores. By integrating entropybased confidence weighting into the distillation loss, CAKD effectively prioritizes high-confidence samples, resulting in improved performance and efficiency. Our experiments on text summarization (using a BART-based model on the CNN/DM dataset) and chatbot tasks (using Llamabased model on the DailyDialog and PersonaChat datasets) demonstrate that CAKD achieves significant performance gains over larger teacher models, with improvements of 10.53, 2.1 and 0.38 ROUGE-L points respectively. Mohammad Basheer Kotit, Omama Hamad, Khaled B. Shaban, Ali Hamdi |
AICCSA | 4 |
| 2025 | Attentional Trajectory Modeling for Text-to-3D Generation with Gaussian Multi-View Diffusion and SDS++abstractThe advancement in converting text into 3D scenes has driven significant improvements in generating realistic and adaptable 3D models. However, existing methods face persistent challenges, including inconsistent multi-view generation, limited scene complexity, and an inability to handle real-world datasets with varying camera trajectories. To address these limitations, we introduce a novel approach utilizing a four-part system: the Cinematographer (Trajectory Diffusion Transformer - Traj-DiT), Decorator (Gaussian-driven Multi-view Latent Diffusion Model - GM-LDM, and Detailer (SDS++ loss). Our model enhances 3D scene generation by aligning 3D Gaussians with pixel data, refining 3D structures, and applying realistic surface properties while ensuring view-to-view consistency and accommodating complex scenes. Our research methodology integrates dense-view trajectories processed through BERT, employing multi-head selfattention to handle intricate, real-world camera movements. We conducted extensive experimental comparisons with state-of-theart models, including DreamFusion, Magic3D, LatentNeRF, SJC, Fantasia3D, ProlificDreamer, and Director3D, using BRISQUE, NIQE, and CLIP-Score metrics. Our approach achieved a BRISQUE score of 23.3, NIQE score of 4.34, and CLIP-Score score of 86.1, significantly outperforming all competing methods. These results demonstrate our model’s superior visual clarity, multi-view consistency, geometric accuracy, and photo-realistic rendering. This work represents a substantial advancement in text-to-3D generation, with promising applications in gaming, simulation, and virtual reality. Marena Anis Labib, Ali Hamdi, Khaled B. Shaban |
AICCSA | 2 |
| 2025 | MultiFuzz: A Dense Retrieval-based Multi-Agent System for Network Protocol FuzzingabstractTraditional protocol fuzzing techniques, such as those employed by AFL-based systems, often lack effectiveness due to a limited semantic understanding of complex protocol grammars and rigid seed mutation strategies. Recent works, such as ChatAFL, have integrated Large Language Models (LLMs) to guide protocol fuzzing and address these limitations, pushing protocol fuzzers to wider exploration of the protocol state space. But ChatAFL still faces issues like unreliable output, LLM hallucinations, and assumptions of LLM knowledge about protocol specifications. This paper introduces MultiFuzz, a novel dense retrieval-based multi-agent system designed to overcome these limitations by integrating semantic-aware context retrieval, specialized agents, and structured tool-assisted reasoning. MultiFuzz utilizes agentic chunks of protocol documentation (RFC Documents) to build embeddings in a vector database for a retrieval-augmented generation (RAG) pipeline, enabling agents to generate more reliable and structured outputs, enhancing the fuzzer in mutating protocol messages with enhanced state coverage and adherence to syntactic constraints. The framework decomposes the fuzzing process into modular groups of agents that collaborate through chain-of-thought reasoning to dynamically adapt fuzzing strategies based on the retrieved contextual knowledge. Experimental evaluations on the Real-Time Streaming Protocol (RTSP) demonstrate that MultiFuzz significantly improves branch coverage and explores deeper protocol states and transitions over state-of-the-art (SOTA) fuzzers such as NSFuzz, AFLNet, and ChatAFL. By combining dense retrieval, agentic coordination, and language model reasoning, MultiFuzz establishes a new paradigm in autonomous protocol fuzzing, offering a scalable and extensible foundation for future research in intelligent agentic-based fuzzing systems. Youssef Maklad, Fares Wael, Ali Hamdi, Wael Elsersy, Khaled B. Shaban |
AICCSA | 3 |
| 2025 | Weather-Aware Transformer for Real-Time Route Optimization in Drone-as-a-Service OperationsabstractThis paper presents a novel framework to accelerate route prediction in Drone-as-a-Service operations through weather-aware deep learning models. While classical pathplanning algorithms, such as $\mathrm{A}^{*}$ and Dijkstra, provide optimal solutions, their computational complexity limits real-time applicability in dynamic environments. We address this limitation by training machine learning and deep learning models on synthetic datasets generated from classical algorithm simulations. Our approach incorporates transformer-based and attention-based architectures that utilize weather heuristics to predict optimal nextnode selections while accounting for meteorological conditions affecting drone operations. The attention mechanisms dynamically weight environmental factors including wind patterns, wind bearing, and temperature to enhance routing decisions under adverse weather conditions. Experimental results demonstrate that our weather-aware models achieve significant computational speedup over traditional algorithms while maintaining route optimization performance, with transformer-based architectures showing superior adaptation to dynamic environmental constraints. The proposed framework enables real-time, weatherresponsive route optimization for large-scale DaaS operations, representing a substantial advancement in the efficiency and safety of autonomous drone systems. Kamal Mohamed, Lillian Wassim, Ali Hamdi, Khaled B. Shaban |
AICCSA | 3 |
| 2025 | Efficient Segmentation of Solar Panel Defects Using Knowledge DistillationabstractEfficient identification of anomalies in solar panels is essential for ensuring optimal energy generation and long-term system reliability. This work introduces a two-stage automated framework leveraging drone-captured infrared imagery to detect and segment common defects such as hotspots, dirt accumulation, and shadowing. In the first stage, multiple YOLO-based detection models were evaluated to localize defective regions. YOLOv10 emerged as the best-performing model, achieving a mAP50 of $85 \%$. The detected regions were then used to guide the segmentation stage. A combination of advanced data augmentation techniques (e.g., flipping, lighting variations, and rotation) and a knowledge distillation strategy - where YOLOv11-Seg (student) learned from YOLOv8-Seg (teacher) - significantly improved segmentation accuracy.The final results of the evaluation avaluation of the test set demonstrated strong segmentation performance. The model achieved a class-averaged mask mAP50 of 0.891, with individual class results as follows: 0.975 for Serious Hot Spot, 0.931 for Slight Hot Spot, and 0.766 for Dirt. These results highlight the effectiveness of combining detection-guided segmentation with targeted training enhancements for real-world solar panel inspection. Shahd Tarek, Ali Hamdi, Khaled B. Shaban |
AICCSA | 2 |
| 2025 | MSLEF: Multi-Segment LLM Ensemble Finetuning in RecruitmentabstractThis paper presents MSLEF, a multi-segment ensemble framework that employs LLM fine-tuning to enhance resume parsing in recruitment automation. It integrates finetuned Large Language Models (LLMs) using weighted voting, with each model specializing in a specific resume segment to boost accuracy. Building on MLAR [1], MSLEF introduces a segmentaware architecture that leverages field-specific weighting tailored to each resume part, effectively overcoming the limitations of single-model systems by adapting to diverse formats and structures. The framework incorporates Gemini-2.5-Flash LLM as a high-level aggregator for complex sections and utilizes Gemma 9B, LLaMA 3.1 8B, and Phi-4 14B. MSLEF achieves significant improvements in Exact Match (EM), F1 score, BLEU, ROUGE, and Recruitment Similarity (RS) metrics, outperforming the best single model by up to +7 % in RS. Its segment-aware design enhances generalization across varied resume layouts, making it highly adaptable to real-world hiring scenarios while ensuring precise and reliable candidate representation. Omar Walid, Mohamed T. Younes, Khaled B. Shaban, Mai Hassan, Ali Hamdi |
AICCSA | 5 |
| 2025 | Augmented Fine-Tuned LLMs for Enhanced Recruitment AutomationabstractThis paper presents a novel approach to recruitment automation. Large Language Models (LLMs) were fine-tuned to improve accuracy and efficiency. Building upon our previous work on the Multilayer Large Language Model-Based Robotic Process Automation Applicant Tracking (MLAR) system [1]. This work introduces a novel methodology. Training fine-tuned LLMs specifically tuned for recruitment tasks. The proposed framework addresses the limitations of generic LLMs by creating a synthetic dataset that uses a standardized JSON format. This helps ensure consistency and scalability. In addition to the synthetic data set, the resumes were parsed using DeepSeek, a high-parameter LLM. The resumes were parsed into the same structured JSON format and placed in the training set. This will help improve data diversity and realism. Through experimentation, we demonstrate significant improvements in performance metrics, such as exact match, F1 score, BLEU score, ROUGE score, and overall similarity compared to base models and other state-of-the-art LLMs. In particular, the fine-tuned Phi-4 model achieved the highest F1 score of $90.62 \%$, indicating exceptional precision and recall in recruitment tasks. This study highlights the potential of fine-tuned LLMs. Furthermore, it will revolutionize recruitment workflows by providing more accurate candidate-job matching. Mohamed T. Younes, Omar Walid, Khaled B. Shaban, Ali Hamdi, Mai Hassan |
AICCSA | 4 |
| 2025 | Intelligent Spectral Efficient Communications Underlay 6G VHetnetsabstractSixth-generation (6G) wireless networks are designed to deliver superior performance by enhancing spectral efficiency and minimizing latency. It enables next-generation applications such as virtual reality, autonomous transportation, and the Internet of Things (IoT). To achieve this, 6G integrates multiple advanced technologies, including satellite systems, unmanned aerial vehicles (UAVs), and stratospheric platforms, forming vertical heterogeneous networks (VHetNets). Despite these advancements, challenges persist in resource allocation, interference control, and capacity optimization. This paper proposes an intelligent spectral efficiency algorithm tailored for 6G VHetNets to enhance network resilience and high-speed connectivity. Simulation results indicate that the proposed approach not only improves overall network capacity but also ensures low computational complexity and rapid convergence. Sawsan Selmi, Leila Aissaoui Ferhi, Wyssem Fathallah, Ali Hamdi, Dhaou Bouchouicha, Hedi Sakli, Ridha Bouallègue |
IWCMC | 4 |
| 2025 | LexiSem: A re-ranker balancing lexical and semantic quality for enhanced abstractive summarizationabstractSequence-to-sequence neural networks have recently achieved significant success in abstractive summarization, especially through fine-tuning large pre-trained language models on downstream datasets. However, these models frequently suffer from exposure bias, which can impair their performance. To address this, re-ranking systems have been introduced, but their potential remains underexplored despite some demonstrated performance gains. Most prior work relies on ROUGE scores and aligned candidate summaries for ranking, exposing a substantial gap between semantic similarity and lexical overlap metrics. In this study, we demonstrate that a second-stage model can be trained to re-rank a set of summary candidates, significantly enhancing performance. Our novel approach leverages a re-ranker that balance lexical and semantic quality. Additionally, we introduce a new strategy for defining negative samples in ranking models. Through experiments on the CNN/DailyMail, XSum and Reddit TIFU datasets, we show that our method effectively estimates the semantic content of summaries without compromising lexical quality. In particular, our method sets a new performance benchmark on the CNN/DailyMail dataset (48.18 R1, 24.46 R2, 45.05 RL) and on Reddit TIFU (30.37 R1,RL 23.87). Eman Aloraini, Hozaifa Kassab, Ali Hamdi, Khaled B. Shaban |
Neurocomputing | 3 |
| 2025 | Drone-as-a-Service: Research Challenges and DirectionsabstractWe conduct a survey on drones used as a service, denoted as drone-as-a-service (DaaS). We develop a novel taxonomy based on DaaS functions, research tasks, and application domains. We provide a discussion on drones and their associated capabilities based on their type of use. We propose a three-layered DaaS system architecture that vertically integratescloudcomputing,drones, andservicesas a reference framework to compare existing drone service implementations. Additionally, we propose a representative uncertainty-aware DaaS model for delivery scenarios, illustrating how service definitions can incorporate both functional and nonfunctional attributes under dynamic environmental conditions. Finally, we identify and discuss future research directions and open problems related to the use of drones for service delivery. Ali Hamdi, Balsam Alkouz, Babar Shahzaad, Athman Bouguettaya, Azadeh Ghari Neiat, Flora D. Salim, Du Yong Kim |
Proc. IEEE | 1 |
| 2024 | ASEM: Enhancing Empathy in Chatbot through Attention-based Sentiment and Emotion ModelingabstractEffective feature representations play a critical role in enhancing the performance of text generation models that rely on deep neural networks. However, current approaches suffer from several drawbacks, such as the inability to capture the deep semantics of language and sensitivity to minor input variations, resulting in significant changes in the generated text. In this paper, we present a novel solution to these challenges by employing a mixture of experts, multiple encoders, to offer distinct perspectives on the emotional state of the user’s utterance while simultaneously enhancing performance. We propose an end-to-end model architecture called ASEM that performs emotion analysis on top of sentiment analysis for open-domain chatbots, enabling the generation of empathetic responses that are fluent and relevant. In contrast to traditional attention mechanisms, the proposed model employs a specialized attention strategy that uniquely zeroes in on sentiment and emotion nuances within the user’s utterance. This ensures the generation of context-rich representations tailored to the underlying emotional tone and sentiment intricacies of the text. Our approach outperforms existing methods for generating empathetic embeddings, providing empathetic and diverse responses. The performance of our proposed model significantly exceeds that of existing models, enhancing emotion detection accuracy by 6.2% and lexical diversity by 1.4%. ASEM code is released at https://github.com/MIRAH-Official/Empathetic-Chatbot-ASEM.git Omama Hamad, Khaled B. Shaban, Ali Hamdi |
LREC/COLING | 3 |
| 2022 | Drone-as-a-Service Composition Under UncertaintyabstractWe propose an uncertainty-aware service approach to provide drone-based delivery services called Drone-as-a-Service (DaaS) effectively. Specifically, we propose a service model of DaaS based on the dynamic spatiotemporal features of drones and their in-flight contexts. The proposed DaaS service approach consists of three components: scheduling, route-planning, and composition. First, we develop a DaaS scheduling model to generate DaaS itineraries through a Skyway network. Second, we propose anuncertainty-aware DaaS route-planning algorithmthat selects the optimal Skyways under weather uncertainties. Third, we develop two DaaS composition techniques to select an optimal DaaS composition at each station of the planned route. Aspatiotemporal DaaS composerfirst selects the optimal DaaSs based on their spatiotemporal availability and drone capabilities. Apredictive DaaS composerthen utilises the outcome of the first composer to enable fast and accurate DaaS composition using several Machine Learning classification methods. We train the classifiers using a new set of spatiotemporal features which are in addition to other DaaS QoS properties. Our experiments results show the effectiveness and efficiency of the proposed approach. Ali Hamdi, Flora D. Salim, Du Yong Kim, Azadeh Ghari Neiat, Athman Bouguettaya |
IEEE Trans. Serv. Comput. | 1 |
| 2021 | MARL: Multimodal Attentional Representation Learning for Disease Prediction
Ali Hamdi, Amr Aboeleneen, Khaled B. Shaban |
ICVS | 1 |
| 2020 | DroTrack: High-speed Drone-based Object Tracking Under UncertaintyabstractWe present DroTrack, a high-speed visual single-object tracking framework for drone-captured video sequences. Most of the existing object tracking methods are designed to tackle well-known challenges, such as occlusion and cluttered backgrounds. The complex motion of drones, i.e., multiple degrees of freedom in three-dimensional space, causes high uncertainty. The uncertainty problem leads to inaccurate location predictions and fuzziness in scale estimations. DroTrack solves such issues by discovering the dependency between object representation and motion geometry. We implement an effective object segmentation based on Fuzzy C Means (FCM). We incorporate the spatial information into the membership function to cluster the most discriminative segments. We then enhance the object segmentation by using a pre-trained Convolution Neural Network (CNN) model. DroTrack also leverages the geometrical angular motion to estimate a reliable object scale. We discuss the experimental results and performance evaluation using two datasets of 51,462 drone-captured frames. The combination of the FCM segmentation and the angular scaling increased DroTrack precision by up to 9% and decreased the centre location error by 162 pixels on average. DroTrack outperforms all the high-speed trackers and achieves comparable results in comparison to deep learning trackers. DroTrack offers high frame rates up to 1000 frame per second (fps) with the best location precision, more than a set of state-of-the-art real-time trackers. Ali Hamdi, Flora D. Salim, Du Yong Kim |
FUZZ-IEEE | 1 |
| 2019 | A Survey of Opinion Mining in Arabic: A Comprehensive System Perspective Covering Challenges and Advances in Tools, Resources, Models, Applications, and VisualizationsabstractOpinion-mining or sentiment analysis continues to gain interest in industry and academics. While there has been significant progress in developing models for sentiment analysis, the field remains an active area of research for many languages across the world, and in particular for the Arabic language, which is the fifth most-spoken language and has become the fourth most-used language on the Internet. With the flurry of research activity in Arabic opinion mining, several researchers have provided surveys to capture advances in the field. While these surveys capture a wealth of important progress in the field, the fast pace of advances in machine learning and natural language processing (NLP) necessitates a continuous need for a more up-to-date literature survey. The aim of this article is to provide a comprehensive literature survey for state-of-the-art advances in Arabic opinion mining. The survey goes beyond surveying previous works that were primarily focused on classification models. Instead, this article provides a comprehensive system perspective by covering advances in different aspects of an opinion-mining system, including advances in NLP software tools, lexical sentiment and corpora resources, classification models, and applications of opinion mining. It also presents future directions for opinion mining in Arabic. The survey also covers latest advances in the field, including deep learning advances in Arabic Opinion Mining. The article provides state-of-the-art information to help new or established researchers in the field as well as industry developers who aim to deploy an operational complete opinion-mining system. Key insights are captured at the end of each section for particular aspects of the opinion-mining system giving the reader a choice of focusing on particular aspects of interest. Gilbert Badaro, Ramy Baly, Hazem M. Hajj, Wassim El-Hajj, Khaled B. Shaban, Nizar Habash, Ahmad A. Al Sallab, Ali Hamdi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 8 |
| 2018 | CLASENTI: A Class-Specific Sentiment Analysis FrameworkabstractArabic text sentiment analysis suffers from low accuracy due to Arabic-specific challenges (e.g., limited resources, morphological complexity, and dialects) and general linguistic issues (e.g., fuzziness, implicit sentiment, sarcasm, and spam). The limited resources problem requires efforts to build new and improved Arabic corpora and lexica. We propose a class-specific sentiment analysis (CLASENTI) framework. The framework includes a new annotation approach to build multi-faceted Arabic corpus and lexicon allowing for simultaneous annotation of different facets, including domains, dialects, linguistic issues, and polarity strengths. Each of these facets has multiple classes (e.g., the nine classes representing dialects found in the Arab world). The new corpus and lexicon annotations facilitate the development of new class-specific classification models and polarity strength calculation. For the new sentiment classification models, we propose a hybrid model combining corpus-based and lexicon-based models. The corpus-based model has two interrelated phases to build; (1) full-corpus classification models for all facets; and (2) class-specific models trained on filtered subsets of the corpus according to the performances of the full-corpus models. To calculate polarity strengths, the lexicon-based model filters the annotated lexicon based on the specific classes of the domain and dialect. As a case study, we collect and annotate 15274 reviews from various sources, including surveys, Facebook comments, and Twitter posts, pertaining to governmental services. In addition, we develop a new web-based application to apply the proposed framework on the case study. CLASENTI framework reaches up to 95% accuracy and 93% F1-Score surpassing the best-known sentiment classifiers implemented in Scikit-learn library that achieve 82% accuracy and 81% F1-Score for Arabic when tested on the same dataset. Ali Hamdi, Khaled B. Shaban, Anazida Zainal |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |