VLDB 2026 Research / reviewers in the wild / expert
Ali Shariq Imran
dblp:89/7455
· DBLP profile ↗
30ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-2416-2878ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond prompt-only segmentation: Foundation model-guided trust-region label refinement for single-stage bone fracture instance segmentationabstractAccurate instance-level delineation of bone fractures on radiographs is challenging because fractures are ultra-thin, low-contrast trajectories, and dense pixel-wise annotation is expensive. Naively using foundation-model masks as pseudo-label supervision can further degrade performance by introducing boundary noise and off-target regions. We propose a controlled foundation-model integration strategy for thin-structure learning: expert annotations define a narrow trust region, foundation-model predictions act only as a local proposal signal, an instance-wise intersection-over-union quality gate rejects unreliable cases, and minimal morphological closing repairs small gaps while preserving connectivity. This converts noisy foundation-model outputs into a constrained near-boundary refinement signal, always anchored to and preserving the original expert annotations. We additionally evaluate the Medical Segment Anything Model without task-specific training and introduce a skeleton-aware measure that quantifies centerline coverage while penalizing boundary halo over-segmentation. We provide a reproducible benchmark for multi-region bone-fracture instance segmentation across one-stage, two-stage, and transformer-based methods, targeting clinician-facing overlays for triage and reporting. Across five-fold cross-validation, the proposed ground-truth-anchored refined-mask strategy provides a modest improvement in sample-count-matched ablations, while unconstrained MedSAM mask supervision degrades performance. Together, these findings support the importance of trust-region constrained foundation-model guidance for thin-structure supervision, and our final model achieves a mean average precision of 0.63 at an intersection-over-union threshold of 0.50 on a held-out test set. Ali Shariq Imran, Zenun Kastrati, Sher Muhammad Daudpota |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Unlocking language barriers: Assessing pre-trained large language models across multilingual tasks and unveiling the black box with Explainable Artificial IntelligenceabstractLarge Language Models (LLMs) have revolutionized many industrial applications and paved the way for fostering a new research direction in many fields. Conventional Natural Language Processing (NLP) techniques, for instance, are no longer necessary for many text-based tasks, including polarity estimation, sentiment and emotion classification, and hate speech detection. However, training a language model for domain-specific tasks is hugely costly and requires high computational power, thereby restricting its true potential for standard tasks. This study, therefore, provides a comprehensive analysis of the latest pre-trained LLMs for various NLP-related applications without fine-tuning them to evaluate their effectiveness. Five language models are thus employed in this study on six distinct NLP tasks (including emotion recognition, sentiment analysis , hate speech detection, irony detection, offensiveness detection, and stance detection) for 12 languages from low- to medium- and high-resource. Generative Pre-trained Transformer 4 (GPT-4) and Gemini Pro outperform state-of-the-art models, achieving average F1 scores of 70.6% and 68.8% on the Tweet Sentiment Multilingual dataset compared to the state-of-the-art average F1 score of 66.8%. The study further interprets the findings obtained by the LLMs using Explainable Artificial Intelligence (XAI). To the best of our knowledge, it is the first time any study has employed explainability on pre-trained language models. Muhamet Kastrati, Ali Shariq Imran, Ehtesham Hashmi, Zenun Kastrati, Sher Muhammad Daudpota, Marenglen Biba |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Aspect-based sentiment analysis for software requirements elicitation using fine-tuned Bidirectional Encoder Representations from Transformers and Explainable Artificial IntelligenceabstractAspect-Based Sentiment Analysis (ABSA) of app reviews allows a better understanding of user preferences regarding specific product features and helps the development team elicit requirements effectively. The existing literature faces challenges such as limited focus on the automation of Requirement Elicitation (RE), insufficient task-specific fine-tuning of models such as Bidirectional Encoder Representations from Transformers (BERT), and lack of interpretability owing to the black-box nature of these models. Therefore, our work makes the following significant contributions to address these challenges: (1) development and evaluation of a robust method based on ABSA for the automation of the RE process; (2) optimization of ABSA using BERT fine-tuning for enhanced performance, which includes conducting a comprehensive ablation study to obtain the best hyperparameters that guarantee the best model performance and robustness; and (3) integration of Explainable Artificial Intelligence (XAI) techniques for enhanced BERT model interpretability. Our work was evaluated on the ABSA Warehouse of Apps REviews (AWARE) dataset, a specifically tailored dataset for the RE process. Our study outperformed baseline models such as the Support Vector Machine (SVM), Convolutional Neural Network (CNN), and BERT, and achieved an average F1-Score of 0.83 for the Aspect Category Detection (ACD) task and 0.94 for the Aspect Category Polarity (ACP) task. In addition, we employed XAI using Locally Interpretable Model-Agnostic Explanations (LIME) to explain the BERT model prediction results, which aids in the improved visualization and interpretability of the app review analysis for the automated RE process. • Fine-tuned Bidirectional Encoder Representations for Aspect-Based Sentiment Analysis. • Aspect-Based Sentiment Analysis on the ABSA Warehouse of Apps REviews dataset. • Explainable Artificial Intelligence for Bidirectional Encoder Representations. Soonh Taj, Sher Muhammad Daudpota, Ali Shariq Imran, Zenun Kastrati |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Random transformations to improve mitigation of query-based black-box attacksabstractThis paper proposes methods to upstage the best-known defences against query-based black-box attacks. These benchmark defences incorporate gaussian noise into input data during inference to achieve state-of-the-art performance in protecting image classification models against the most advanced query-based black-box attacks. Even so there is a need to improve upon them; for example, the widely benchmarked Random noise defense (RND) method has demonstrated limited robustness – achieving only 53.5% and 18.1% with a ResNet-50 model on the CIFAR-10 and ImageNet datasets, respectively – against the square attack, which is commonly regarded as the state-of-the-art black-box attack. Therefore, in this work, we propose two alternatives to gaussian noise addition at inference time: random crop-resize and random rotation of the input images. Although these transformations are generally used for data augmentation while training to improve model invariance and generalisation, their protective potential against query-based black-box attacks at inference time is unexplored. Therefore, for the first time, we report that for such well-trained models either of the two transformations can also blunt powerful query-based black-box attacks when used at inference time on three popular datasets. The results show that the proposed randomised transformations outperform RND in terms of robust accuracy against a strong adversary that uses a high budget of 100,000 queries based on expectation over transformation (EOT) of 10, by 0.9% on the CIFAR-10 dataset, 9.4% on the ImageNet dataset and 1.6% on the Tiny ImageNet dataset. Crucially, in two even tougher attack settings, that is, high-confidence adversarial examples and EOT-50 adversary, these transformations are even more effective as the margin of improvement over the benchmarks increases further. Ziad Tariq Muhammad Ali, R. Muhammad Atif Azad, Muhammad Ajmal Azad, James Holyhead, Iain Rice, Ali Shariq Imran |
Expert Syst. Appl. | 6 |
| 2024 | Navigating Limitations With Precision: A Fine-Grained Ensemble Approach To Wrist Pathology Recognition On A Limited X-Ray DatasetabstractThe exploration of automated wrist fracture recognition has gained considerable research attention in recent years. In practical medical scenarios, physicians and surgeons may lack the specialized expertise required for accurate X-ray interpretation, highlighting the need for machine vision to enhance diagnostic accuracy. However, conventional recognition techniques face challenges in discerning subtle differences in X-rays when classifying wrist pathologies, as many of these pathologies, such as fractures, can be small and hard to distinguish. This study tackles wrist pathology recognition as a fine-grained visual recognition (FGVR) problem, utilizing a limited, custom-curated dataset that mirrors real-world medical constraints, relying solely on image-level annotations. We introduce a specialized FGVR-based ensemble approach to identify discriminative regions within X rays. We employ an Explainable AI (XAI) technique called Grad-CAM to pinpoint these regions. Our ensemble approach outperformed many conventional SOTA and FGVR techniques, underscoring the effectiveness of our strategy in enhancing accuracy in wrist pathology recognition. Ali Shariq Imran, Mohib Ullah, Zenun Kastrati, Sher Muhammad Daudpota |
ICIP | 2 |
| 2024 | A Self-Supervised Diffusion Framework For Facial Emotion RecognitionabstractIn this paper, we introduced a novel Facial Emotion Recognition (FER) framework that utilizes a diffusion-based approach and an attention mechanism. The model is efficiently trained through self-supervised learning, leveraging labeled and unlabelled data. The proposed framework has been rigorously tested on the FER2013 and AffectNet datasets, achieving promising accuracies of $67.2 \%$ and $68.1 \%$, respectively. The quantitative results not only surpass the performance of existing state-of-the-art FER models but also demonstrate the synergistic effect of combining diffusion-based modeling with self-supervised learning and attention mechanisms within a solid architectural framework. Our approach sets a new benchmark in the field, offering a significant step forward in the accurate and efficient recognition of facial expressions. Saif Hassan, Mohib Ullah, Ali Shariq Imran, Ghulam Mujtaba 0001, Muhammad Mudassar Yamin, Ehtesham Hashmi, Faouzi Alaya Cheikh, Azeddine Beghdadi |
ICIP | 3 |
| 2024 | Analyzing and Predicting the Helpfulness of Reviews in MOOCs Context Using Deep LearningabstractStudents’ feedback is an essential part of the teaching-learning process and serves as an effective instrument for continuous improvement in educational environments. The insights gathered from students’ experiences and perceptions expressed in reviews provide instructors with a valuable resource to enhance their teaching methods, instructional design, and overall classroom dynamics. However, students’ reviews are often unclear, contradictory, and conflicting with each other, making their interpretation and use challenging. Therefore, this study proposes a novel deep learning-based approach that helps course designers and instructors effectively identify constructive and useful reviews. The approach leverages the integration of several attributes, including textual review, student satisfaction, meta-data of the course, and review-derived information such as sentiment, readability, and review depth. The approach is tested on a real-life dataset comprising 38,717 reviews gathered from the Coursera learning platform for the purpose of this study. The experimental results, with an F1-score of 0.91, suggest that the approach can be an effective tool for educators and instructional designers to identify helpful student reviews. Zenun Kastrati, Sana Fatima, Arianit Kurti, Sher Muhammad Daudpota, Ali Shariq Imran |
KES | 5 |
| 2024 | Synthetic Image Generation Using Deep Learning: A Systematic Literature ReviewabstractABSTRACT The advent of deep neural networks and improved computational power have brought a revolutionary transformation in the fields of computer vision and image processing. Within the realm of computer vision, there has been a significant interest in the area of synthetic image generation, which is a creative side of AI. Many researchers have introduced innovative methods to identify deep neural network‐based architectures involved in image generation via different modes of input, like text, scene graph layouts and so forth to generate synthetic images. Computer‐generated images have been found to contribute a lot to the training of different machine and deep‐learning models. Nonetheless, we have observed an immediate need for a comprehensive and systematic literature review that encompasses a summary and critical evaluation of current primary studies' approaches toward image generation. To address this, we carried out a systematic literature review on synthetic image generation approaches published from 2018 to February 2023. Moreover, we have conducted a systematic review of various datasets, approaches to image generation, performance metrics for existing methods, and a brief experimental comparison of DCGAN (deep convolutional generative adversarial network) and cGAN (conditional generative adversarial network) in the context of image generation. Additionally, we have identified applications related to image generation models with critical evaluation of the primary studies on the subject matter. Finally, we present some future research directions to further contribute to the field of image generation using deep neural networks. Aisha Zulfiqar, Sher Muhammad Daudpota, Ali Shariq Imran, Zenun Kastrati, Mohib Ullah, Suraksha Sadhwani |
Comput. Intell. | 3 |
| 2024 | Leveraging distant supervision and deep learning for twitter sentiment and emotion classificationabstractAbstract Nowadays, various applications across industries, healthcare, and security have begun adopting automatic sentiment analysis and emotion detection in short texts, such as posts from social media. Twitter stands out as one of the most popular online social media platforms due to its easy, unique, and advanced accessibility using the API. On the other hand, supervised learning is the most widely used paradigm for tasks involving sentiment polarity and fine-grained emotion detection in short and informal texts, such as Twitter posts. However, supervised learning models are data-hungry and heavily reliant on abundant labeled data, which remains a challenge. This study aims to address this challenge by creating a large-scale real-world dataset of 17.5 million tweets. A distant supervision approach relying on emojis available in tweets is applied to label tweets corresponding to Ekman’s six basic emotions. Additionally, we conducted a series of experiments using various conventional machine learning models and deep learning, including transformer-based models, on our dataset to establish baseline results. The experimental results and an extensive ablation analysis on the dataset showed that BiLSTM with FastText and an attention mechanism outperforms other models in both classification tasks, achieving an F1-score of 70.92% for sentiment classification and 54.85% for emotion detection. Muhamet Kastrati, Zenun Kastrati, Ali Shariq Imran, Marenglen Biba |
J. Intell. Inf. Syst. | 3 |
| 2023 | Attention-Guided Self-supervised Framework for Facial Emotion Recognition
Saif Hassan, Mohib Ullah, Ali Shariq Imran, Faouzi Alaya Cheikh |
PRICAI (3) | 3 |
| 2023 | Improving news headline text generation quality through frequent POS-Tag patterns analysisabstractOriginal synthetic content writing is one of the human abilities that algorithms aspire to emulate. The advent of sophisticated algorithms, especially based on neural networks has shown promising results in recent times. A watershed moment was witnessed when the attention mechanism was introduced which paved the way for transformers, a new exciting architecture in natural language processing. Recent sensations like GPT and BERT for synthetic text generation rely on NLP transformers. Although, GPT and BERT-based models are capable of generating creative text given they are properly trained on abundant data, however, the generated text suffers the quality aspect when limited data is available. This is especially an issue for low-resource languages where labeled data is still scarce. In such cases, the generated text, more often than not, lacks the proper sentence structure, thus unreadable. This study proposes a post-processing step in text generation that improves the quality of generated text through the GPT model. The proposed post-processing step is based on the analysis of POS tagging patterns in the original text and accepts only those generated sentences from GPT which satisfy POS patterns that are originally learned from the data. We exploit the GPT model to generate English headlines by utilizing Australian Broadcasting Corporation (ABC) news dataset. Furthermore, for assessing the applicability of the model in low-resource languages, we also train the model on the Urdu news dataset for Urdu news headlines generation. The experiments presented in this paper on these datasets from high- and low-resource languages show that the performance of generated headlines has a significant improvement by using the proposed headline POS pattern extraction. We evaluate the performance through subjective evaluation as well as using text generation quality metrics like BLEU and ROUGE. Noureen Fatima, Sher Muhammad Daudpota, Zenun Kastrati, Ali Shariq Imran, Saif Hassan, Nouh Sabri Elmitwally |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Depression Screening in Humans With AI and Deep Learning TechniquesabstractSocial media platforms have been widely used as a communication tool where most of the population expresses their feelings and shares life experiences. Along with general information about the public, these platforms hold an ample amount of content related to depressed users and thus can generate sensitive social signals indicating if a person is suffering from some serious issues, such as self-harm, suicidal thoughts, or intention for an unlawful act. Early depression detection using advanced natural language processing (NLP), deep machine learning, and transfer learning techniques can assist in designing an efficient system to detect major depressive systems at an early stage. The current depression detection models are not enough to capture sensitive social signals indicating the true mood, personality, and behavior of an individual. Thus, making the current systems unsatisfactory. To address this life-threatening human-health problem, we propose an efficient artificial intelligence (AI) and deep learning (DL)-based model for identifying depressed individuals on social media platforms. The model employs hybrid feature-based behavioral-biometric signals captured using Word2Vec, term frequency-inverse document frequency (TF-IDF) models to learn a convolutional neural network (CNN) and long-short term memory (LSTM) models. The data are captured from multiple sources using advanced crawling strategies to have data variety in the corpus. Thus, making the proposed system effective across platforms. The Dataset produced by this study is the first of its kind with a variety of depressive signals from online social network (OSN) platforms including Facebook, Twitter, and YouTube. The experiments have shown that both DL models LSTM and CNN, and the hybrid (CNN + LSTM) models achieved promising results on all individual as well as combined datasets. Out of 24 experiments for Word2Vec LSTM and Word2Vec (CNN + LSTM) models, we achieved the accuracy of 99.02% and 99.01%, respectively, and recorded as best results outperforming all the existing approaches on performance measures such as recall, precision, accuracy, and${F}1$-score. The Word2Vec-based features have been proved optimal features for detecting depressions symptoms on Facebook corpus (FC) and YouTube corpus (YC) by achieving an accuracy of 95.02% (with CNN) and 98.15% (with CNN + LSTM), respectively. Mudasir Ahmad Wani, Mohammed Ahmed El-Affendi, Kashish Ara Shakil, Ali Shariq Imran, Ahmed A. Abd El-Latif 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2022 | Sentiment Polarity and Emotion Detection from Tweets Using Distant Supervision and Deep Learning Models
Muhamet Kastrati, Marenglen Biba, Ali Shariq Imran, Zenun Kastrati |
ISMIS | 3 |
| 2021 | A Two-Stage Deep Modeling Approach to Articulatory InversionabstractThis paper proposes a two-stage deep feed-forward neural network (DNN) to tackle the acoustic-to-articulatory inversion (AAI) problem. DNNs are a viable solution for the AAI task, but the temporal continuity of the estimated articulatory values has not been exploited properly when a DNN is employed. In this work, we propose to address the lack of any temporal constraints while enforcing a parameter-parsimonious solution by deploying a two-stage solution based only on DNNs: (i) Articulatory trajectories are estimated in a first stage using DNN, and (ii) a temporal window of the estimated trajectories is used in a follow-up DNN stage as a refinement. The first stage estimation could be thought of as an auxiliary additional information that poses some constraints on the inversion process. Experimental evidence demonstrates an average error reduction of 7.51% in terms of RMSE compared to the baseline, and an improvement of 2.39% with respect to Pearson correlation is also attained. Finally, we should point out that AAI is still a highly challenging problem, mainly due to the non-linearity of the acoustic-to-articulatory and one-to-many mapping. It is thus promising that a significant improvement was attained with our simple yet elegant solution. Abdolreza Sabzi Shahrebabaki, Negar Olfati, Ali Shariq Imran, Magne Hallstein Johnsen, Sabato Marco Siniscalchi, Torbjørn Svendsen |
ICASSP | 3 |
| 2021 | Human action recognition using attention based LSTM network with dilated CNN features
Khan Muhammad 0001, Mustaqeem Khan 0001, Amin Ullah, Ali Shariq Imran, Mustafa Servet Kiran, Giovanna Sannino, Victor Hugo C. de Albuquerque |
Future Gener. Comput. Syst. | 4 |
| 2021 | Light-DehazeNet: A Novel Lightweight CNN Architecture for Single Image DehazingabstractDue to the rapid development of artificial intelligence technology, industrial sectors are revolutionizing in automation, reliability, and robustness, thereby significantly increasing quality and productivity. Most of the surveillance and industrial sectors are monitored by visual sensor networks capturing different surrounding environment images. However, during tempestuous weather conditions, the visual quality of the images is reduced due to contaminated suspended atmospheric particles that affect the overall surveillance systems. To tackle these challenges, this article presents a computationally efficient lightweight convolutional neural network referred to as Light-DehazeNet (LD-Net) for the reconstruction of hazy images. Unlike other learning-based approaches, which separately measure the transmission map and the atmospheric light, our proposed LD-Net jointly estimates both the transmission map and the atmospheric light using a transformed atmospheric scattering model. Furthermore, a color visibility restoration method is proposed to evade the color distortion in the dehaze image. Finally, we conduct extensive experiments using synthetic and natural hazy images. The quantitative and qualitative evaluation on different benchmark hazy datasets verify the superiority of the proposed method over other state-of-the-art image dehazing techniques. Moreover, additional experimentation validates the applicability of the proposed method in the object detection tasks. Considering the lightweight architecture with minimal computational cost, the proposed system is encouraged to be incorporated as an integral part of the vision-based monitoring systems to improve the overall performance. Hayat Ullah, Khan Muhammad 0001, Saeed Anwar, Ali Shariq Imran, Victor Hugo C. de Albuquerque |
IEEE Trans. Image Process. | 6 |
| 2019 | Evaluating learners' emotional states by monitoring brain waves for comparing game-based learning approach to pen-and-paperabstractA new interest in the use of game factors while acquiring new knowledge has emerged, and a number of researchers are investigating the effectiveness of the game-based approach in education systems. Recent research in game-based learning suggests that this approach imparts learning by involving learners in the learning process. The game factors generate affective-cognitive reactions that absorb users in playing the game and positively influence the learning. This paper offers a comparison of the learning processes between the game-based learning and pen-and-paper approaches. In this paper the analysis of both learning approaches is realized through a brain-controlled technology, using the Emotiv EEG Tech headset, by analyzing the stress, excitement, relaxation, focus, interest, and engagement that the learner is experiencing while going through both approaches. Krenare Pireva Nuci, Rabail Tahir, Ali Shariq Imran, Niraj Chaudhary |
FIE | 3 |
| 2019 | Text-Independent Speaker ID Employing 2D-CNN for Automatic Video Lecture Categorization in a MOOC SettingabstractA new form of distance and blended education has hit the market in recent years with the advent of massive open online courses (MOOCs) which have brought many opportunities to the educational sector. Consequently, the availability of learning content to vast demographics of people and across locations has opened up a plethora of possibilities for everyone to gain new knowledge through MOOCs. This poses an immense issue to the content providers as the amount of manual effort required to structure properly and to organize the content automatically for millions of video lectures daily become incredibly challenging. This paper, therefore, addresses this issue as a small part of our proposed personalized content management system by exploiting the voice pattern of the lecturer for identification and for classifying video lectures to the right speaker category. The use of Mel frequency Cepstral coefficients (MFCC) as 2D input features maps to 2D-CNN has shown promising results in contrast to machine learning and deep learning classifiers - making text-independent speaker identification plausible in MOOC setting for automatic video lecture categorization. It will not only help categorize educational videos efficiently for easy search and retrieval but will also promote effective utilization of micro-lectures and multimedia video learning objects (MLO). Ali Shariq Imran, Zenun Kastrati, Torbjørn Svendsen, Arianit Kurti |
ICTAI | 1 |
| 2019 | A Phonetic-Level Analysis of Different Input Features for Articulatory InversionabstractThe challenge of articulatory inversion is to determine the tem- poral movement of the articulators from the speech waveform, or from acoustic-phonetic knowledge, e.g. derived from infor- mation about the linguistic content of the utterance. The actual position of the articulators is typically obtained from measured data, in our case position measurements obtained using EMA (Electromagnetic articulography). In this paper, we investigate the impact on articulatory inversion problem by using features derived from the acoustic waveform relative to using linguis- tic features related to the time aligned phone sequence of the utterance. Filterbank energies (FBE) are used as acoustic fea- tures, while phoneme identities and (binary) phonetic attributes are used as linguistic features. Experiments are performed on a speech corpus with synchronously recorded EMA measure- ments and employing a bidirectional long short-term memory (BLSTM) that estimates the articulators’ position. Acoustic FBE features performed better for vowel sounds. Phonetic fea- tures attained better results for nasal and fricative sounds except for /h/. Further improvements were obtained by combining FBE and linguistic features, which led to an average relative RMSE reduction of 9.8%, and a 3% relative improvement of the Pearson correlation coefficient. Abdolreza Sabzi Shahrebabaki, Negar Olfati, Ali Shariq Imran, Sabato Marco Siniscalchi, Torbjørn Svendsen |
INTERSPEECH | 3 |
| 2019 | Performance analysis of machine learning classifiers on improved concept vector space modelsabstractThis paper provides a comprehensive performance analysis of parametric and non-parametric machine learning classifiers including a deep feed-forward multi-layer perceptron (MLP) network on two variants of improved Concept Vector Space (iCVS) model. In the first variant, a weighting scheme enhanced with the notion of concept importance is used to assess weight of ontology concepts. Concept importance shows how important a concept is in an ontology and it is automatically computed by converting the ontology into a graph and then applying one of the Markov based algorithms. In the second variant of iCVS, concepts provided by the ontology and their semantically related terms are used to construct concept vectors in order to represent the document into a semantic vector space. We conducted various experiments using a variety of machine learning classifiers for three different models of document representation. The first model is a baseline concept vector space (CVS) model that relies on an exact/partial match technique to represent a document into a vector space. The second and third model is an iCVS model that employs an enhanced concept weighting scheme for assessing weights of concepts (variant 1), and the acquisition of terms that are semantically related to concepts of the ontology for semantic document representation (variant 2), respectively. Additionally, a comparison between seven different classifiers is performed for all three models using precision, recall, and F1 score. Results for multiple configurations of deep learning architecture are obtained by varying the number of hidden layers and nodes in each layer, and are compared to those obtained with conventional classifiers. The obtained results show that the classification performance is highly dependent upon the choice of a classifier, and that the Random Forest, Gradient Boosting, and Multilayer Perceptron are among the classifiers that performed rather well for all three models. Zenun Kastrati, Ali Shariq Imran |
Future Gener. Comput. Syst. | 2 |
| 2019 | The impact of deep learning on document classification using semantically rich representations
Zenun Kastrati, Ali Shariq Imran, Sule Yildirim Yayilgan |
Inf. Process. Manag. | 2 |
| 2019 | Integrating word embeddings and document topics with deep learning in a video classification framework
Zenun Kastrati, Ali Shariq Imran, Arianit Kurti |
Pattern Recognit. Lett. | 2 |
| 2018 | MOOC dropout prediction using machine learning techniques: Review and research challengesabstractMOOC represents an ultimate way to deliver educational content in higher education settings by providing high-quality educational material to the students throughout the world. Considering the differences between traditional learning paradigm and MOOCs, a new research agenda focusing on predicting and explaining dropout of students and low completion rates in MOOCs has emerged. However, due to different problem specifications and evaluation metrics, performing a comparative analysis of state-of-the-art machine learning architectures is a challenging task. In this paper, we provide an overview of the MOOC student dropout prediction phenomenon where machine learning techniques have been utilized. Furthermore, we highlight some solutions being used to tackle with dropout problem, provide an analysis about the challenges of prediction models, and propose some valuable insights and recommendations that might lead to developing useful and effective machine learning solutions to solve the MOOC dropout problem. Fisnik Dalipi, Ali Shariq Imran, Zenun Kastrati |
EDUCON | 2 |
| 2018 | Performance Analysis of Vehicle Detection Techniques: A Concise Survey
Adnan Hanif, Atif Bin Mansoor, Ali Shariq Imran |
WorldCIST (2) | 3 |
| 2016 | HoG based real-time multi-target tracking in Bayesian frameworkabstractMulti-target tracking is one of the most challenging tasks in computer vision. Several complex techniques have been proposed in the literature to tackle the problem. The main idea of such approaches is to find an optimal set of trajectories within a temporal window. The performance of such approaches are fairly good but their computational complexity is too high making them unpractical. In this paper, we propose a novel tracking-by-detection approach in a Bayesian filtering framework. The appearance of a target is modeled through HoG descriptor and the critical problem of target association is solved through combinatorial optimization. It is a simple yet very efficient approach and experimental results show that it achieves state-of-the-art performance in real time. Mohib Ullah, Faouzi Alaya Cheikh, Ali Shariq Imran |
AVSS | 3 |
| 2016 | SEMCON: A Semantic and Contextual Objective Metric for Enriching Domain Ontology ConceptsabstractThis paper presents a novel concept enrichment objective metric combining contextual and semantic information of terms extracted from the domain documents. The proposed metric is called SEMCON which stands for semantic and contextual objective metric. It employs a hybrid learning approach utilizing functionalities from statistical and linguistic ontology learning techniques. The metric also introduced for the first time two statistical features that have shown to improve the overall score ranking of highly relevant terms for concept enrichment. Subjective and objective experiments are conducted in various domains. Experimental results (F1) from computer domain show that SEMCON achieved better performance in contrast to tf*idf, and LSA methods, with 12.2%, 21.8%, and 24.5% improvement over them respectively. Additionally, an investigation into how much each of contextual and semantic components contributes to the overall task of concept enrichment is conducted and the obtained results suggest that a balanced weight gives the best performance. Zenun Kastrati, Ali Shariq Imran, Sule Yildirim Yayilgan |
Int. J. Semantic Web Inf. Syst. | 2 |
| 2015 | Using Context-Aware and Semantic Similarity Based Model to Enrich Ontology Concepts
Zenun Kastrati, Sule Yildirim Yayilgan, Ali Shariq Imran |
NLDB | 3 |
| 2014 | Adaptive Concept Vector Space Representation Using Markov Chain Model
Zenun Kastrati, Ali Shariq Imran |
EKAW | 2 |
| 2011 | Blackboard content classification for lecture videosabstractIn this paper, we propose a novel approach to understand the high level semantics of instructional video by identifying mid-level features from the lecture content. The lecture content in instructional videos can be divided into text, equations and figures. In unscripted lecture video, these visual contents can be useful visual cues to understand the high level semantics. For example, it could help us achieve efficient structuring and indexing of multimedia learning material. To understand the high level semantics from the content itself is however not a trivial task. To this end, we propose visual content classification system (VCCS) for multimedia lecture videos. We propose hybrid approach by combining support vector machine (SVM) and optical character recognition (OCR) to classify visual content into figures, text and equations. The initial results show overall classification accuracy above 85 percent. Ali Shariq Imran, Faouzi Alaya Cheikh |
ICIP | 1 |
| 2009 | Interactive media learning object in distance and blended educationabstractLecture videos contain most of the instructional content. These videos contents are in general non-scripted and unedited and thus do not provide the required level of interactivity from such material. Therefore, these videos fail to captivate students' attention for long and thus their effective use remains a challenge. In this regard, Media Learning Object (MLO) can play an important role in future flexible education. MLOs are media rich documents that have pedagogical values encapsulated with all needed features such as navigation structures and surrogates. The aim of this research work is to develop a framework for the creation, storage, distribution and evaluation of MLOs. We will use the developed framework in pilot courses and evaluate the learning experience and outcome of teaching experience based on MLOs. The ultimate goal is to develop this framework as a platform with a set of tools for processing and analyzing the media collected from the different courses and their annotation with automatically extracted meta-data. Additionally, media browsing and interaction structures will be developed and embedded in the created MLOs to facilitate the interaction and use of such objects in everyday studies. Ali Shariq Imran |
ACM Multimedia | 1 |