Matloob Khushi

dblp:212/0065 · DBLP profile ↗
← Back
33ranked-venue papers
3as first author
20since 2021 · last 2025
0000-0001-7792-2327ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Leveraging BiLSTM-GAT for enhanced stock market prediction: a dual-graph approach to portfolio optimization
abstract
Abstract Stock price prediction remains a critical challenge in financial research due to its potential to inform strategic decision-making. Existing approaches predominantly focus on two key tasks: (1) regression, which forecasts future stock prices, and (2) classification, which identifies trading signals such as buy, sell, or hold. However, the inherent limitations of financial data hinder effective model training, often leading to suboptimal performance. To mitigate this issue, prior studies have expanded datasets by aggregating historical data from multiple companies. This strategy, however, fails to account for the unique characteristics and interdependencies among individual stocks, thereby reducing predictive accuracy. To address these limitations, we propose a novel BiLSTM-GAT-AM model that integrates bidirectional long short-term memory (BiLSTM) networks with graph attention networks (GAT) and an attention mechanism (AM). Unlike conventional graph-based models that define edges based solely on technical or fundamental relationships, our approach employs a dual-graph structure: one graph captures technical similarities, while the other encodes fundamental industry relationships. These two representations are aligned through an attention mechanism, enabling the model to exploit both technical and fundamental insights for enhanced stock market predictions. We conduct extensive experiments, including ablation studies and comparative evaluations against baseline models. The results demonstrate that our model achieves superior predictive performance. Furthermore, leveraging the model’s forecasts, we construct an optimized portfolio and conduct backtesting on the test dataset. Empirical results indicate that our portfolio consistently outperforms both baseline models and the S&P 500 index, highlighting the effectiveness of our approach in stock market prediction and portfolio optimization.
Xiaobin Lu, Josiah Poon, Matloob Khushi
Appl. Intell.3
2024 Vaccine Misinformation Detection in X using Cooperative Multimodal Framework
abstract
Identifying social media posts that spread vaccine misinformation can inform emerging public health risks and aid in designing effective communication interventions. Existing studies, while promising, often rely on single user posts, potentially leading to flawed conclusions. This highlights the necessity to model users' historical posts for a comprehensive understanding of their stance towards vaccines. However, users' historical posts may contain a diverse range of content that adds noise and leads to low performance. To address this gap, in this study, we present VaxMine, a cooperative multi-agent reinforcement learning method that automatically selects relevant textual and visual content from a user's posts, reducing noise. To evaluate the performance of the proposed method, we create and release a new dataset of 2,072 users with historical posts due to the unavailability of publicly available datasets. The experimental results show that our approach outperforms state-of-the-art methods with an F1-Score of 0.94 (an absolute increase of 13%), demonstrating that extracting relevant content from users' historical posts and understanding both modalities are essential to detecting anti-vaccine users on social media. We further analyze the robustness and generalizability of VaxMine, showing that extracting relevant textual and visual content from a user's posts improves performance. We conclude with a discussion of the practical implications of our study by explaining how computational methods used in surveillance can benefit from our work, with flow-on effects on the design of health communication interventions to counter vaccine misinformation on social media.
Usman Naseem, Adam G. Dunn, Matloob Khushi, Jinman Kim
ACM Multimedia3
2024 A Linguistic Grounding-Infused Contrastive Learning Approach for Health Mention Classification on Social Media
abstract
Social media users use disease and symptoms words in different ways, including describing their personal health experiences figuratively or in other general discussions. The health mention classification (HMC) task aims to separate how people use terms, which is important in public health applications. Existing HMC studies address this problem using pretrained language models (PLMs). However, the remaining gaps in the area include the need for linguistic grounding, the requirement for large volumes of labelled data, and that solutions are often only tested on Twitter or Reddit, which provides limited evidence of the transportability of models. To address these gaps, we propose a novel method that uses a transformer-based PLM to obtain a contextual representation of target (disease or symptom) terms coupled with a contrastive loss to establish a larger gap between target terms' literal and figurative uses using linguistic theories. We introduce the use of a simple and effective approach for harvesting candidate instances from the broad corpus and generalising the proposed method using self-training to address the label scarcity challenge. Our experiments on publicly available health-mention datasets from Twitter (HMC2019) and Reddit (RHMD) demonstrate that our method outperforms the state-of-the-art HMC methods on both datasets for the HMC task. We further analyse the transferability and generalisability of our method and conclude with a discussion on the empirical and ethical considerations of our study.
Usman Naseem, Jinman Kim, Matloob Khushi, Adam G. Dunn
WSDM3
2024 Hybrid Text Representation for Explainable Suicide Risk Identification on Social Media
abstract
Social media data that characterize users can provide mental health signals, including suicide risks. Existing methods for suicide risk identification on social media have demonstrated promising results; however, the limitation of existing methods is that they are unable to capture low-and high-level features with complex structured data on social media and are incapable of explaining the predicted labels. Explainable models are more useful when translated, so we aimed to evaluate a novel method that would produce explainable models. This article presents a hybrid text representation method that integrates word and document-level text representations to explain suicide risk identification on social media. The proposed method is then fed to a transformer-based encoder with ordinal classification to determine suicide risk. Our results show that our method outperforms state-of-the-art baselines with an FScore of 0.79 (an absolute increase of 15%) on a public suicide dataset. Our method shows that an explainable model can perform at a comparable level to the best nonexplainable models but has advantages if translated for use in clinical and public health practice.
Usman Naseem, Matloob Khushi, Jinman Kim, Adam G. Dunn
IEEE Trans. Comput. Soc. Syst.2
2024 K-PathVQA: Knowledge-Aware Multimodal Representation for Pathology Visual Question Answering
abstract
Pathology imaging is routinely used to detect the underlying effects and causes of diseases or injuries. Pathology visual question answering (PathVQA) aims to enable computers to answer questions about clinical visual findings from pathology images. Prior work on PathVQA has focused on directly analyzing the image content using conventional pretrained encoders without utilizing relevant external information when the image content is inadequate. In this paper, we present a knowledge-driven PathVQA (K-PathVQA), which uses a medical knowledge graph (KG) from a complementary external structured knowledge base to infer answers for the PathVQA task. K-PathVQA improves the question representation with external medical knowledge and then aggregates vision, language, and knowledge embeddings to learn a joint knowledge-image-question representation. Our experiments using a publicly available PathVQA dataset showed that our K-PathVQA outperformed the best baseline method with an increase of 4.15% in accuracy for the overall task, an increase of 4.40% in open-ended question type and an absolute increase of 1.03% in closed-ended question types. Ablation testing shows the impact of each of the contributions. Generalizability of the method is demonstrated with a separate medical VQA dataset.
Usman Naseem, Matloob Khushi, Adam G. Dunn, Jinman Kim
IEEE J. Biomed. Health Informatics2
2023 Fed-mSSA: A Federated Approach for Spatio-Temporal Data Modeling Using Multivariate Singular Spectrum Analysis
abstract
In modern cyber-physical systems, the vast interconnected processes generated from sensor networks necessitate advanced modeling techniques to exploit decentralized data considering edge computation and data access issues. As sensors emit correlated real-life time series, successful forecasting hinges on revealing the spatio-temporal structures and qualities of data. Matrix Estimation-based (ME) methods, as state-of-the-art techniques, excel at denoising and forecasting high-dimensional correlated time series by representing spatio-temporal data as a temporal matrix. However, ME methods face challenges in handling the decentralized data and access restrictions, due to existing licensing agreements and the inherent burden of centralized modeling. To address this limitation, we propose the Federated Multivariate Singular Spectrum Analysis (Fed-mSSA), a federated matrix estimation-based framework, to denoise and predict correlated time series in the presence of noisy and decentralized data. Specifically, we introduce a novel consensus optimization problem to jointly learn the low-rank matrix representation, capturing spatio-temporal patterns to recover latent states and missing data. Furthermore, we present a federated prediction method that privately and efficiently extracts non-linear temporal dynamics using the denoised temporal matrix. Our results show that our proposed framework achieves state-of-the-art prediction performance in a distributed setting, particularly in the presence of missing data
Jiayu He, Matloob Khushi, Tung-Anh Nguyen, Nguyen Hoang Tran
ICDM2
2023 A Multimodal Framework for the Identification of Vaccine Critical Memes on Twitter
abstract
Memes can be a useful way to spread information because they are funny, easy to share, and can spread quickly and reach further than other forms. With increased interest in COVID-19 vaccines, vaccination-related memes have grown in number and reach. Memes analysis can be difficult because they use sarcasm and often require contextual understanding. Previous research has shown promising results but could be improved by capturing global and local representations within memes to model contextual information. Further, the limited public availability of annotated vaccine critical memes datasets limit our ability to design computational methods to help design targeted interventions and boost vaccine uptake. To address these gaps, we present VaxMeme, which consists of 10,244 manually labelled memes. With VaxMeme, we propose a new multimodal framework designed to improve the memes' representation by learning the global and local representations of memes. The improved memes' representations are then fed to an attentive representation learning module to capture contextual information for classification using an optimised loss function. Experimental results show that our framework outperformed state-of-the-art methods with an F1-Score of 84.2%. We further analyse the transferability and generalisability of our framework and show that understanding both modalities is important to identify vaccine critical memes on Twitter. Finally, we discuss how understanding memes can be useful in designing shareable vaccination promotion, myth debunking memes and monitoring their uptake on social media platforms.
Usman Naseem, Jinman Kim, Matloob Khushi, Adam G. Dunn
WSDM3
2023 ConanVarvar: a versatile tool for the detection of large syndromic copy number variation from whole-genome sequencing data
abstract
BACKGROUND: A wide range of tools are available for the detection of copy number variants (CNVs) from whole-genome sequencing (WGS) data. However, none of them focus on clinically-relevant CNVs, such as those that are associated with known genetic syndromes. Such variants are often large in size, typically 1-5 Mb, but currently available CNV callers have been developed and benchmarked for the discovery of smaller variants. Thus, the ability of these programs to detect tens of real syndromic CNVs remains largely unknown. RESULTS: Here we present ConanVarvar, a tool which implements a complete workflow for the targeted analysis of large germline CNVs from WGS data. ConanVarvar comes with an intuitive R Shiny graphical user interface and annotates identified variants with information about 56 associated syndromic conditions. We benchmarked ConanVarvar and four other programs on a dataset containing real and simulated syndromic CNVs larger than 1 Mb. In comparison to other tools, ConanVarvar reports 10-30 times less false-positive variants without compromising sensitivity and is quicker to run, especially on large batches of samples. CONCLUSIONS: ConanVarvar is a useful instrument for primary analysis in disease sequencing studies, where large CNVs could be the cause of disease.
Mikhail Gudkov, Loïc Thibaut, Matloob Khushi, Gillian M. Blue, David S. Winlaw, Sally L. Dunwoodie, Eleni Giannoulatou
BMC Bioinform.3
2023 RHMD: A Real-World Dataset for Health Mention Classification on Reddit
abstract
People on social media share their thoughts and experiences using diseases and symptoms words other than to mention their health, which can introduce biases in data-driven public health applications. For the advancement of HMC research, in this study, we present a Reddit health mention dataset (RHMD), a new dataset of multi-domain Reddit data for the HMC. RHMD is composed of 10015 manually annotated Reddit posts that include 15 common disease or symptom terms and are labeled with four labels: personal health mentions (HMs), nonpersonal HMs, figurative HMs, and hyperbolic HMs. Empirical evaluation using recently proposed methods demonstrates the challenge of labeling user-generated text across these four types. Contributions to this work include the public release of a robustly annotated Reddit dataset (RHMD) for HM tasks and a comprehensive performance analysis of baseline methods. We expect the release of the dataset, and the evaluations will help facilitate the development of new methods for detecting HMs in the user-generated text. The dataset is available athttps://github.com/usmaann/RHMD-Health-Mention-Dataset.
Usman Naseem, Matloob Khushi, Jinman Kim, Adam G. Dunn
IEEE Trans. Comput. Soc. Syst.2
2023 Vision-Language Transformer for Interpretable Pathology Visual Question Answering
abstract
Pathology visual question answering (PathVQA) attempts to answer a medical question posed by pathology images. Despite its great potential in healthcare, it is not widely adopted because it requires interactions on both the image (vision) and question (language) to generate an answer. Existing methods focused on treating vision and language features independently, which were unable to capture the high and low-level interactions that are required for VQA. Further, these methods failed to offer capabilities to interpret the retrieved answers, which are obscure to humans where the models' interpretability to justify the retrieved answers has remained largely unexplored. Motivated by these limitations, we introduce a vision-language transformer that embeds vision (images) and language (questions) features for an interpretable PathVQA. We present an interpretable transformer-based Path-VQA (TraP-VQA), where we embed transformers' encoder layers with vision and language features extracted using pre-trained CNN and domain-specific language model (LM), respectively. A decoder layer is then embedded to upsample the encoded features for the final prediction for PathVQA. Our experiments showed that our TraP-VQA outperformed the state-of-the-art comparative methods with public PathVQA dataset. Our experiments validated the robustness of our model on another medical VQA dataset, and the ablation study demonstrated the capability of our integrated transformer-based vision-language model for PathVQA. Finally, we present the visualization results of both text and images, which explain the reason for a retrieved answer in PathVQA.
Usman Naseem, Matloob Khushi, Jinman Kim
IEEE J. Biomed. Health Informatics2
2022 Early Identification of Depression Severity Levels on Reddit Using Ordinal Classification
abstract
User-generated text on social media is a promising avenue for public health surveillance and has been actively explored for its feasibility in the early identification of depression. Existing methods in the identification of depression have shown promising results; however, these methods were all focused on treating the identification as a binary classification problem. To date, there has been little effort towards identifying users’ depression severity level and disregard the inherent ordinal nature across these fine-grain levels. This paper aims to make early identification of depression severity levels on social media data. To accomplish this, we built a new dataset based on the inherent ordinal nature over depression severity levels using clinical depression standards on Reddit posts. The posts were classified into 4 depression severity levels covering the clinical depression standards on social media. Accordingly, we reformulate the early identification of depression as an ordinal classification task over clinical depression standards such as Beck’s Depression Inventory and the Depressive Disorder Annotation scheme to identify depression severity levels. With these, we propose a hierarchical attention method optimized to factor in the increasing depression severity levels through a soft probability distribution. We experimented using two datasets (a public dataset having more than one post from each user and our built dataset with a single user post) using real-world Reddit posts that have been classified according to questionnaires built by clinical experts and demonstrated that our method outperforms state-of-the-art models. Finally, we conclude by analyzing the minimum number of posts required to identify depression severity level followed by a discussion of empirical and practical considerations of our study.
Usman Naseem, Adam G. Dunn, Jinman Kim, Matloob Khushi
WWW4
2022 Identification of Disease or Symptom terms in Reddit to Improve Health Mention Classification
abstract
In a user-generated text such as on social media platforms and online forums, people often use disease or symptom terms in ways other than to describe their health. In data-driven public health surveillance, the health mention classification (HMC) task aims to identify posts where users are discussing health conditions rather than using disease and symptom terms for other reasons. Existing computational research typically only studies health mentions in Twitter, with limited coverage of disease or symptom terms, ignore user behavior information, and other ways people use disease or symptom terms. To advance the HMC research, we present a Reddit health mention dataset (RHMD), a new dataset of multi-domain Reddit data for the HMC. RHMD consists of 10,015 manually labeled Reddit posts that mention 15 common disease or symptom terms and are annotated with four labels: namely personal health mentions, non-personal health mentions, figurative health mentions, and hyperbolic health mentions. With RHMD, we propose HMCNET that combines a target keyword (disease or symptom term) identification and user behavior hierarchically to improve HMC. Experimental results demonstrate that the proposed approach outperforms state-of-the-art methods with an F1-Score of 0.75 (an increase of 11% over the state-of-the-art) and shows that our new dataset poses a strong challenge to the existing HMC methods.
Usman Naseem, Jinman Kim, Matloob Khushi, Adam G. Dunn
WWW3
2022 A patient network-based machine learning model for disease prediction: The case of type 2 diabetes mellitus
Haohui Lu, Shahadat Uddin, Farshid Hajati, Mohammad Ali Moni, Matloob Khushi
Appl. Intell.5
2022 Benchmarking for biomedical natural language processing tasks with a domain specific ALBERT
abstract
BACKGROUND: The abundance of biomedical text data coupled with advances in natural language processing (NLP) is resulting in novel biomedical NLP (BioNLP) applications. These NLP applications, or tasks, are reliant on the availability of domain-specific language models (LMs) that are trained on a massive amount of data. Most of the existing domain-specific LMs adopted bidirectional encoder representations from transformers (BERT) architecture which has limitations, and their generalizability is unproven as there is an absence of baseline results among common BioNLP tasks. RESULTS: We present 8 variants of BioALBERT, a domain-specific adaptation of a lite bidirectional encoder representations from transformers (ALBERT), trained on biomedical (PubMed and PubMed Central) and clinical (MIMIC-III) corpora and fine-tuned for 6 different tasks across 20 benchmark datasets. Experiments show that a large variant of BioALBERT trained on PubMed outperforms the state-of-the-art on named-entity recognition (+ 11.09% BLURB score improvement), relation extraction (+ 0.80% BLURB score), sentence similarity (+ 1.05% BLURB score), document classification (+ 0.62% F1-score), and question answering (+ 2.83% BLURB score). It represents a new state-of-the-art in 5 out of 6 benchmark BioNLP tasks. CONCLUSIONS: The large variant of BioALBERT trained on PubMed achieved a higher BLURB score than previous state-of-the-art models on 5 of the 6 benchmark BioNLP tasks. Depending on the task, 5 different variants of BioALBERT outperformed previous state-of-the-art models on 17 of the 20 benchmark datasets, showing that our model is robust and generalizable in the common BioNLP tasks. We have made BioALBERT freely available which will help the BioNLP community avoid computational cost of training and establish a new set of baselines for future efforts across a broad range of BioNLP tasks.
Usman Naseem, Adam G. Dunn, Matloob Khushi, Jinman Kim
BMC Bioinform.3
2022 Comorbidity and multimorbidity prediction of major chronic diseases using machine learning and network analytics
Shahadat Uddin, Shangzhou Wang, Haohui Lu, Arif Khan 0001, Farshid Hajati, Matloob Khushi
Expert Syst. Appl.6
2021 Classifying vaccine sentiment tweets by modelling domain-specific representation and commonsense knowledge into context-aware attentive GRU
abstract
Vaccines are an important public health measure, but vaccine hesitancy and refusal can create clusters of low vaccine coverage and reduce the effectiveness of vaccination programs. Social media provides an opportunity to estimate emerging risks to vaccine acceptance by including geographical location and detailing vaccine-related concerns. Methods for classifying social media posts, such as vaccine-related tweets, use language models (LMs) trained on general domain text. However, challenges to measuring vaccine sentiment at scale arise from the absence of tonal stress and gestural cues and may not always have additional information about the user, e.g., past tweets or social connections. Another challenge in LMs is the lack of ‘commonsense’ knowledge that are apparent in users' metadata, i.e., emoticons, positive and negative words etc. In this study, to classify vaccine sentiment tweets with limited information, we present a novel end-to-end framework consisting of interconnected components that use domain-specific LM trained on vaccine-related tweets and models commonsense knowledge into a bidirectional gated recurrent network (CK-BiGRU) with context-aware attention. We further leverage syntactical, user metadata and sentiment information to capture the sentiment of a tweet. We experimented using two popular vaccine-related Twitter datasets and demonstrate that our proposed approach outperforms state-of-the-art models in identifying pro-vaccine, anti-vaccine and neutral tweets.
Usman Naseem, Matloob Khushi, Jinman Kim, Adam G. Dunn
IJCNN2
2021 BioALBERT: A Simple and Effective Pre-trained Language Model for Biomedical Named Entity Recognition
abstract
In recent years, with the growing amount of biomedical documents, coupled with advancement in natural language processing algorithms, the research on biomedical named entity recognition (BioNER) has increased exponentially. However, BioNER research is challenging as NER in the biomedical domain are: (i) often restricted due to limited amount of training data, (ii) an entity can refer to multiple types and concepts depending on its context and, (iii) heavy reliance on acronyms that are sub-domain specific. Existing BioNER approaches often neglect these issues and directly adopt the state-of-the-art (SOTA) models trained in general corpora, which often yields unsatisfactory results. We propose biomedical ALBERT (A Lite Bidirectional Encoder Representations from Transformers for Biomedical Text Mining) - bioALBERT - an effective domain-specific pre-trained language model trained on a huge biomedical corpus designed to capture biomedical context-dependent NER. We adopted a self-supervised loss function used in ALBERT that targets modelling inter-sentence coherence to better learn context-dependent representations and incorporated parameter reduction strategies to minimise memory usage and enhance the training time in BioNER. In our experiments, BioALBERT outperformed comparative SOTA BioNER models on 8 biomedical NER benchmark datasets with 4 different entity types. The performance is increased for; (i) disease type corpora by 7.47% (NCBI- disease) and 10.63% (BC5CDR-disease); (ii) drug-chem type corpora by 4.61 % (BC5CDR-Chem) and 3.89% (BC4CHEMD); (iii) gene-protein type corpora by 12.25% (BC2GM) and 6.42% (JNLPBA); and (iv) species type corpora by 6.19% (LINNAEUS) and 23.71 % (Species-800) is observed which leads to a state-of-the-art results. The performance of a proposed model on four different biomedical entity types shows that our model is robust and generalisable in recognising biomedical entities in text.
Usman Naseem, Matloob Khushi, Vinay Reddy, Sakthivel Rajendran, Muhammad Imran Razzak, Jinman Kim
IJCNN2
2021 Robust Dual Recurrent Neural Networks for Financial Time Series Prediction
abstract
Various recurrent neural network (RNN) architectures have been implemented successfully for time series prediction in recent years.However, real-world time series data usually contain noise, which decreases the performance of the neural networks.Despite the substantial efforts to understand the pattern of time series, there is a lack of research on detecting and filtering out the inherent noise when predicting time series based on training RNN models.We propose a dual RNN strategy, namely Robust Dual Recurrent Neural Networks (RDRNN), for noisy time series prediction.We designed and trained two RNNs simultaneously and used the loss value to classify different samples into noise-free samples and noisy samples.We exchanged the small-loss samples (which were likely to be noise-free data) to fit the main pattern of time series data, and re-weighted the large-loss samples (which were likely to be noisy data) to alleviate the impact of noise.Empirical results on three popular Chinese stock market indexes demonstrate that the new learning paradigm significantly outperforms baseline approaches.Our code is available at https://jiayuheusyd.github.io/
Jiayu He, Matloob Khushi, Nguyen Hoang Tran, Tongliang Liu
SDM2
2021 Corporate Bankruptcy Prediction: An Approach Towards Better Corporate World
abstract
Abstract The area of corporate bankruptcy prediction attains high economic importance, as it affects many stakeholders. The prediction of corporate bankruptcy has been extensively studied in economics, accounting and decision sciences over the past two decades. The corporate bankruptcy prediction has been a matter of talk among academic literature and professional researchers throughout the world. Different traditional approaches were suggested based on hypothesis testing and statistical modeling. Therefore, the primary purpose of the research is to come up with a model that can estimate the probability of corporate bankruptcy by evaluating its occurrence of failure using different machine learning models. As the dataset was not well prepared and contains missing values, various data mining and data pre-processing techniques were utilized for data preparation. Within this research, the task of resolving the issues induced by the imbalance between the two classes is approached by applying different data balancing techniques. We address the problem of imbalanced data with the random undersampling and Synthetic Minority Over Sampling Technique (SMOTE). We used five machine learning models (support vector machine, J48 decision tree, Logistic model tree, random forest and decision forest) to predict corporate bankruptcy earlier to the occurrence. We use data from 2009 to 2013 on Poland manufacturing corporates and selected the 64 financial indicators to be broken down. The main finding of the study is a significant improvement in predictive accuracy using machine learning techniques. We also include other economic indicators ratios, along with Altman’s Z-score variables related to profitability, liquidity, leverage and solvency (short/long term) to propose an efficient model. Machine learning models give better results while balancing the data through SMOTE as compared to random undersampling. The machine learning technique related to decision forest led to 99% accuracy, whereas support vector machine (SVM), J48 decision tree, Logistic Model Tree (LMT) and Random Forest (RF) led to 92%, 92.3%, 93.8% and 98.7% accuracy, respectively, with all predictive financial indicators. We find that the decision forest outperforms the other techniques and previous techniques discussed in the literature. The proposed method is also deployed on the web to assist regulators, investors, creditors and scholars to predict corporate bankruptcy.
Talha Mahboob Alam, Kamran Shaukat, Mubbashar Mushtaq, Matloob Khushi, Suhuai Luo
Comput. J.5
2021 COVIDSenti: A Large-Scale Benchmark Twitter Data Set for COVID-19 Sentiment Analysis
abstract
Social media (and the world at large) have been awash with news of the COVID-19 pandemic. With the passage of time, news and awareness about COVID-19 spread like the pandemic itself, with an explosion of messages, updates, videos, and posts. Mass hysteria manifest as another concern in addition to the health risk that COVID-19 presented. Predictably, public panic soon followed, mostly due to misconceptions, a lack of information, or sometimes outright misinformation about COVID-19 and its impacts. It is thus timely and important to conduct anex post factoassessment of the early information flows during the pandemic on social media, as well as a case study of evolving public opinion on social media which is of general interest. This study aims to inform policy that can be applied to social media platforms; for example, determining what degree of moderation is necessary to curtail misinformation on social media. This study also analyzes views concerning COVID-19 by focusing on people who interact and share social media on Twitter. As a platform for our experiments, we present a new large-scale sentiment data set COVIDSENTI, which consists of 90 000 COVID-19-related tweets collected in the early stages of the pandemic, from February to March 2020. The tweets have been labeled into positive, negative, and neutral sentiment classes. We analyzed the collected tweets for sentiment classification using different sets of features and classifiers. Negative opinion played an important role in conditioning public sentiment, for instance, we observed that people favored lockdown earlier in the pandemic; however, as expected, sentiment shifted by mid-March. Our study supports the view that there is a need to develop a proactive and agile public health presence to combat the spread of negative sentiment on social media following a pandemic.
Usman Naseem, Muhammad Imran Razzak, Matloob Khushi, Peter W. Eklund, Jinman Kim
IEEE Trans. Comput. Soc. Syst.3
2020 Protein-Protein Interactions Prediction Based on Bi-directional Gated Recurrent Unit and Multimodal Representation
Kanchan Jha, Sriparna Saha 0001, Matloob Khushi
ICONIP (5)3
2020 Data Mining ENCODE Data Predicts a Significant Role of SINA3 in Human Liver Cancer
Matloob Khushi, Usman Naseem, Jonathan Du, Anis Khan, Simon K. Poon
ICONIP (3)1
2020 Machine Learned Pulse Transit Time (MLPTT) Measurements from Photoplethysmography
Philip Mehrgardt, Matloob Khushi, Anusha Withana, Simon K. Poon
ICONIP (3)2
2020 Diabetic Retinopathy Detection Using Multi-layer Neural Networks and Split Attention with Focal Loss
Usman Naseem, Matloob Khushi, Shah Khalid Khan, Nazar Waheed, Adnan Mir, Atika Qazi, Bandar AlShammari, Simon K. Poon
ICONIP (3)2
2020 Classification of Neuroblastoma Histopathological Images Using Machine Learning
Adhish Panta, Matloob Khushi, Usman Naseem, Paul J. Kennedy, Daniel R. Catchpoole
ICONIP (3)2
2020 Statistical and Geometrical Alignment using Metric Learning in Domain Adaptation
abstract
Domain adapted machine learning is driven by the possibilities of learning from source data distribution to understand different target data distributions. An assumption is made that one application (source) domain always has enough labeled information, but the other related application (target) may contain information that is partially labeled or completely unlabeled. Therefore, it is necessary to train the target domain classifier using enough labeled information of the source domain. However, contrary to primitive assumptions, the source domain and target domain data need not have the same distribution. Therefore, we can't directly use data of source domain to train classifier for data of target domain. Existing approaches can be deprived of one or more objectives: perform geometric diffusion on the manifold, align the cross-domain distributions, preserve the discriminative information using metric learning. Here, we have proposed a novel framework that aims to meet all such objectives. In this framework, we proposed two methods, statistical and geometrical alignment using metric learning with pseudo labels (SGA-MDAP) and without pseudo labels (SGA-MDA) in visual domain adaptation. It has been demonstrated through various experiments that our framework outperforms various state-of-the-art methods over four different real-world cross-domain visual identification datasets such as PIE face, ORL face, Yale face, and Office Caltech.
Rakesh Kumar Sanodiya, Alwyn Mathew, Jimson Mathew, Matloob Khushi
IJCNN4
2020 Wavelet Denoising and Attention-based RNN- ARIMA Model to Predict Forex Price
abstract
Every change of trend in the forex market presents a great opportunity as well as a risk for investors. Accurate forecasting of forex prices is a crucial element in any effective hedging or speculation strategy. However, the complex nature of the forex market makes the predicting problem challenging, which has prompted extensive research from various academic disciplines. In this paper, a novel approach that integrates the wavelet denoising, Attention-based Recurrent Neural Network (ARNN), and Autoregressive Integrated Moving Average (ARIMA) are proposed. Wavelet transform removes the noise from the time series to stabilize the data structure. ARNN model captures the robust and non-linear relationships in the sequence and ARIMA can well fit the linear correlation of the sequential information. By hybridization of the three models, the methodology is capable of modelling dynamic systems such as the forex market. Our experiments on USD/JPY five-minute data outperforms the baseline methods. Root-Mean-Squared-Error (RMSE) of the hybrid approach was found to be 1.65 with a directional accuracy of ~76%.
Matloob Khushi
IJCNN2
2020 GA-MSSR: Genetic Algorithm Maximizing Sharpe and Sterling Ratio Method for RoboTrading
abstract
Foreign exchange is the largest financial market in the world, and it is also one of the most volatile markets. Technical analysis plays an important role in the forex market and trading algorithms are designed utilizing machine learning techniques. Most literature used historical price information and technical indicators for training. However, the noisy nature of the market affects the consistency and profitability of the algorithms. To address this problem, we designed trading rule features that are derived from technical indicators and trading rules. The parameters of technical indicators are optimized to maximize trading performance. We also proposed a novel cost function that computes the risk-adjusted return, Sharpe and Sterling Ratio (SSR), in an effort to reduce the variance and the magnitude of drawdowns. An automatic robotic trading (RoboTrading) strategy is designed with the proposed Genetic Algorithm Maximizing Sharpe and Sterling Ratio model (GA-MSSR) model. The experiment was conducted on intraday data of 6 major currency pairs from 2018 to 2019. The results consistently showed significant positive returns and the performance of the trading system is superior using the optimized rule-based features. The highest return obtained was 320% annually using 5-minute AUDUSD currency pair. Besides, the proposed model achieves the best performance on risk factors, including maximum drawdowns and variance in return, comparing to benchmark models. The code can be accessed at https://github.com/zzzac/rule-based-forextrading-system.
Zezheng Zhang, Matloob Khushi
IJCNN2
2019 Machine Learning Based Method for Huntington's Disease Gait Pattern Recognition
Xiuyu Huang, Matloob Khushi, Mark Latt, Clement Loy, Simon K. Poon
ICONIP (4)2
2019 Semi-supervised Regularized Coplanar Discriminant Analysis
Rakesh Kumar Sanodiya, Michelle Davies Thalakottur, Jimson Mathew, Matloob Khushi
ICONIP (5)4
2019 IMDB-Attire: A Novel Dataset for Attire Detection and Localization
Saad Bin Yousuf, Hasan Sajid, Simon K. Poon, Matloob Khushi
ICONIP (2)4
2018 Predicting Functional Interactions Among DNA-Binding Proteins
Matloob Khushi, Nazim Choudhury, Jonathan W. Arthur, Christine L. Clarke, J. Dinny Graham
ICONIP (5)1
2017 Automated classification and characterization of the mitotic spindle following knockdown of a mitosis-related protein
abstract
BACKGROUND: Cell division (mitosis) results in the equal segregation of chromosomes between two daughter cells. The mitotic spindle plays a pivotal role in chromosome alignment and segregation during metaphase and anaphase. Structural or functional errors of this spindle can cause aneuploidy, a hallmark of many cancers. To investigate if a given protein associates with the mitotic spindle and regulates its assembly, stability, or function, fluorescence microscopy can be performed to determine if disruption of that protein induces phenotypes indicative of spindle dysfunction. Importantly, functional disruption of proteins with specific roles during mitosis can lead to cancer cell death by inducing mitotic insult. However, there is a lack of automated computational tools to detect and quantify the effects of such disruption on spindle integrity. RESULTS: We developed the image analysis software tool MatQuantify, which detects both large-scale and subtle structural changes in the spindle or DNA and can be used to statistically compare the effects of different treatments. MatQuantify can quantify various physical properties extracted from fluorescence microscopy images, such as area, lengths of various components, perimeter, eccentricity, fractal dimension, satellite objects and orientation. It can also measure textual properties including entropy, intensities and the standard deviation of intensities. Using MatQuantify, we studied the effect of knocking down the protein clathrin heavy chain (CHC) on the mitotic spindle. We analysed 217 microscopy images of untreated metaphase cells, 172 images of metaphase cells transfected with small interfering RNAs targeting the luciferase gene (as a negative control), and 230 images of metaphase cells depleted of CHC. Using the quantified data, we trained 23 supervised machine learning classification algorithms. The Support Vector Machine learning algorithm was the most accurate method (accuracy: 85.1%; area under the curve: 0.92) for classifying a spindle image. The Kruskal-Wallis and Tukey-Kramer tests demonstrated that solidity, compactness, eccentricity, extent, mean intensity and number of satellite objects (multipolar spindles) significantly differed between CHC-depleted cells and untreated/luciferase-knockdown cells. CONCLUSION: MatQuantify enables automated quantitative analysis of images of mitotic spindles. Using this tool, researchers can unambiguously test if disruption of a protein-of-interest changes metaphase spindle maintenance and thereby affects mitosis.
Matloob Khushi, Imraan M. Dean, Erdahl T. Teber, Megan Chircop, Jonathan W. Arthur, Neftali Flores-Rodriguez
BMC Bioinform.1