EDBT 2026 Demo / reviewers in the wild / expert
Matloob Khushi
dblp:212/0065
· DBLP profile ↗
6ranked-venue papers in the field
0as first author
6since 2021 · last 2024
0000-0001-7792-2327ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 4Information Retrieval & Web Search · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Linguistic Grounding-Infused Contrastive Learning Approach for Health Mention Classification on Social MediaabstractSocial media users use disease and symptoms words in different ways, including describing their personal health experiences figuratively or in other general discussions. The health mention classification (HMC) task aims to separate how people use terms, which is important in public health applications. Existing HMC studies address this problem using pretrained language models (PLMs). However, the remaining gaps in the area include the need for linguistic grounding, the requirement for large volumes of labelled data, and that solutions are often only tested on Twitter or Reddit, which provides limited evidence of the transportability of models. To address these gaps, we propose a novel method that uses a transformer-based PLM to obtain a contextual representation of target (disease or symptom) terms coupled with a contrastive loss to establish a larger gap between target terms' literal and figurative uses using linguistic theories. We introduce the use of a simple and effective approach for harvesting candidate instances from the broad corpus and generalising the proposed method using self-training to address the label scarcity challenge. Our experiments on publicly available health-mention datasets from Twitter (HMC2019) and Reddit (RHMD) demonstrate that our method outperforms the state-of-the-art HMC methods on both datasets for the HMC task. We further analyse the transferability and generalisability of our method and conclude with a discussion on the empirical and ethical considerations of our study. Usman Naseem, Jinman Kim, Matloob Khushi, Adam G. Dunn |
WSDM | 3 |
| 2023 | Fed-mSSA: A Federated Approach for Spatio-Temporal Data Modeling Using Multivariate Singular Spectrum AnalysisabstractIn modern cyber-physical systems, the vast interconnected processes generated from sensor networks necessitate advanced modeling techniques to exploit decentralized data considering edge computation and data access issues. As sensors emit correlated real-life time series, successful forecasting hinges on revealing the spatio-temporal structures and qualities of data. Matrix Estimation-based (ME) methods, as state-of-the-art techniques, excel at denoising and forecasting high-dimensional correlated time series by representing spatio-temporal data as a temporal matrix. However, ME methods face challenges in handling the decentralized data and access restrictions, due to existing licensing agreements and the inherent burden of centralized modeling. To address this limitation, we propose the Federated Multivariate Singular Spectrum Analysis (Fed-mSSA), a federated matrix estimation-based framework, to denoise and predict correlated time series in the presence of noisy and decentralized data. Specifically, we introduce a novel consensus optimization problem to jointly learn the low-rank matrix representation, capturing spatio-temporal patterns to recover latent states and missing data. Furthermore, we present a federated prediction method that privately and efficiently extracts non-linear temporal dynamics using the denoised temporal matrix. Our results show that our proposed framework achieves state-of-the-art prediction performance in a distributed setting, particularly in the presence of missing data Jiayu He, Matloob Khushi, Tung-Anh Nguyen, Nguyen Hoang Tran |
ICDM | 2 |
| 2023 | A Multimodal Framework for the Identification of Vaccine Critical Memes on TwitterabstractMemes can be a useful way to spread information because they are funny, easy to share, and can spread quickly and reach further than other forms. With increased interest in COVID-19 vaccines, vaccination-related memes have grown in number and reach. Memes analysis can be difficult because they use sarcasm and often require contextual understanding. Previous research has shown promising results but could be improved by capturing global and local representations within memes to model contextual information. Further, the limited public availability of annotated vaccine critical memes datasets limit our ability to design computational methods to help design targeted interventions and boost vaccine uptake. To address these gaps, we present VaxMeme, which consists of 10,244 manually labelled memes. With VaxMeme, we propose a new multimodal framework designed to improve the memes' representation by learning the global and local representations of memes. The improved memes' representations are then fed to an attentive representation learning module to capture contextual information for classification using an optimised loss function. Experimental results show that our framework outperformed state-of-the-art methods with an F1-Score of 84.2%. We further analyse the transferability and generalisability of our framework and show that understanding both modalities is important to identify vaccine critical memes on Twitter. Finally, we discuss how understanding memes can be useful in designing shareable vaccination promotion, myth debunking memes and monitoring their uptake on social media platforms. Usman Naseem, Jinman Kim, Matloob Khushi, Adam G. Dunn |
WSDM | 3 |
| 2022 | Early Identification of Depression Severity Levels on Reddit Using Ordinal ClassificationabstractUser-generated text on social media is a promising avenue for public health surveillance and has been actively explored for its feasibility in the early identification of depression. Existing methods in the identification of depression have shown promising results; however, these methods were all focused on treating the identification as a binary classification problem. To date, there has been little effort towards identifying users’ depression severity level and disregard the inherent ordinal nature across these fine-grain levels. This paper aims to make early identification of depression severity levels on social media data. To accomplish this, we built a new dataset based on the inherent ordinal nature over depression severity levels using clinical depression standards on Reddit posts. The posts were classified into 4 depression severity levels covering the clinical depression standards on social media. Accordingly, we reformulate the early identification of depression as an ordinal classification task over clinical depression standards such as Beck’s Depression Inventory and the Depressive Disorder Annotation scheme to identify depression severity levels. With these, we propose a hierarchical attention method optimized to factor in the increasing depression severity levels through a soft probability distribution. We experimented using two datasets (a public dataset having more than one post from each user and our built dataset with a single user post) using real-world Reddit posts that have been classified according to questionnaires built by clinical experts and demonstrated that our method outperforms state-of-the-art models. Finally, we conclude by analyzing the minimum number of posts required to identify depression severity level followed by a discussion of empirical and practical considerations of our study. Usman Naseem, Adam G. Dunn, Jinman Kim, Matloob Khushi |
WWW | 4 |
| 2022 | Identification of Disease or Symptom terms in Reddit to Improve Health Mention ClassificationabstractIn a user-generated text such as on social media platforms and online forums, people often use disease or symptom terms in ways other than to describe their health. In data-driven public health surveillance, the health mention classification (HMC) task aims to identify posts where users are discussing health conditions rather than using disease and symptom terms for other reasons. Existing computational research typically only studies health mentions in Twitter, with limited coverage of disease or symptom terms, ignore user behavior information, and other ways people use disease or symptom terms. To advance the HMC research, we present a Reddit health mention dataset (RHMD), a new dataset of multi-domain Reddit data for the HMC. RHMD consists of 10,015 manually labeled Reddit posts that mention 15 common disease or symptom terms and are annotated with four labels: namely personal health mentions, non-personal health mentions, figurative health mentions, and hyperbolic health mentions. With RHMD, we propose HMCNET that combines a target keyword (disease or symptom term) identification and user behavior hierarchically to improve HMC. Experimental results demonstrate that the proposed approach outperforms state-of-the-art methods with an F1-Score of 0.75 (an increase of 11% over the state-of-the-art) and shows that our new dataset poses a strong challenge to the existing HMC methods. Usman Naseem, Jinman Kim, Matloob Khushi, Adam G. Dunn |
WWW | 3 |
| 2021 | Robust Dual Recurrent Neural Networks for Financial Time Series PredictionabstractVarious recurrent neural network (RNN) architectures have been implemented successfully for time series prediction in recent years.However, real-world time series data usually contain noise, which decreases the performance of the neural networks.Despite the substantial efforts to understand the pattern of time series, there is a lack of research on detecting and filtering out the inherent noise when predicting time series based on training RNN models.We propose a dual RNN strategy, namely Robust Dual Recurrent Neural Networks (RDRNN), for noisy time series prediction.We designed and trained two RNNs simultaneously and used the loss value to classify different samples into noise-free samples and noisy samples.We exchanged the small-loss samples (which were likely to be noise-free data) to fit the main pattern of time series data, and re-weighted the large-loss samples (which were likely to be noisy data) to alleviate the impact of noise.Empirical results on three popular Chinese stock market indexes demonstrate that the new learning paradigm significantly outperforms baseline approaches.Our code is available at https://jiayuheusyd.github.io/ Jiayu He, Matloob Khushi, Nguyen Hoang Tran, Tongliang Liu |
SDM | 2 |