VLDB 2026 Research / reviewers in the wild / expert
Jumanah Alshehri
dblp:245/4099 · also Jumanah S. Alshehri
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2022
0000-0002-0077-7173ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | On Label Quality in Class Imbalance Setting -A Case StudyabstractProducing high-quality labeled data is a challenge in any supervised learning problem, where in many cases, human involvement is necessary to ensure the label quality. However, human annotations are not flawless, especially in the case of a challenging problem. In nontrivial problems, the high disagreement among annotators results in noisy labels, which affect the performance of any machine learning model. In this work, we consider three noise reduction strategies to improve the label quality in the Article-Comment Alignment Problem, where the main task is to classify article-comment pairs according to their relevancy level. The first considered labeling disagreement reduction strategy utilizes annotators’ background knowledge during the label aggregation step. The second strategy utilizes user disagreement during the training process. In the third and final strategy, we ask annotators to perform corrections and relabel the examples with noisy labels. We deploy these strategies and compare them to a resampling strategy for addressing the class imbalance, another common supervised learning challenge. These alternatives were evaluated on ACAP, a multiclass text pairs classification problem with highly imbalanced data, where one of the classes represents at most 15% of the dataset’s entire population. Our results provide evidence that considered strategies can reduce disagreement between annotators. However, data quality improvement is insufficient to enhance classification accuracy in the article-comment alignment problem, which exhibits a high-class imbalance. The model performance is enhanced for the same problem by addressing the imbalance issue with a weight loss-based class distribution resampling. We show that allowing the model to pay more attention to the minority class during the training process with the presence of noisy examples improves the test accuracy by 3%. Jumanah Alshehri, Marija Stanojevic, Eduard C. Dragut, Zoran Obradovic |
ICMLA | 1 |
| 2022 | MultiLayerET: A Unified Representation of Entities and Topics Using Multilayer Graphs
Jumanah Alshehri, Marija Stanojevic, Parisa Khan, Benjamin Rapp, Eduard C. Dragut, Zoran Obradovic |
ECML/PKDD (2) | 1 |
| 2021 | Distinguishability of graphs: a case for quantum-inspired measuresabstractThe question of graph similarity or graph distinguishability arises often in natural systems and their analysis over graphical networks. In many domains, graph similarity is used for graph classification, outlier detection or the identification of distinguished interaction patterns. Several methods have been proposed on how to address this topic, but graph comparison still presents many challenges. Recently, information physics has emerged as a promising theoretical foundation for complex networks. In many applications, it has been demonstrated that natural complex systems exhibit features that can be described and interpreted by measures typically applied in quantum mechanical systems. Therefore, a natural starting point for the identification of network similarity measures is information physics and a series of measures of distance for quantum states. In this work, we report experiments on synthetic and real-world data sets, and compare quantum-inspired measures to a series of state-of-the-art and well-established methods of graph distinguishability. We show that quantum-inspired methods satisfy the mathematical and intuitive requirements for graph similarities, while offering high interpretability. Athanasia Polychronopoulou, Jumanah Alshehri, Zoran Obradovic |
ASONAM | 2 |
| 2021 | Stay on Topic, Please: Aligning User Comments to the Content of a News Article
Jumanah Alshehri, Marija Stanojevic, Eduard C. Dragut, Zoran Obradovic |
ECIR (1) | 1 |
| 2019 | Surveying public opinion using label prediction on social media dataabstractIn this study, a procedure is proposed for surveying public opinion from big social media domain-specific textual data to minimize the difficulties associated with modeling public behavior. Strategies for labeling posts relevant to a topic are discussed. A two-part framework is proposed in which semi-automatic labeling is applied to a small subset of posts, referred to as the "seed" in further text. This seed is used as bases for semi-supervised labeling of the rest of the data. The hypothesis is that the proposed method will achieve better labeling performance than existing classification models when applied to small amounts of labeled data. The seed is labeled using posts of users with a known and consistent view on the topic. A semi-supervised multi-class prediction model labels the remaining data iteratively. In each iteration, it adds context-label pairs to the training set if softmax-based label probabilities are above the threshold. The proposed method is characterized on four datasets by comparison to the three popular text modeling algorithms (n-grams + tfidf, fastText, VDCNN) for different sizes of labeled seeds (5,000 and 50,000 posts) and for several label-prediction significance thresholds. Our proposed semi-supervised method outperformed alternative algorithms by capturing additional contexts from the unlabeled data. The accuracy of the algorithm was increasing by (3-10%) when using a larger fraction of data as the seed. For the smaller seed, lower label probability threshold was clearly a better choice, while for larger seeds no predominant threshold was observed. The proposed framework, using fastText library for efficient text classification and representation learning, achieved the best results for a smaller seed, while VDCNN wrapped in the proposed framework achieved the best results for the bigger seed. The performance was negatively influenced by the number of classes. Finally, the model was applied to characterize a biased dataset of opinions related to gun control/rights advocacy. The proposed semi-automatic seed labeling is used to label 8,448 twitter posts of 171 advocates for guns control/rights. On this application, our approach performed better than existing models and it achieves 96.5% accuracy and 0.68 F1 score. Marija Stanojevic, Jumanah Alshehri, Zoran Obradovic |
ASONAM | 2 |