VLDB 2026 Research / reviewers in the wild / expert
Fatima Zahra Qachfar
dblp:318/3920
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-6254-3364ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Real-Time, Evidence-Based Alerts for Protection From Phishing AttacksabstractDespite two decades of research on automatic filtering systems, phishing attacks remain a serious problem. To alleviate risks from filtering failures, we design and evaluate the effectiveness of a new warning system on users’ susceptibility to phishing. Our proposed technique highlights key sentences based on an analysis of the persuasive techniques used. An online mixed-design study ($n=604$) shows that adding our highlighting technique outperforms existing warning solutions. It also identifies the relative efficacy of different appeals and the characteristics of susceptible users. Results show that adding our highlighting techniqueis useful even with false positives and false negatives. Inspired by this result, we propose an automatic warning generator. We created a small labeled dataset of suspicious sentences and used data augmentation. Our best models achieve F1 score of 99.95% in detecting phishing emails and 88% in detecting suspicious sentences. Shahryar Baki, Fatima Zahra Qachfar, Rakesh M. Verma, Ryan Kennedy, Daniel Jones 0003 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | Domain-Agnostic Adapter Architecture for Deception Detection: Extensive Evaluations with the DIFrauD BenchmarkabstractDespite significant strides in training expansive transformer models, their deployment for niche tasks remains intricate. This paper delves into deception detection, assessing domain adaptation methodologies from a cross-domain lens using transformer Large Language Models (LLMs). We roll out a new corpus with roughly 100,000 honest and misleading statements in seven domains, designed to serve as a benchmark for multidomain deception detection. As a primary contribution, we present a novel parameter-efficient finetuning adapter, PreXIA, which was proposed and implemented as part of this work. The design is model-, domain- and task-agnostic, with broad applications that are not limited by the confines of deception or classification tasks. We comprehensively analyze and rigorously evaluate LLM tuning methods and our original design using the new benchmark, highlighting their strengths, pointing out weaknesses, and suggesting potential areas for improvement. The proposed adapter consistently outperforms all competition on the DIFrauD benchmark used in this study. To the best of our knowledge, it improves on the state-of-the-art in its class for the deception task. In addition, the evaluation process leads to unexpected findings that, at the very least, cast doubt on the conclusions made in some of the recently published research regarding reasoning ability’s unequivocal dominance over representations quality with respect to the relative contribution of each one to a model’s performance and predictions. Dainis Boumber, Fatima Zahra Qachfar, Rakesh M. Verma |
LREC/COLING | 2 |
| 2024 | All Your LLMs Belong to Us: Experiments with a New Extortion Phishing Dataset
Fatima Zahra Qachfar, Rakesh M. Verma |
DBSec | 1 |
| 2024 | Blue Sky: Multilingual, Multimodal Domain Independent Deception DetectionabstractDeception, a pervasive aspect of communication, has undergone a significant transformation in the digital age. With the globalization of online interactions, individuals are communicating in multiple languages, mixing languages on social media. A variety of data is now available in many languages, while the techniques for detecting deception are similar across the board. Recent studies have shown the possibility of the existence of universal linguistic cues to deception across domains within the English language; however, the existence of such cues in other languages remains unknown. Furthermore, the practical task of deception detection in low-resource languages is not a well-studied problem due to the lack of labeled data. Another dimension of deception is multimodality. For example, in fake news or disinformation, there may be a picture with an altered caption. This paper calls for a comprehensive investigation into the complexities of deceptive language across linguistic boundaries and modalities, and raises the possibility of use of multilingual transformer models and labeled data in a variety of languages to universally address the task of deception detection. Dainis Boumber, Rakesh M. Verma, Fatima Zahra Qachfar |
SDM | 3 |
| 2022 | Leveraging Synthetic Data and PU Learning For Phishing Email DetectionabstractImbalanced data classification has always been one of the most challenging problems in data science especially in the cybersecurity field, where we observe an out-of-balance proportion between benign and phishing examples in security datasets. Even though there are many phishing detection methods in literature, most of them neglect the imbalanced nature of phishing email datasets. In this paper, we examine the imbalanced property by varying legitimate to phishing class ratios. We generate new synthetic instances using a generative adversarial network model for long sentences (LeakGAN) to balance out the training process and ameliorate its impact on classification. These synthetic instances are labeled by positive-unlabeled learning and added to the initial imbalanced training set. The resulting dataset is given to the Bidirectional Encoder Representations from Transformers (BERT) model for sequence classification. We compare several state-of-the-art methods from the literature against our approach, which achieves a high performance throughout all the imbalanced ratios reaching an F1-score of 99.6% for the most extreme imbalanced ratio and an F1-score of 99.8% for balanced cases. Fatima Zahra Qachfar, Rakesh M. Verma, Arjun Mukherjee |
CODASPY | 1 |