EDBT 2026 Demo / reviewers in the wild / expert
Guy-Vincent Jourdan
dblp:99/2759
· DBLP profile ↗
58ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0001-6067-6545ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 15 · 2 first-author · 3 since 2021Security and privacy · 14 · 8 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Theory of computation · 7 · 4 first-author · 2 since 2021Computer networks · 6 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domains of deception: phishing through the lens of ownershipabstractPhishing attacks exploit a variety of hosting and domain ownership models, making them a dynamic and challenging threat to address. Many existing proactive detection methods focus on identifying newly registered domains, often relying on data sources like Certificate Transparency (CT) logs and Domain Name System (DNS) records. Although effective, these approaches capture only part of the attacks, leaving unanswered questions about the proportion of threats they can detect. To answer this question, we conducted a year-long study, collecting phishing reports from multiple sources and using a machine-learning classifier to analyze phishing domains. As a result of this study, we introduce two new datasets—PhishXtract and PhishXtract-Class—designed to support and advance future research in this area. By analyzing features such as the Internet presence and historical activity, we classified these domains into three types: attacker-owned, compromised, and third-party hosted. This helped us build a clearer framework for understanding the phishing domain ownership patterns observed during the year-long study. Our findings show that most phishing attacks are not hosted on attacker-owned domains. Instead, attackers often abuse legitimate infrastructures and third-party platforms to create their campaigns. Furthermore, by analyzing the reporting feeds, we were able to assess their unique and overlapping contributions to phishing detection and identify the platforms most frequently exploited by attackers. Overall, our findings provide an understanding of current phishing strategies, highlighting the reliance on legitimate infrastructures over attacker-owned domains and emphasizing the need for targeted detection strategies tailored to different hosting models exploited by attackers. Mina Erfan, Paula Branco, Guy-Vincent Jourdan |
Comput. Secur. | 3 |
| 2025 | VulPatrol: Interprocedural Vulnerability Detection and Localization through Semantic Graph LearningabstractThe growing complexity of software systems and the management of large, rapidly evolving codebases necessitate the analysis of immense volumes of lines per day due to code modifications and refactoring. Despite the use of static and dynamic analysis, test coverage, and rigorous code reviews, traditional methods often fail to accurately detect all security vulnerabilities, resulting in significant risks in production software. Recently, deep learning models have shown promising possibilities for improving vulnerability detection. Yet, there remains a clear gap between the abilities of current deep learning approaches and the level of performance required for precise source code vulnerability detection. To bridge this gap, it is crucial to develop enhancements in two fundamental areas: a code representation that accurately captures the semantics of programs and a model architecture with adequate expressiveness to analyze this representation effectively. We introduce VulPatrol, a semantic-aware, deep neural network-based system that constructs LLVM-IR interprocedural code property graphs from C/C++ source code. VulPatrol employs message-passing neural networks to capture complex dependencies and dynamic interactions within the code. As a result, it enhances the model's ability to classify potential vulnerabilities. Furthermore, we generate the first Vulnerability database based on compilable C/C++ open-source software to LLVM-IR, along with an obfuscated version. Our extensive evaluation on different benchmark datasets, including real-world programs, shows that VulPatrol outperforms the state-of-the-art baselines, improving the F1 measure by up to 12% for identifying vulnerable functions. Additionally, we evaluate VulPatrol on obfuscated code, which yields superior results regarding string variation and dissimilarity of the original codebase. Asmaa Hailane, Paria Shirani, Guy-Vincent Jourdan |
CODASPY | 3 |
| 2025 | Lost in the Noise: Evading and Detecting Backdoors in Conditional Diffusion Models
William Aiken, Paula Branco, Guy-Vincent Jourdan |
IEA/AIE (1) | 3 |
| 2025 | DistResampleR-Lite: Light Distributed Resampler for Imbalanced Regression Problems
Alexander De Furia, Anushree Kulai, Paula Branco, Guy-Vincent Jourdan |
IEA/AIE (1) | 4 |
| 2025 | Comparing Macro and Micro Approaches for Detecting Phishing Where It SpreadsabstractPhishing attacks have evolved, increasingly leveraging legitimate platforms to bypass traditional domain and URL-based detection methods. Existing proactive approaches such as monitoring Domain Name System (DNS) anomalies and Certificate Transparency (CT) logs have become less effective, as attackers now exploit trusted websites and services. To address this challenge, we present TelePhish, a dataset of annotated Telegram messages collected from public groups in high-risk categories such as Cryptocurrency, Games, and Darknet. Unlike prior works, TelePhish emphasizes the propagational behavior of phishing by capturing Telegram group interactions, message patterns, and user engagement, rather than relying on easily manipulated textual or URL features. Using this dataset, we investigate how modeling granularity impacts phishing detection performance by adopting a macro-to-micro approach. We start from a general model that captures broad phishing patterns across Telegram, then progressively narrow to more specialized, category-aware models. We evaluate four modeling strategies: (i) a general model trained on all data, (ii) general-cluster models that segment the dataset via clustering and then train models, (iii) category models trained on each category, and (iv) categorycluster models that discover sub-patterns within each category. We show that localized models can outperform general models in capturing category-specific patterns. These findings provide guidelines for designing detection models that can be tailored to either capture general phishing patterns across the entire platform or address the unique characteristics of individual categories. Mina Erfan, Paula Branco, Guy-Vincent Jourdan |
PST | 3 |
| 2025 | SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability AnalysisabstractAs Large Language Models (LLMs) evolve in understanding and generating code, accurately evaluating their reliability in analyzing source code vulnerabilities becomes in-creasingly vital. While studies have examined LLM capabilities in tasks like vulnerability detection and repair, they often over-look the importance of both structure and semantic reasoning crucial for trustworthy vulnerability analysis. To address this gap, we introduce SV-TRUSTEVAL-C, a benchmark designed to evaluate LLMs' abilities for vulnerability analysis of code written in the C programming language through two key di-mensions: structure reasoning-assessing how models identify relationships between code elements under varying data and control flow complexities; and semantic reasoning-examining their logical consistency in scenarios where code is structurally and semantically perturbed. Our results show that current LLMs are far from satisfactory in understanding complex code relationships and that their vulnerability analyses rely more on pattern matching than on robust logical reasoning. These findings underscore the effectiveness of the SV-TRUSTEVAL-C benchmark and highlight critical areas for enhancing the reasoning capabilities and trustworthiness of LLMs in real-world vulnerability analysis tasks. Our ini-tial benchmark dataset is available at https://huggingface.co/datasets/LLMs4CodeSecurity/SV-TrustEval-C-1.0 Paula Branco, Alexander M. Hoole, Manish Marwah, Hari M. Koduvely, Guy-Vincent Jourdan, Stephan Jou |
SP | 6 |
| 2025 | Advanced IDS: a comparative study of datasets and machine learning algorithms for network flow-based intrusion detection systemsabstractGlobally, cyberattacks are growing and mutating each month. Intelligent Intrusion Network Detection Systems are developed to analyze and detect anomalous traffic to face these threats. A way to address this is by using network flows, an aggregated version of communications between devices. Network Flow datasets are used to train Artificial Intelligence (AI) models to classify specific attacks. Training these models requires threat samples usually generated synthetically in labs as capturing them on operational network is a challenging task. As threats are fast-evolving, new network flows are continuously developed and shared. However, using old datasets is still a popular procedure when testing models, hindering a more comprehensive characterization of the advantages and opportunities of recent solutions on new attacks. Moreover, a standardized benchmark is missing rendering a poor comparison between the models produced by algorithms. To address these gaps, we present a benchmark with fourteen recent and preprocessed datasets and study seven categories of algorithms for Network Intrusion Detection based on Network Flows. We provide a centralized source of pre-processed datasets to researchers for easy download. All dataset are also provided with a train, validation and test split to allow a straightforward and fair comparison between existing and new solutions. We selected open state-of-the-art publicly available algorithms, representatives of diverse approaches. We carried out an experimental comparison using the Macro F1 score of these algorithms. Our results highlight each model operation on dataset scenarios and provide guidance on competitive solutions. Finally, we discuss the main characteristics of the models and benchmarks, focusing on practical implications and recommendations for practitioners and researchers. Jose Carlos Mondragon, Paula Branco, Guy-Vincent Jourdan, Andres Eduardo Gutierrez-Rodriguez, Rajesh Roshan Biswal |
Appl. Intell. | 3 |
| 2025 | X-HEART: eXplainable heterogeneous log anomaly detection using robust transformers
Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Herna L. Viktor |
Knowl. Inf. Syst. | 3 |
| 2024 | Improving GNN-Based Methods for Scam Detection in Bitcoin Transactions - A Practical Case StudyabstractThe Bitcoin generator scam is one example of existing deceptive schemes enticing users with promises of free or effortless Bitcoin generation. These scams predominantly exploit individuals unfamiliar with cryptocurrency seeking low-effort avenues to obtain Bitcoin without financial investment. In this paper, we propose and analyze methods to improve the performance of Graph Neural Networks (GNNs) in detecting fraudulent cases within Bitcoin transactional data. We explore multiple GNN variants, alongside various graph sampling methodologies. To overcome the shortcomings of these sampling methods, we propose a new sampling method BFRON—a hybrid approach mixing Breadth-First Search and Frontier Sampling. Additionally, we introduce an enhanced optimization pipeline and a new metric to improve fraudulent node detection. Evaluation metrics, including Instance Information Gain and Group Distance Ratio, are employed to analyze the challenges of over-smoothing in Graph Neural Networks and the efficacy of diverse graph sampling techniques. Our results show that overall BFRON is the best solution and RGGCN is the best-performing GNN. Moreover, we show that our enhanced pipeline and the usage of graph normalization have important advantages. Karanjot Singh Saggu, Paula Branco, Guy-Vincent Jourdan |
DSAA | 3 |
| 2024 | DevilDiffusion: Embedding Hidden Noise Backdoors into Diffusion ModelsabstractDiffusion models represent state-of-the-art deep learning architectures behind many popular and powerful image-synthesizing generative Artificial Intelligence (AI) systems. Their underlying approach, which relies on scheduled noise addition and optimized noise removal, has been applicable to the syn-thesis of sophisticated and nearly photorealistic images across various domains; however, the potential impacts of compro-mised diffusion models remain an underexplored area in the literature. While prior studies have investigated the insertion of explicit triggers into the noise or prompt spaces of diffusion models, it is unrealistic to assume that benign users would intentionally input such triggers into their own diffusion process. In our DevilDiffusion approach, we demonstrate the capability to surreptitiously embed triggers that leverage uncommon but naturally -occurring characteristics of Gaussian noise directly into the noise space of conditional diffusion models. When these specific characteristics occur in naturally -occurring noise, the model will instead construct our target backdoor image. By adjusting the trigger size and the ratio of poisoned images, we can control the trigger rates of the specified target image, ranging from less than 0.01 % to 25 % of all generated images, while still maintaining performance on the target task. William Aiken, Paula Branco, Guy-Vincent Jourdan |
PST | 3 |
| 2023 | Going Haywire: False Friends in Federated Learning and How to Find ThemabstractFederated Learning (FL) promises to offer a major paradigm shift in the way deep learning models are trained at scale, yet malicious clients can surreptitiously embed backdoors into models via trivial augmentation on their own subset of the data. This is especially true in small- and medium-scale FL systems, which consist of dozens, rather than millions, of clients. In this work, we investigate a novel attack scenario for an FL architecture consisting of multiple non-i.i.d. silos of data in which each distribution has a unique backdoor attacker and where the model convergences of adversaries are not more similar than those of benign clients. We propose a new method, dubbed Haywire, as a security-in-depth approach to respond to this novel attack scenario. Our defense utilizes a combination of kPCA dimensionality reduction of fully-connected layers in the network, KMeans anomaly detection to drop anomalous clients, and server aggregation robust to outliers via the Geometric Median. Our solution prevents the contamination of the global model despite having no access to the backdoor triggers. We evaluate the performance of Haywire from model-accuracy, defense-performance, and attack-success perspectives against multiple baselines. Through an extensive set of experiments, we find that Haywire produces the best performances at preventing backdoor attacks while simultaneously not unfairly penalizing benign clients. We carried out additional in-depth experiments across multiple runs that demonstrate the reliability of Haywire. William Aiken, Paula Branco, Guy-Vincent Jourdan |
AsiaCCS | 3 |
| 2023 | Unsupervised Graph Neural Networks for Source Code Similarity Detection
Julien Cassagne, Ettore Merlo, Paula Branco, Guy-Vincent Jourdan, Iosif-Viorel Onut |
DS | 4 |
| 2023 | HEART: Heterogeneous Log Anomaly Detection Using Robust Transformers
Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Herna L. Viktor |
DS | 3 |
| 2023 | Are GNNs the Right Tool to Mine the Blockchain? The Case of the Bitcoin Generator ScamabstractA Bitcoin Generator Scam (BGS) is a type of cyberattack in which scammers promise to provide individuals with free cryptocurrencies if they pay a mining fee. Although graph neural networks (GNNs) have been used for detecting other cryptocurrency frauds, the usefulness of these methods for BGS detection has not been studied. In this paper, we carry out extensive experiments to assess the use of both standard machine learning (ML) methods and GNNs to detect Bitcoin transactions associated with activities stemming from Bitcoin Generator Scams. We observe that the over-smoothing problem exists in GNNs designed for BGS detection and show that Random Walk Positional Encoding (RWPE) allows representing long-range interactions between far-away transactions in GNNs without causing over-smoothing. We show that the General, Powerful, Scalable (GPS) Graph Transformer with RWPE outperforms both GNN and ML based state-of-the-art fraud detection methods in Bitcoin Generator Scams. We also analyze the effectiveness of Breadth First Search (BFS) for graph sampling and show that it should not be used as it induces bias toward the subnetwork structure. We propose the Random First Search (RFS) sampling alternative and show that this is a more suitable solution. Zhikun Yuen, Paula Branco, Aaron Chew, Guy-Vincent Jourdan, Fabian Lim, Laura Wynter |
DSAA | 4 |
| 2023 | Measuring Improvement of F1-Scores in Detection of Self-Admitted Technical DebtabstractArtificial Intelligence and Machine Learning have witnessed rapid, significant improvements in Natural Language Processing (NLP) tasks. Utilizing Deep Learning, researchers have taken advantage of repository comments in Software Engineering to produce accurate methods for detecting Self-Admitted Technical Debt (SATD) from 20 open-source Java projects’ code. In this work, we improve SATD detection with a novel approach that leverages the Bidirectional Encoder Representations from Transformers (BERT) architecture. For comparison, we re-evaluated previous deep learning methods and applied stratified 10-fold cross-validation to report reliable F1-scores. We examine our model in both cross-project and intra-project contexts. For each context, we use re-sampling and duplication as augmentation strategies to account for data imbalance. We find that our trained BERT model improves over the best performance of all previous methods in 19 of the 20 projects in cross-project scenarios. However, the data augmentation techniques were not sufficient to overcome the lack of data present in the intra-project scenarios, and existing methods still perform better. Future research will look into ways to diversify SATD datasets in order to maximize the latent power in large BERT models. William Aiken, Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Mehrdad Sabetzadeh, Herna L. Viktor |
TechDebt@ICSE | 4 |
| 2022 | Phishing Kits Source Code Similarity Distribution: A Case StudyabstractAttackers (“phishers”) typically deploy source code in some host website to impersonate a brand or in general a situation in which a user is expected to provide some personal information of interest to phishers (e.g. credentials, credit card number). Phishing kits are ready-to-deploy sets of files that can be simply copied on a web server and used almost as they are. In this paper, we consider the static similarity analysis of the source code of 20871 phishing kits totalling over 182 million lines of PHP, Javascript and HTML code, that have been collected during phishing attacks and recovered by forensics teams. Reported experimental results show that as much as 90% of the analyzed kits share 90% or more of their source code with at least another kit. Differences are small, less than about 1000 programming words – identifiers, constants, strings and so on – in 40% of cases. A plausible lineage of phishing kits is presented by connecting together kits with the highest similarity. Obtained results show a very different reconstructed lineage for phishing kits when compared to a publicly available application such as Wordpress. Observed kits similarity distribution is consistent with the assumed hypothesis that kit propagation is often based on identical or near-identical copies at low cost changes. The proposed approach may help classifying new incoming phishing kits as “near-copy” or “intellectual leaps” from known and already encountered kits. This could facilitate the identification and classification of new kits as derived from older known kits. Ettore Merlo, Mathieu Margier, Guy-Vincent Jourdan, Iosif-Viorel Onut |
SANER | 3 |
| 2022 | The "Bitcoin Generator" ScamabstractThe “Bitcoin Generator Scam” (BGS) is a cyberattack in which scammers promise to provide victims with free cryptocurrencies in exchange for a small mining fee. In this paper, we present a data-driven system to detect, track, and analyze the BGS. It works as follows: we first formulate search queries related to BGS and use search engines to find potential instances of the scam. We then use a crawler to access these pages and a classifier to differentiate actual scam instances from benign pages. Last, we automatically monitor the BGS instances to extract the cryptocurrency addresses used in the scam. A unique feature of our system is that it proactively searches for and detects the scam pages. Thus, we can find addresses that have not yet received any transactions. Our data collection project spanned 16 months, from November 2019 to February 2021. We uncovered more than 8,000 cryptocurrency addresses directly associated with the scam, hosted on over 1,000 domains. Overall, these addresses have received around 8.7 million USD, with an average of 49.24 USD per transaction. Over 70% of the active addresses that we are capturing are detected before they receive any transactions, that is, before anyone is victimized. We also present some post-processing analysis of the dataset that we have captured to aggregate attacks that can be reasonably confidently linked to the same attacker or group. Our system is one of the first academic feeds to the APWG eCrime Exchange database. It has been actively and automatically feeding the database since November 2020. Emad Badawi, Guy-Vincent Jourdan, Iosif-Viorel Onut |
Blockchain Res. Appl. | 2 |
| 2022 | COVID-19 malicious domain names classificationabstractDue to the rapid technological advances that have been made over the years, more people are changing their way of living from traditional ways of doing business to those featuring greater use of electronic resources. This transition has attracted (and continues to attract) the attention of cybercriminals, referred to in this article as "attackers", who make use of the structure of the Internet to commit cybercrimes, such as phishing, in order to trick users into revealing sensitive data, including personal information, banking and credit card details, IDs, passwords, and more important information via replicas of legitimate websites of trusted organizations. In our digital society, the COVID-19 pandemic represents an unprecedented situation. As a result, many individuals were left vulnerable to cyberattacks while attempting to gather credible information about this alarming situation. Unfortunately, by taking advantage of this situation, specific attacks associated with the pandemic dramatically increased. Regrettably, cyberattacks do not appear to be abating. For this reason, cyber-security corporations and researchers must constantly develop effective and innovative solutions to tackle this growing issue. Although several anti-phishing approaches are already in use, such as the use of blacklists, visuals, heuristics, and other protective solutions, they cannot efficiently prevent imminent phishing attacks. In this paper, we propose machine learning models that use a limited number of features to classify COVID-19-related domain names as either malicious or legitimate. Our primary results show that a small set of carefully extracted lexical features, from domain names, can allow models to yield high scores; additionally, the number of subdomain levels as a feature can have a large influence on the predictions. Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Herna L. Viktor |
Expert Syst. Appl. | 3 |
| 2021 | Proactive Detection of Phishing Kit Traffic
Guy-Vincent Jourdan, Gregor von Bochmann, Iosif-Viorel Onut |
ACNS (2) | 2 |
| 2021 | Using CGAN to Deal with Class Imbalance and Small Sample Size in Cybersecurity ProblemsabstractPredictive modelling in cybersecurity domains usually involves dealing with complex settings. The class imbalance problem is a well-know challenge typically present in the cybersecurity domain. For instance, in a real-world intrusion detection scenario, the number of attacks is expected to be a a very small percentage of the normal cases. Moreover, in these applications, the number of available examples labelled is also small due to the complexity and cost of the labelling process: teams of domain experts need to be involved in the process which becomes expensive, time consuming and prone to errors. To address these problems is critical to the success of predictive modelling in cybersecurity applications. In this paper we tackle the class imbalance and small sample size through the use of a CGAN-based up-sampling procedure. We carry out an extensive set of experiments that show the positive impact of applying this solution to address the class imbalance and small sample size problems. A large data repository is built and freely provided to the research community containing 114 binary datasets based on real-world cybersecurity problems that are generated with diversified levels of imbalance and sample size. Our experiments show a clear advantage of using the CGAN-based up-sampling method specially for situations where the sample size is small and there is a large imbalance between the problem classes. In the most critical scenarios associated with extreme rarity and very small sample size, an impressive performance boost is achieved. We also explore the behaviour of this approach when the presence of these problems is less marked and we found that, while CGAN-based up-sampling is not able to further improve the minority class performance, it also has no negative impact. Thus, it is a safe to use solution, also in these scenarios. Ehsan Nazari, Paula Branco, Guy-Vincent Jourdan |
PST | 3 |
| 2021 | Minimizing Characterizing sets
Uraz Cengiz Türker, Robert M. Hierons, Guy-Vincent Jourdan |
Sci. Comput. Program. | 3 |
| 2019 | The "Game Hack" Scam
Emad Badawi, Guy-Vincent Jourdan, Gregor von Bochmann, Iosif-Viorel Onut, Jason Flood |
ICWE | 2 |
| 2019 | Domain Classifier: Compromised Machines Versus Malicious Registrations
Sophie Le Page, Guy-Vincent Jourdan, Gregor von Bochmann, Iosif-Viorel Onut, Jason Flood |
ICWE | 2 |
| 2019 | Victim or Attacker? A Multi-dataset Domain Classification of Phishing AttacksabstractIn “phishing attacks”, phishing websites disguised as trustworthy websites attempt to steal sensitive information from end users. Remediation options differ depending on whether the phishing website is hosted on a legitimate but compromised domain, in which case the domain owner is also a victim, or whether the domain itself is maliciously registered by an attacker. We propose here a novel machine-learning domain classifier, introducing features based on the internet presence and history of a domain, using only publicly available information. Using a phish feed and malicious domain feed from the Anti-Phishing Working Group (APWG), evaluation of our domain classifier achieves 94% accuracy on future malicious domains, while maintaining 88% and 92% accuracy on malicious and compromised datasets respectively from two other sources. To increase our training set we introduce a semi-supervised technique to label part of APWG's phish feed. For the rest of the feed we use our classifier and show that 62% of the websites hosting attacks are compromised while the remaining 38% belong to the attackers. The result of this research is a tool which gives some important insights on phisher's current modus operandi. It also provides a quick mechanism, entirely based on freely available data, to assess crucial information about the server and better respond to an attack by having it taken down as quickly as possible. Sophie Le Page, Guy-Vincent Jourdan |
PST | 2 |
| 2019 | Automatic Detection and Analysis of the "Game Hack" ScamabstractThe “Game Hack” Scam (GHS) is a mostly unreported cyberattack in which attackers attempt to convince victims that they will be provided with free, unlimited “resources” or other advantages for their favorite game. The endgame of the scammers ranges from monetizing for themselves the victims time and resources by having them click through endless “surveys”, filing out “market research” forms, etc., to collecting personal information, getting the victims to subscribe to questionable services, up to installing questionable executable files on their machines. Other scams such as the “Technical Support Scam”, the “Survey Scam”, and the “Romance Scam” have been analyzed before but to the best of our knowledge, GHS has not been well studied so far and is indeed mostly unknown. In this paper, our aim is to investigate and gain more knowledge on this type of scam by following a data-driven approach; we formulate GHS-related search queries, and used multiple search engines to collect data about the websites to which GHS victims are directed when they search online for various game hacks and tricks. We analyze the collected data to provide new insight into GHS and research the extent of this scam. We show that despite its low profile, the click traffic generated by the scam is in the hundreds of millions. We also show that GHS attackers use social media, streaming sites, blogs, and even unrelated sites such as change.org or jeuxvideo.com to carry out their attacks and reach a large number of victims. Our data collection spans a year; in that time, we uncovered 65,905 different GHS URLs, mapped onto over 5,900 unique domains.We were able to link attacks to attackers and found that they routinely target a vast array of games. Furthermore, we find that GHS instances are on the rise, and so is the number of victims. Our low-end estimation is that these attacks have been clicked at least 150 million times in the last five years. Finally, in keeping with similar large-scale scam studies, we find that the current public blacklists are inadequate and suggest that our method is more effective at detecting these attacks. Emad Badawi, Guy-Vincent Jourdan, Gregor von Bochmann, Iosif-Viorel Onut |
J. Web Eng. | 2 |
| 2018 | Phishing Attacks Modifications and Evolutions
Guy-Vincent Jourdan, Gregor von Bochmann, Iosif-Viorel Onut, Jason Flood |
ESORICS (1) | 2 |
| 2018 | Using AP-TED to Detect Phishing Attack VariationsabstractIt is well known that many phishing attacks are variations of previous phishing attacks. We evaluate here the feasibility of applying Pawlik and Augsten's recent implementation of Tree Edit Distance (AP-TED) calculations as a way to compare DOMs and identify similar phishing attack instances. We also compare this tree method with an existing method that uses the distance between tag vectors to quantity similarity between phishing sites. We observe that no single distance method perfectly detects all types of phishing attack variations. We find that the tree method is more demanding for computing equipment, but it better discriminates the similarity with known attacks. We also introduce a method to reduce the volume of calculations by 99.4% when calculating pairwise edit distance on trees with respect to AP-TED calculations on all data. Sophie Le Page, Gregor von Bochmann, Jason Flood, Guy-Vincent Jourdan, Iosif-Viorel Onut |
PST | 5 |
| 2017 | Tracking Phishing Attacks Over TimeabstractThe so-called ``phishing'' attacks are one of the important threats to individuals and corporations in today's Internet. Combatting phishing is thus a top-priority, and has been the focus of much work, both on the academic and on the industry sides. In this paper, we look at this problem from a new angle. We have monitored a total of 19,066 phishing attacks over a period of ten months and found that over 90% of these attacks were actually replicas or variations of other attacks in the database. This provides several opportunities and insights for the fight against phishing: first, quickly and efficiently detecting replicas is a very effective prevention tool. We detail one such tool in this paper. Second, the widely held belief that phishing attacks are dealt with promptly is but an illusion. We have recorded numerous attacks that stay active throughout our observation period. This shows that the current prevention techniques are ineffective and need to be overhauled. We provide some suggestions in this direction. Third, our observation give a new perspective into the modus operandi of attackers. In particular, some of our observations suggest that a small group of attackers could be behind a large part of the current attacks. Taking down that group could potentially have a large impact on the phishing attacks observed today. Guy-Vincent Jourdan, Gregor von Bochmann, Russell Couturier, Iosif-Viorel Onut |
WWW | 2 |
| 2016 | Reconstructing Interactions with Rich Internet Applications from HTTP Traces
Sara Baghbanzadeh, Salman Hooshmand, Gregor von Bochmann, Guy-Vincent Jourdan, Seyed M. Mirtaheri, Muhammad Faheem 0002, Iosif-Viorel Onut |
IFIP Int. Conf. Digital Forensics | 4 |
| 2015 | Using Multiple Adaptive Distinguishing Sequences for Checking Sequence Generation
Canan Güniçen, Guy-Vincent Jourdan, Hüsnü Yenigün |
ICTSS | 2 |
| 2015 | Reduced checking sequences using unreliable reset
Guy-Vincent Jourdan, Hasan Ural, Hüsnü Yenigün |
Inf. Process. Lett. | 1 |
| 2014 | Indexing Rich Internet Applications Using Components-Based Crawling
Ali Moosavi, Salman Hooshmand, Sara Baghbanzadeh, Guy-Vincent Jourdan |
ICWE | 4 |
| 2014 | PDist-RIA Crawler: A Peer-to-Peer Distributed Crawler for Rich Internet Applications
Seyed M. Mirtaheri, Gregor von Bochmann, Guy-Vincent Jourdan, Iosif-Viorel Onut |
WISE (2) | 3 |
| 2014 | Model-Based Rich Internet Applications Crawling: "Menu" and "Probability" Models
Suryakant Choudhary, Mustafa Emre Dincturk, Seyed M. Mirtaheri, Gregor von Bochmann, Guy-Vincent Jourdan, Iosif-Viorel Onut |
J. Web Eng. | 5 |
| 2014 | A Model-Based Approach for Crawling Rich Internet ApplicationsabstractNew Web technologies, like AJAX, result in more responsive and interactive Web applications, sometimes called Rich Internet Applications (RIAs). Crawling techniques developed for traditional Web applications are not sufficient for crawling RIAs. The inability to crawl RIAs is a problem that needs to be addressed for at least making RIAs searchable and testable. We present a new methodology, called “model-based crawling”, that can be used as a basis to design efficient crawling strategies for RIAs. We illustrate model-based crawling with a sample strategy, called the “hypercube strategy”. The performances of our model-based crawling strategies are compared against existing standard crawling strategies, including breadth-first, depth-first, and a greedy strategy. Experimental results show that our model-based crawling approach is significantly more efficient than these standard strategies. Mustafa Emre Dincturk, Guy-Vincent Jourdan, Gregor von Bochmann, Iosif-Viorel Onut |
ACM Trans. Web | 2 |
| 2014 | Bluetooth scatternet formation from a time-efficiency perspective
Ahmed Jeddah, Arnaud Casteigts, Guy-Vincent Jourdan, Hussein T. Mouftah |
Wirel. Networks | 3 |
| 2013 | Building Rich Internet Applications Models: Example of a Better Strategy
Suryakant Choudhary, Mustafa Emre Dincturk, Seyed M. Mirtaheri, Guy-Vincent Jourdan, Gregor von Bochmann, Iosif-Viorel Onut |
ICWE | 4 |
| 2013 | BSF-UED: A new time-efficient Bluetooth Scatternet Formation algorithm based on Unnecessary-Edges DeletionabstractWe introduce a new time-efficient Bluetooth Scatternet Formation (BSF) algorithm, called BSF-UED (Unnecessary-Edges Deletion). BSF-UED forms connected scatternets deterministically. Heuristics are added to make these scatternets outdegree limited (that is, with no more than 7 slaves per piconet). The performance of the algorithm is evaluated through a range of simulation experiments. BSF-UED is compared against some of the most common BSF algorithms which are BlueStars, BlueMIS I, BlueMIS II, and BlueMesh. We show that BSF-UED provides a good balance between the usual scatternets performance metrics, while being time efficient (nearly 1/3 of the execution time of BlueMesh). BlueStars remains a faster algorithm, but with the major flaw of generating scatternets whose piconets have a large number of slaves. Ahmed Jeddah, Arnaud Casteigts, Guy-Vincent Jourdan, Hussein T. Mouftah |
ISCC | 3 |
| 2012 | Evaluating Reliability-Testing Usage ModelsabstractTesting the reliability of an application usually requires a good usage model that accurately captures the likely sequences of inputs that the application will receive from the environment. Markov usage models and their variations have been found to be well suited for generating test cases that are statistically close to what the application is expected to receive when in production. In this article, we study the specific case of web applications. We present an evaluation method for estimating the accuracy of various reliability-testing usage models. The method is based on comparison between observed users' traces and traces inferred from the usage model. Our method gauges the accuracy of the reliability-testing usage model by calculating the sum of goodness-of-fit values of each traces and scaling the result between 0 and 1. Gregor von Bochmann, Guy-Vincent Jourdan |
COMPSAC | 3 |
| 2012 | Solving Some Modeling Challenges when Testing Rich Internet Applications for SecurityabstractCrawling is a necessary step for testing web applications for security. An important concept that impacts the efficiency of crawling is state equivalence. This paper proposes two techniques to improve any state equivalence mechanism. The first technique detects parts of the pages that are unimportant for crawling. The second technique helps identifying session parameters. We also present a summary of our research on crawling techniques for the new generation of web applications, so-called Rich Internet Applications (RIAs). RIAs present new security and crawling challenges that cannot be addressed by traditional techniques. Solving these issues is a must if we want to continue benefitting from automated tools for testing web applications. Suryakant Choudhary, Mustafa Emre Dincturk, Gregor von Bochmann, Guy-Vincent Jourdan, Iosif-Viorel Onut, Paul Ionescu |
ICST | 4 |
| 2012 | A Statistical Approach for Efficient Crawling of Rich Internet Applications
Mustafa Emre Dincturk, Suryakant Choudhary, Gregor von Bochmann, Guy-Vincent Jourdan, Iosif-Viorel Onut |
ICWE | 4 |
| 2012 | Time-efficient algorithms for the outdegree limited bluetooth scatternet formation problemabstractWe present in this paper three Bluetooth Scatternet Formation (BSF) algorithms which forms deterministically connected and outdegree limited scatternets. Our algorithms improve the execution time of algorithm BlueMIS, which consists of two consecutive phases called BlueMIS I and BlueMIS II. The first phase, BlueMIS I, is a BSF algorithm that forms connected and outdegree limited scatternets in a time-efficient manner. Our algorithms improve BlueMIS by improving BlueMIS I. First, we introduce a time-efficient implementation of communication rounds, called OrderedExchange, that is more suitable for Bluetooth networks, where in a communication round each node sends and receives a message to and from all its neighbors. Using OrderedExchange, we introduce algorithms ComputeMIS I and ComputeMIS II which form the same scatternet as those formed by BlueMIS I but with less execution time. We also introduce algorithm Eliminate. Instead of letting each node u forms a maximal independent set of all its neighbors and then considers them as slaves, as in the case of BlueMIS I, Eliminate let each node u forms a maximal independent set of all its neighbors that have smaller identifiers. Eliminate forms scatternets that are similiar to those formed by BlueMIS I, and more efficient with respect to some performance metrics. Eliminate improves the execution time of ComputeMIS I and ComputeMIS II by about 60%. Ahmed Jeddah, Guy-Vincent Jourdan, Hussein T. Mouftah |
ISCC | 2 |
| 2011 | A Strategy for Efficient Crawling of Rich Internet Applications
Kamara Benjamin, Gregor von Bochmann, Mustafa Emre Dincturk, Guy-Vincent Jourdan, Iosif-Viorel Onut |
ICWE | 4 |
| 2011 | Enabling dynamic linkage of linguistic census data at Statistics Canada (extended abstract)abstractResearch in population health consists in studying the impact of various factors (determinants) on health, with the longterm objective of yielding better policies, programs, and services. Researchers of Official Language Minority Communities (OLMCs) focus specifically on determinants related to speaking a minority language, such as English in Quebec, or French in the rest of Canada. Investigations of this type require the possibility of associating health data to linguistic information. Unfortunately, the largest health databases in Ontario, held at the Institute for Clinical Evaluative Sciences (ICES), do not contain usable linguistic variables to date. High-quality language variables however exist at Statistics Canada (2006 Census), and we are interested in enabling its linkage to ICES health data in a dynamic way. The linkage we consider is intrinsically transient and aggregated: it consists in allowing ICES to learn interactively how many Francophones are present in a given sample of individuals (sum queries). We suggest two possible privacy-preserving mechanisms to enable dynamic sum queries: 1) by constraining the dataflow itself; 2) by adapting recent results ([1]) to characterize what leakage is at play in our scenario and what parameters impact the tradeoff between leakage and utility. We rely on these results to argue that a safe exposition of linguistic data could indeed be envisioned, and beyond, that similar techniques could be used to enrich provincial health databases in general with a range of federal census data, making it possible to perform fine-grained community-based studies in Canada. Arnaud Casteigts, Marie-Hélène Chomienne, Louise Bouchard, Guy-Vincent Jourdan |
ISI | 4 |
| 2011 | Improved Usage Model for Web Application Reliability Testing
Gregor von Bochmann, Guy-Vincent Jourdan |
ICTSS | 2 |
| 2010 | Some side effects of FHSS on Bluetooth networks distributed algorithmsabstractWe present, in this paper, two case studies showing the substantial impact of some features in the Bluetooth specifications on Bluetooth networks distributed algorithms. These features were often neglected in previous research studies, and are related to the Frequency Hopping Spread Spectrum (FHSS) technique used in Bluetooth. The first case study shows that some small changes in the specifications of Bluetooth can have a substantial positive impact static Bluetooth Scatternet Formation (BSF) algorithms but not the dynamic BSF algorithms. The second case shows that a small change in the implementation of the Bluetooth specifications can have a substantial negative impact on static BSF algorithms. This study calls for more research on the side-effects of using the FHSS technique in Bluetooth and calls for more publishing of similar results to help to understand more distributed algorithms running over Bluetooth networks. Ahmed Jeddah, Guy-Vincent Jourdan, Nejib Zaguia |
AICCSA | 2 |
| 2010 | Lightweight protection against brute force login attacks on Web applicationsabstractPassword-based systems and, more generally, authentication systems based on something you know, are commonplace on the Internet. Web applications using these systems can be the target of brute force login attacks, in which an attacker tries to compromise a given account or any user account on the system. These applications rarely implement effective protection mechanisms against these attacks. In this paper, we review the situation and propose a practical, simple, security mechanism. Our system is non-intrusive and can be incorporated into most web applications with very little modification to the application code. Carlisle M. Adams, Guy-Vincent Jourdan, Jean-Pierre Levac, François Prevost |
PST | 2 |
| 2010 | Lower bounds on lengths of checking sequencesabstractAbstract Lower bounds on the lengths of checking sequences constructed for testing from Finite State Machine-based specifications are established. These bounds consider the case where a distinguishing sequence is used in forming state recognition and transition verification subsequences and identify the effects of overlapping among such subsequences. Empirical results show that the existing methods for construction of checking sequences provide checking sequences with lengths that are within acceptable distance to these lower bounds. Guy-Vincent Jourdan, Hasan Ural, Hüsnü Yenigün, Ji Chao Zhang |
Formal Aspects Comput. | 1 |
| 2009 | A Note on the Study of Bluetooth Networks' Distributed AlgorithmsabstractBluetooth has provided many features that enable wireless ad-hoc networks, but it has also introduced many problems. The root cause of these problems lies in its communication mechanisms. We argue in this paper that the models that have been used to study distributed algorithms on Bluetooth networks do not adequately model these networks in most cases, and were often oversimplified. This is mainly due to how the many restrictions that the Bluetooth specifications impose on such networks are taken into account, and the lack of ldquoshared knowledgerdquo of these restrictions among the researchers of this field. We give some examples to back our argument. We give also some suggestions and proposals to overcome these issues. Ahmed Jeddah, Nejib Zaguia, Guy-Vincent Jourdan |
MASS | 3 |
| 2009 | Checking Sequence Construction Using Adaptive and Preset Distinguishing SequencesabstractMethods for testing from finite state machine-based specifications often require the existence of a preset distinguishing sequence for constructing checking sequences. It has been shown that an adaptive distinguishing sequence is sufficient for these methods. This result is significant because adaptive distinguishing sequences are strictly more common and up to exponentially shorter than preset ones. However, there has been no study on the actual effect of using adaptive distinguishing sequences on the length of checking sequences. This paper describes experiments that show that checking sequences constructed using adaptive distinguishing sequences are almost consistently shorter than those based on preset distinguishing sequences. This is investigated for three different checking sequence generation methods and the results obtained from an extensive experimental study are given. Robert M. Hierons, Guy-Vincent Jourdan, Hasan Ural, Hüsnü Yenigün |
SEFM | 2 |
| 2009 | On Testing 1-Safe Petri NetsabstractFormal models are often considered for software systems specification, and are helpful for verifying that certain properties are respected, or for automatically generating the implementation code corresponding to the model, or again for conformance testing, for the automatic generation of test cases to check an implementation against the formal specification. Variations of finite state machine (FSM) models have been mostly used for conformance testing, while the otherwise very popular formal model of Petri nets is seldom mentioned in this context. In this paper, we look at the question of conformance testing when the model is provided in the form of a 1-safe Petri net. We provide a general framework for conformance testing, and give algorithms for deriving test cases under different assumptions: besides the adaptation of methods originally developed for FSMs which lead to exponentially long test sequences, we have identified cases for which polynomial testing algorithms for free-choice Petri nets can be provided. These results are significant when modeling concurrent systems, as exemplified by workflow modeling. Guy-Vincent Jourdan, Gregor von Bochmann |
TASE | 1 |
| 2007 | Recovering Repetitive Sub-functions from Observations
Guy-Vincent Jourdan, Hasan Ural, Hüsnü Yenigün |
FORTE | 1 |
| 2006 | Minimizing Coordination Channels in Distributed Testing
Guy-Vincent Jourdan, Hasan Ural, Hüsnü Yenigün |
FORTE | 1 |
| 2005 | Minimizing the number of inputs while applying adaptive test cases
Guy-Vincent Jourdan, Hasan Ural, Nejib Zaguia |
Inf. Process. Lett. | 1 |
| 2003 | Degree Navigator™: The Journey of a Visualization Software
Guy-Vincent Jourdan, Ivan Rival, Nejib Zaguia |
GD | 1 |
| 1995 | On-The-Fly Analysis of Distributed Computations
Eddy Fromentin, Claude Jard, Guy-Vincent Jourdan, Michel Raynal |
Inf. Process. Lett. | 3 |
| 1994 | A General Approach to Trace-Checking in Distributed Computing SystemsabstractThe problem of checking the correctness of distributed computations arises when debugging distributed algorithms, and more generally when testing protocols or distributed applications. For that purpose, one describes the expected behavior (or suspected errors) by a global property: for example, a predicate on process variables, or the set of admissible orderings on observable events. The problem is to check whether this property is satisfied or not during the execution. A relevant model for this study is the partial order of message causality and the associated state graph, called "lattice of consistent cuts". In this paper, we propose a general approach to trace checking, based on partial order theory.> Claude Jard, Thierry Jéron, Guy-Vincent Jourdan, Jean-Xavier Rampon |
ICDCS | 3 |
| 1994 | On the Coding of Dependencies in Distributed Computations (Abstract)abstractNo abstract available. Claude Jard, Guy-Vincent Jourdan |
PODC | 2 |