EDBT 2026 Demo / reviewers in the wild / expert
Nhien-An Le-Khac
dblp:37/6424 · also Nhien An LeKhac, NhienAn LeKhac
· DBLP profile ↗
60ranked-venue papers
6as first author
27since 2021 · last 2026
0000-0003-4373-2212ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 34 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Computer networks · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Binary Classification: A Semi-supervised Approach to Generalized AI-generated Image DetectionabstractThe rapid advancement of generators (e.g., StyleGAN, Midjourney, DALL-E) has produced highly realistic synthetic images, posing significant challenges to digital media authenticity. These generators are typically based on a few core architectural families, primarily Generative Adversarial Networks (GANs) and Diffusion Models (DMs). A critical vulnerability in current forensics is the failure of detectors to achieve cross-generator generalization, especially when crossing architectural boundaries (e.g., from GANs to DMs). We hypothesize that this gap stems from fundamental differences in the artifacts produced by these distinct architectures. In this work, we provide a theoretical analysis explaining how the distinct optimization objectives of the GAN and DM architectures lead to different manifold coverage behaviors. We demonstrate that GANs permit partial coverage, often leading to boundary artifacts, while DMs enforce complete coverage, resulting in over-smoothing patterns. Motivated by this analysis, we propose the Triarchy Detect or (TriDetect), a semi-supervised approach that enhances binary classification by discovering latent architectural patterns within the "fake" class. TriDetect employs balanced cluster assignment via the Sinkhorn-Knopp algorithm and a cross-view consistency mechanism, encouraging the model to learn fundamental architectural distincts. We evaluate our approach on two standard benchmarks and three in-the-wild datasets against 13 baselines to demonstrate its generalization capability to unseen generators. Hong-Hanh Nguyen-Le, Van-Tuan Tran, Thuc D. Nguyen, Nhien-An Le-Khac |
AAAI | 4 |
| 2025 | LogaLookup: Efficient Multivariate Lookup Argument for Accelerated Proof Generation
Dien H. A. Tran, Tam N. B. Nguyen, Nhien-An Le-Khac, Thuc D. Nguyen |
AsiaCCS | 3 |
| 2025 | Stabilizing Data-Free Model ExtractionabstractModel extraction is a severe threat to Machine Learning-as-a-Service systems, especially through data-free approaches, where dishonest users can replicate the functionality of a black-box target model without access to realistic data. Despite recent advancements, existing data-free model extraction methods suffer from the oscillating accuracy of the substitute model. This oscillation, which could be attributed to the constant shift in the generated data distribution during the attack, makes the attack impractical since the optimal substitute model cannot be determined without access to the target model’s in-distribution data. Hence, we propose MetaDFME, a novel data-free model extraction method that employs meta-learning in the generator training to reduce the distribution shift, aiming to mitigate the substitute model’s accuracy oscillation. In detail, we train our generator to iteratively capture the meta-representations of the synthetic data during the attack. These meta-representations can be adapted with a few steps to produce data that facilitates the substitute model to learn from the target model while reducing the effect of distribution shifts. Our experiments on popular baseline image datasets, MNIST, SVHN, CIFAR-10, and CIFAR-100, demonstrate that MetaDFME outperforms the current state-of-the-art data-free model extraction method while exhibiting a more stable substitute model’s accuracy during the attack. Dat-Thinh Nguyen, Kim-Hung Le, Nhien-An Le-Khac |
ECAI | 3 |
| 2025 | Think Twice Before Adaptation: Improving Adaptability of DeepFake Detection via Online Test-Time AdaptationabstractDeepfake (DF) detectors face significant challenges when deployed in real-world environments, particularly when encountering test samples deviated from training data through either postprocessing manipulations or distribution shifts. We demonstrate postprocessing techniques can completely obscure generation artifacts presented in DF samples, leading to performance degradation of DF detectors. To address these challenges, we propose Think Twice before Adaptation (T2A), a novel online test-time adaptation method that enhances the adaptability of detectors during inference without requiring access to source training data or labels. Our key idea is to enable the model to explore alternative options through an Uncertainty-aware Negative Learning objective rather than solely relying on its initial predictions as commonly seen in entropy minimization (EM)-based approaches. We also introduce an Uncertain Sample Prioritization strategy and Gradients Masking technique to improve the adaptation by focusing on important samples and model parameters. Our theoretical analysis demonstrates that the proposed negative learning objective exhibits complementary behavior to EM, facilitating better adaptation capability. Empirically, our method achieves state-of-the-art results compared to existing test-time adaptation (TTA) approaches and significantly enhances the resilience and generalization of DF detectors during inference. Hong-Hanh Nguyen-Le, Van-Tuan Tran, Dinh-Thuc Nguyen, Nhien-An Le-Khac |
IJCAI | 4 |
| 2025 | SoK: Systematic analysis of adversarial threats against deep learning approaches for autonomous anomaly detection systems in SDN-IoT networks
Tharindu Lakshan Yasarathna, Nhien-An Le-Khac |
J. Inf. Secur. Appl. | 2 |
| 2025 | Variational Deep Clustering approaches for anomaly-based cyber-attack detection
Van Quan Nguyen, Long Thanh Ngo, Minh Le Nguyen 0001, Nhien-An Le-Khac |
J. Netw. Comput. Appl. | 5 |
| 2025 | Large Language Model XAI approach for illicit activity Investigation in Bitcoin
Jack Nicholls, Aditya Kuppa, Nhien-An Le-Khac |
Neural Comput. Appl. | 3 |
| 2025 | Privacy-preserving speaker verification system using Ranking-of-Element hashing
Hong-Hanh Nguyen-Le, Lam Tran, Dinh Song An Nguyen, Nhien-An Le-Khac, Thuc Nguyen |
Pattern Recognit. | 4 |
| 2024 | Driven to Evidence: The Digital Forensic Trail of Vehicles
Catherine McKay, Nhien-An Le-Khac |
ICDF2C (2) | 2 |
| 2024 | SoK: Behind the Accuracy of Complex Human Activity Recognition Using Deep LearningabstractHuman Activity Recognition (HAR) is a well-studied field with research dating back to the 1980s. Over time, HAR technologies have evolved significantly from manual feature extraction, rule-based algorithms, and simple machine learning models to powerful deep learning models, from one sensor type to a diverse array of sensing modalities. The scope has also expanded from recognising a limited set of activities to encompassing a larger variety of both simple and complex activities. However, there still exist many challenges that hinder advancement in complex activity recognition using modern deep learning methods. In this paper, we comprehensively systematise factors leading to inaccuracy in complex HAR, such as data variety and model capacity. Among many sensor types, we give more attention to wearable and camera due to their prevalence. Through this Systematisation of Knowledge (SoK) paper, readers can gain a solid understanding of the development history and existing challenges of HAR, different categorisations of activities, obstacles in deep learning-based complex HAR that impact accuracy, and potential research directions. Nhien-An Le-Khac |
IJCNN | 2 |
| 2024 | D-CAPTCHA++: A Study of Resilience of Deepfake CAPTCHA under Transferable Imperceptible Adversarial AttackabstractThe advancements in generative AI have enabled the improvement of audio synthesis models, including text-to-speech and voice conversion. This raises concerns about its potential misuse in social manipulation and political interference, as synthetic speech has become indistinguishable from natural human speech. Several speech-generation programs are utilized for malicious purposes, especially impersonating individuals through phone calls. Therefore, detecting fake audio is crucial to maintain social security and safeguard the integrity of information. Recent research has proposed a D-CAPTCHA system based on the challenge-response protocol to differentiate fake phone calls from real ones. In this work, we study the resilience of this system and introduce a more robust version, D-CAPTCHA++, to defend against fake calls. Specifically, we first expose the vulnerability of the D-CAPTCHA system under the transferable imperceptible adversarial attack. Secondly, we mitigate such vulnerability by improving the robustness of the system by using adversarial training in D-CAPTCHA deepfake detectors and task classifiers. Hong-Hanh Nguyen-Le, Van-Tuan Tran, Dinh-Thuc Nguyen, Nhien-An Le-Khac |
IJCNN | 4 |
| 2024 | Cross-Validation for Detecting Label Poisoning Attacks: A Study on Random Forest Algorithm
Tharindu Lakshan Yasarathna, Lankeshwara Munasinghe, Harsha K. Kalutarage, Nhien-An Le-Khac |
SEC | 4 |
| 2024 | Manipulating Prompts and Retrieval-Augmented Generation for LLM Service Providers
Aditya Kuppa, Jack Nicholls, Nhien-An Le-Khac |
SECRYPT | 3 |
| 2024 | Improving Security in Internet of Medical Things through Hierarchical Cyberattacks ClassificationabstractThe development of Internet of Medical Things (IoMT) devices has significantly enhanced healthcare quality but has also introduced critical cybersecurity vulnerabilities. While effective in some cases, traditional flat classification models often struggle to maintain accuracy when dealing with closely related IoMT cyber threats. This work addresses the issue by proposing a novel approach using hierarchical classification techniques for cyberattack detection and classification. Unlike traditional flat classification methods, our hierarchical approach considers the relationships between attack classes, organizing them into a structured hierarchy. We utilize the CICIoMT2024 dataset, the first specific dataset for IoMT cyberattacks, to develop and evaluate our model. Our results indicate that the hierarchical approach demonstrated high accuracy and proposed the hierarchical consistency of the data, compared to baseline flat classification approaches. The study concludes that hierarchical classification techniques offer a more nuanced method for detecting closely related attack types and hold significant potential for improving IoMT security, particularly in more complex hierarchical scenarios. These findings contribute valuable insights to the field of IoMT cybersecurity, suggesting that further research into more intricate hierarchical structures could yield even more effective security solutions. Vince Noort, Nhien-An Le-Khac, Hong-Hanh Nguyen-Le |
TrustCom | 2 |
| 2023 | FraudLens: Graph Structural Learning for Bitcoin Illicit Activity IdentificationabstractIllicit activity in cryptocurrency has increased dramatically over the years. Bitcoin mechanics allow for users to mask their identity through obfuscation techniques. Much research has been published in the domain of identifying illicit activity in cryptocurrency, and in particular the emergence of Graph Neural Networks (GNNs) has shown great promise in this area. In this paper, we propose two graph preprocessing methods to improve performance and robustness of our node classification GNN models in identifying illicit transactions in the Bitcoin network. Our methods focus on graph restructuring through measuring the connectivity of nodes in a graph, and the similarity of the underlying features each node possesses. We demonstrate the graph restructuring methodologies on five GNN architectures and empirically show an improvement of evaluation metrics when compared against the unprocessed graph dataset. We compare our proposed methods against other imbalanced node classification techniques on a common graph dataset. This methodology has great opportunity in the transaction monitoring landscape for exchanges and financial institutions attempting to capture potential illicit activity taking place on their networks including money laundering. Jack Nicholls, Aditya Kuppa, Nhien-An Le-Khac |
ACSAC | 3 |
| 2023 | SoK: The Next Phase of Identifying Illicit Activity in BitcoinabstractIdentifying illicit behavior in the Bitcoin network is a well explored topic. The methods proposed over time have generated great insights into the deanonymization of the Bitcoin user base through the clustering of inputs and outputs. With advanced techniques being deployed by Bitcoin users, these heuristics are now being challenged in their ability to aid in the detection of illicit activity. In this SoK, we provide a comprehensive list of methods deployed by malicious actors on the network and illicit transaction mining methods. We highlight the issues associated with conducting law enforcement investigations and propose recommendations for the research community to address these issues. Our recommendations include the release of public data by exchanges to allow researchers and law enforcement to further protect the network from malicious users. We recommend the enhancement of current heuristics through machine learning methods and discuss how researchers can take the fight head-on against expert cyber criminals. Jack Nicholls, Aditya Kuppa, Nhien-An Le-Khac |
ICBC | 3 |
| 2023 | Finding Forensic Artefacts in Long-Term Frequency Band Occupancy Measurements Using Statistics and Machine Learning
Bart Somers, Asanka P. Sayakkara, Darren R. Hayes, Nhien-An Le-Khac |
ICDF2C (1) | 4 |
| 2023 | Identify Users on Dating Applications: A Forensic Perspective
Paul Stenzel, Nhien-An Le-Khac |
ICDF2C (1) | 2 |
| 2023 | Forensic Analysis of the iOS Apple Pay Mobile Payment System
Trevor Nicholson, Darren R. Hayes, Nhien-An Le-Khac |
IFIP Int. Conf. Digital Forensics | 3 |
| 2022 | DACMA: Designing space ordering optimizations to scalably manage aerial imagesabstractAerial images are a special class of remote sensing images, as they are intentionally collected with a high degree of overlap. This high degree of overlap complicates existing index strategies such as R-tree and Space Filling Curve (SFC) based index techniques due to complications in space splitting, granularity of the grid cells and excessive duplication of image object identifiers (IOIs). However, SFC based space ordering can be modified to provide scalable management of overlapping aerial images. This involves overcoming similar IOIs in adjacent grid cells, which would naturally occur in SFC based grids with such data. IOI duplication can be minimized by merging adjacent grid cells through the proposed “Designing Adjacent Cell Merge Algorithm” (DACMA). This work focuses on establishing a proper adjacent cell merge metric and merge percentage value. Using a highly scalable, distributed HBase cluster for both a single aerial mapping project, and multiple aerial mapping projects, experiments evaluated Jaccard Similarity (JS) and Percentage of Overlap (PO) merge metrics. JS had significant advantages: (i) generating smaller merged regions and (ii) obtaining over 21% and 36% improvement in reducing query response times compared to PO. As a result, JS is proposed for the merge metric for DACMA. For the merge percentage two considerations were dominant: (i) substantial storage reductions with respect to both straight forward SFC-based cell space indexing and 4SA based indexing, and (ii) minimal impact on the query response time. The proposed merge percentage value was selected to optimize the storage (i.e. space) needs and response time (i.e. time) herein named the "Space-Time Trade-off Optimization Percentage" value (or STOP value) is presented. Chamin Nalinda Lokugam Hewage, Debra F. Laefer, Michela Bertolotto, Anh-Vu Vo, Nhien-An Le-Khac |
IEEE Big Data | 5 |
| 2022 | 4DHI: An index for approximate kNN search of remotely sensed images in Key-Value databasesabstractState-of-the-art, scalable, indexing techniques in location-based image data retrieval are primarily focused on supporting window and range queries. However, support of these indexes is not well explored when there are multiple spatially similar images to retrieve for a given geographic location. Adoption of existing spatial indexes such as the kD-tree pose major scalability impediments. In response, this work proposes a novel scalable, key-value, database oriented, secondary-memory based, spatial index to retrieve the top$k$most spatially similar images to a given geographic location. The proposed index introduces a 4-dimensional Hilbert index (4DHI). This space filling curve is implemented atop HBase (a key-value database). Experiments performed on both synthetically generated and real world data demonstrate comparable accuracy with MD-HBase (a state of the art, scalable, multidimensional point data management system) and better performance. Specifically, 4DHI yielded 34% - 39% storage improvements compared to the disk consumption of the original index of MD-HBase. The compactness in 4DHI also yielded up to 3.4 and 4.7 fold gains when retrieving 6400 and 12800 neighbours, respectively; compared to the adoption of original index of MD-HBase for respective neighbour searches. An optimization technique termed “Bounding Box Displacement” (BBD) is introduced to improve the accuracy of the top$k$approximations in relation to the results of in-memory kD-tree. Finally, a method of reducing row key length is also discussed for the proposed 4DHI to further improve the storage efficiency and scalability in managing large numbers of remotely sensed images. Chamin Nalinda Lokugam Hewage, Anh-Vu Vo, Nhien-An Le-Khac, Debra F. Laefer, Michela Bertolotto |
IC2E | 3 |
| 2021 | A Hybrid CNN-LSTM Based Approach for Anomaly Detection Systems in SDNsabstractSoftware-Defined Networking (SDN) is a promising technology for the future Internet. However, the SDN paradigm introduces new attack vectors that do not exist in the conventional distributed networks. This paper develops a hybrid Intrusion Detection System (IDS) by combining the Convolutional Neural Network (CNN) and Long Short-Term Memory Network (LSTM). The proposed model is capable of capturing the spatial and temporal features of the network traffic. Two regularization techniques i.e., L2 Regularization () and dropout method are used to overcome with the overfitting problem. The proposed method improves the intrusion detection performance of zero-day attacks. The InSDN dataset — the most recent dataset for SDN networks is used to test and evaluate the performance of the proposed model. The results indicate that integrating the CNN with LSTM improves the intrusion detection performance and achieves an accuracy of 96.32%. The estimated accuracy is higher than the accuracy of each individual model. In addition, it is established that the regularization techniques improves the performance of the CNN algorithms in detecting new intrusions when compared to the standard CNN. The findings of this study facilitates the development of robust IDS systems for SDN environment. Mahmoud Abdallah, Nhien-An Le-Khac, Hamed Z. Jahromi, Anca Jurcut |
ARES | 2 |
| 2021 | Linking CVE's to MITRE ATT&CK TechniquesabstractThe MITRE Corporation is a non-profit organization that has made substantial efforts into creating and maintaining knowledge bases relevant to cybersecurity and has been widely adopted by the community. ATT&CK ”Adversarial Tactics, Techniques, and Common Knowledge” is a popular taxonomy by MITRE, which describes threat actor behaviors. Techniques are the foundation of the ATT&CK model, they are the actions that adversaries perform to accomplish goals, which translate into the model’s tactics. The aim of ATT&CK is to categorize adversary behavior to help improve the post-compromise detection of advanced intrusions. Aditya Kuppa, Lamine M. Aouad, Nhien-An Le-Khac |
ARES | 3 |
| 2021 | Accessing Secure Data on Android Through Application Analysis
Richard Buurke, Nhien-An Le-Khac |
ICDF2C | 2 |
| 2021 | Structural textile pattern recognition and processing based on hypergraphsabstractAbstract The humanities, like many other areas of society, are currently undergoing major changes in the wake of digital transformation. However, in order to make collection of digitised material in this area easily accessible, we often still lack adequate search functionality. For instance, digital archives for textiles offer keyword search, which is fairly well understood, and arrange their content following a certain taxonomy, but search functionality at the level of thread structure is still missing. To facilitate the clustering and search, we introduce an approach for recognising similar weaving patterns based on their structures for textile archives. We first represent textile structures using hypergraphs and extract multisets of k-neighbourhoods describing weaving patterns from these graphs. Then, the resulting multisets are clustered using various distance measures and various clustering algorithms (K-Means for simplicity and hierarchical agglomerative algorithms for precision). We evaluate the different variants of our approach experimentally, showing that this can be implemented efficiently (meaning it has linear complexity), and demonstrate its quality to query and cluster datasets containing large textile samples. As, to the best of our knowledge, this is the first practical approach for explicitly modelling complex and irregular weaving patterns usable for retrieval, we aim at establishing a solid baseline. Vuong M. Ngo, Sven Helmer, Nhien-An Le-Khac, M. Tahar Kechadi |
Inf. Retr. J. | 3 |
| 2021 | A novel hybrid model for intrusion detection systems in SDNs based on CNN and a new regularization techniqueabstractSoftware-defined networking (SDN) is a new networking paradigm that separates the controller from the network devices i.e. routers and switches. The centralized architecture of the SDN facilitates the overall network management and addresses the requirement of current data centers. While there are high benefits offered by the SDN architecture, the risk of new attacks is a critical problem and can prevent the wide adoption of SDNs. The SDN controller is a crucial element, and it is an attractive target for the intruders. In case the attacker successfully accessed the SDN controller, it can route the traffic based on its own requirements, causing severe damage to the entire network. The network intrusion detection systems (NIDSs) are important tools to detect and secure the network environment from malicious activities and anomalous attacks. Deep Learning (DL) has recently shown desirable results in a variety of problems, such as text, speech, and image applications, etc. While several related works deployed DL for NIDSs, most of these approaches ignore the influence of the overfitting problem during the implementation of DL algorithms. As a result, it can impact the robustness of the anomaly detection system and lead to poor model performance for zero-day attacks. In this work, we propose a new hybrid DL approach based on the convolutional neural network (CNN) to classify the flow traffic into normal or attack classes. A new regularizer method, namely SD-Reg, which is based on the standard deviation of the weight matrix, has been used to address the problem of overfitting and to improve the capability of NIDSs in detection of unseen intrusion events. The evaluation results indicate that the SD-Reg outperforms the previous regularizer methods. In addition, the proposed hybrid technique gives a higher performance in all the evaluation metrics compared to the single DL models. Several datasets, including the InSDN – the most recent dataset for SDN – are used to train and evaluate the performance of all techniques. Furthermore, we suggest a lightweight NIDS by training the CNN-based models using a less number of features without causing a significant drop in the model performance. Mahmoud Said Elsayed, Nhien-An Le-Khac, Marwan Ali Albahar, Anca Jurcut |
J. Netw. Comput. Appl. | 2 |
| 2021 | Adversarial XAI Methods in CybersecurityabstractMachine Learning methods are playing a vital role in combating ever-evolving threats in the cybersecurity domain. Explanation methods that shed light on the decision process of black-box classifiers are one of the biggest drivers in the successful adoption of these models. Explaining predictions that address ‘Why?/Why Not?’ questions help users/stakeholders/analysts understand and accept the predicted outputs with confidence and build trust. Counterfactual explanations are gaining popularity as an alternative method to help users to not only understand the decisions of black-box models (why?) but also to provide a mechanism to highlight mutually exclusive data instances that would change the outcomes (why not?). Recent Explainable Artificial Intelligence literature has focused on three main areas: (a) creating and improving explainability methods that help users better understand how the internal of ML models work as well as their outputs; (b) attacks on interpreters with a white-box setting; (c) defining the relevant properties, metrics of explanations generated by models. Nevertheless, there is no thorough study of how the model explanations can introduce new attack surfaces to the underlying systems. A motivated adversary can leverage the information provided by explanations to launch membership inference, and model extraction attacks to compromise the overall privacy of the system. Similarly, explanations can also facilitate powerful evasion attacks such as poisoning and back door attacks. In this paper, we cover this gap by tackling various cybersecurity properties and threat models related to counterfactual explanations. We propose a new black-box attack that leverages Explainable Artificial Intelligence (XAI) methods to compromise the confidentiality and privacy properties of underlying classifiers. We validate our approach with datasets and models used in the cyber security domain to demonstrate that our method achieves the attacker’s goal under threat models which reflect the real-world settings. Aditya Kuppa, Nhien-An Le-Khac |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | SoK: exploring the state of the art and the future potential of artificial intelligence in digital forensic investigationabstractMulti-year digital forensic backlogs have become commonplace in law enforcement agencies throughout the globe. Digital forensic investigators are overloaded with the volume of cases requiring their expertise compounded by the volume of data to be processed. Artificial intelligence is often seen as the solution to many big data problems. This paper summarises existing artificial intelligence based tools and approaches in digital forensics. Automated evidence processing leveraging artificial intelligence based techniques shows great promise in expediting the digital forensic analysis process while increasing case processing capacities. For each application of artificial intelligence highlighted, a number of current challenges and future potential impact is discussed. Xiaoyu Du 0003, Christopher James Hargreaves, John Sheppard 0002, Felix Anda, Asanka P. Sayakkara, Nhien-An Le-Khac, Mark Scanlon |
ARES | 6 |
| 2020 | Effect of Security Controls on Patching Window: A Causal Inference based ApproachabstractIn many organisations there are up to 15 security controls that help defenders accurately identify and prioritise information security risks. Due to the lack of clarity into the effectiveness and capabilities of these defences, and poor visibility to overall risk posture has led to a crisis of prioritisation. Lately, organisations rely on scenario based red teaming exercises which test the contribution of a security control to the security preparedness of the organisation, and testing the resilience of a control. However, these assessments don’t quantify the effect of controls on the security policies already in place. Measuring this effect can help stakeholders to re-calibrate and effectively prioritise their risks. Aditya Kuppa, Lamine M. Aouad, Nhien-An Le-Khac |
ACSAC | 3 |
| 2020 | Retracing the Flow of the Stream: Investigating Kodi Streaming Services
Samuel Todd Bromley, John Sheppard 0002, Mark Scanlon, Nhien-An Le-Khac |
ICDF2C | 4 |
| 2020 | Retrieving E-Dating Application Artifacts from iPhone Backups
Ranul Deelaka Thantilage, Nhien-An Le-Khac |
IFIP Int. Conf. Digital Forensics | 2 |
| 2020 | Black Box Attacks on Explainable Artificial Intelligence(XAI) methods in Cyber SecurityabstractCybersecurity community is slowly leveraging Machine Learning (ML) to combat ever evolving threats. One of the biggest drivers for successful adoption of these models is how well domain experts and users are able to understand and trust their functionality. As these black-box models are being employed to make important predictions, the demand for transparency and explainability is increasing from the stakeholders.Explanations supporting the output of ML models are crucial in cyber security, where experts require far more information from the model than a simple binary output for their analysis. Recent approaches in the literature have focused on three different areas: (a) creating and improving explainability methods which help users better understand the internal workings of ML models and their outputs; (b) attacks on interpreters in white box setting; (c) defining the exact properties and metrics of the explanations generated by models. However, they have not covered, the security properties and threat models relevant to cybersecurity domain, and attacks on explainable models in black box settings.In this paper, we bridge this gap by proposing a taxonomy for Explainable Artificial Intelligence (XAI) methods, covering various security properties and threat models relevant to cyber security domain. We design a novel black box attack for analyzing the consistency, correctness and confidence security properties of gradient based XAI methods. We validate our proposed system on 3 security-relevant data-sets and models, and demonstrate that the method achieves attacker's goal of misleading both the classifier and explanation report and, only explainability method without affecting the classifier output. Our evaluation of the proposed approach shows promising results and can help in designing secure and robust XAI methods. Aditya Kuppa, Nhien-An Le-Khac |
IJCNN | 2 |
| 2020 | Detecting Abnormal Traffic in Large-Scale NetworksabstractWith the rapid technological advancements, organizations need to rapidly scale up their information technology (IT) infrastructure viz. hardware, software, and services, at a low cost. However, the dynamic growth in the network services and applications creates security vulnerabilities and new risks that can be exploited by various attacks. For example, User to Root (U2R) and Remote to Local (R2L) attack categories can cause a significant damage and paralyze the entire network system. Such attacks are not easy to detect due to the high degree of similarity to normal traffic. While network anomaly detection systems are being widely used to classify and detect malicious traffic, there are many challenges to discover and identify the minority attacks in imbalanced datasets. In this paper, we provide a detailed and systematic analysis of the existing Machine Learning (ML) approaches that can tackle most of these attacks. Furthermore, we propose a Deep Learning (DL) based framework using Long Short Term Memory (LSTM) autoencoder that can accurately detect malicious traffics in network traffic. We perform our experiments in a publicly available dataset of Intrusion Detection Systems (IDSs). We obtain a significant improvement in attack detection, as compared to other benchmarking methods. Hence, our method provides great confidence in securing these networks from malicious traffic. Mahmoud Said Elsayed, Nhien-An Le-Khac, Soumyabrata Dev, Anca Jurcut |
ISNCC | 2 |
| 2020 | Towards A New Approach to Identify WhatsApp MessagesabstractToday traditional communication methods, such as SMS or phone calls, are used less often and are replaced by the use of chat applications. WhatsApp is one of the most popular chat applications nowadays. WhatsApp offers different ways of communicating, which include sending text messages and making phone calls. The implementation of encryption makes WhatsApp more challenging for law enforcement agencies to identify when a suspect is sending or receiving messages via this chat application. Most research in literature focused on the analysis of WhatsApp data by obtaining information from a physical device, such as a seized mobile device. However, it is not always possible to extract the data needed from a mobile device for the analysis of the WhatsApp data because of the encryption, or no devices have been seized yet. In addition, the current techniques for real time analysis of WhatsApp messages show that there is a high risk of detection by the suspect. Alternative methods are needed to understand the communication patterns of a suspect and criminal organizations. In this paper, we focused on identifying when a suspect is receiving or sending WhatsApp messages using only wiretap data. Therefore, no seized devices are needed. The pattern analysis has been used to identify patterns of data sent to and received from the WhatsApp servers. The identified patterns were tested against a large dataset created with different mobile devices to determine if the patterns are consistent. By using the technique described in this paper, investigators will obtain more information if and with whom a suspect is communicating. Rick Cents, Nhien-An Le-Khac |
TrustCom | 2 |
| 2020 | Towards a new deep learning based approach for the password predictionabstractAlthough tools for tracking and monitoring illegal networks have been developed for centuries, current methods available at the moment still need continues improvement. This is due to the fact that tracking and monitoring illegal networks in the cyberspace has become increasingly challenging for law enforcement agencies due to sophisticated encryption algorithms and strong passwords. Password predicting approaches will allow investigators to crack passwords used by criminals to protect their data or their communications. Hence, it will help judges to prosecute authors of crimes since they will have at their disposal all evidences needed which are, until now, hidden thanks to strong passwords. In this paper, we introduce a work-in-progress password guessing approach, called PassGuess to guess a missing character in a given password. PassGuess is powered by deep learning and it already predicts passwords with 80/% accuracy on the train set and 73\% accuracy on the test set for a random missing character in any position for a given password. This model is the very first preliminary work of the PassGuess intended for the purpose of the feasibility study and the authors of this paper are confident that the performance of PassGuess can be improved significantly in the future. Manaz Kaleel, Nhien-An Le-Khac |
TrustCom | 2 |
| 2020 | DDoSNet: A Deep-Learning Model for Detecting Network AttacksabstractSoftware-Defined Networking (SDN) is an emerging paradigm, which evolved in recent years to address the weaknesses in traditional networks. The significant feature of the SDN, which is achieved by disassociating the control plane from the data plane, facilitates network management and allows the network to be efficiently programmable. However, the new architecture can be susceptible to several attacks that lead to resource exhaustion and prevent the SDN controller from supporting legitimate users. One of these attacks, which nowadays is growing significantly, is the Distributed Denial of Service (DDoS) attack. DDoS attack has a high impact on crashing the network resources, making the target servers unable to support the valid users. The current methods deploy Machine Learning (ML) for intrusion detection against DDoS attacks in the SDN network using the standard datasets. However, these methods suffer several drawbacks, and the used datasets do not contain the most recent attack patterns - hence, lacking in attack diversity. In this paper, we propose DDoSNet, an intrusion detection system against DDoS attacks in SDN environments. Our method is based on Deep Learning (DL) technique, combining the Recurrent Neural Network (RNN) with autoencoder. We evaluate our model using the newly released dataset CICDDoS2019, which contains a comprehensive variety of DDoS attacks and addresses the gaps of the existing current datasets. We obtain a significant improvement in attack detection, as compared to other benchmarking methods. Hence, our model provides great confidence in securing these networks. Mahmoud Said Elsayed, Nhien-An Le-Khac, Soumyabrata Dev, Anca Jurcut |
WoWMoM | 2 |
| 2020 | Lightweight privacy-Preserving data classification
Ngoc Hong Tran, Nhien-An Le-Khac, M. Tahar Kechadi |
Comput. Secur. | 2 |
| 2020 | Smart vehicle forensics: Challenges and case study
Nhien-An Le-Khac, Daniel Jacobs, John Nijhoff, Karsten Bertens, Kim-Kwang Raymond Choo |
Future Gener. Comput. Syst. | 1 |
| 2019 | Improving Borderline Adulthood Facial Age Estimation through Ensemble LearningabstractAchieving high performance for facial age estimation with subjects in the borderline between adulthood and non-adulthood has always been a challenge. Several studies have used different approaches from the age of a baby to an elder adult and different datasets have been employed to measure the mean absolute error (MAE) ranging between 1.47 to 8 years. The weakness of the algorithms specifically in the borderline has been a motivation for this paper. In our approach, we have developed an ensemble technique that improves the accuracy of underage estimation in conjunction with our deep learning model (DS13K) that has been fine-tuned on the Deep Expectation (DEX) model. We have achieved an accuracy of 68% for the age group 16 to 17 years old, which is 4 times better than the DEX accuracy for such age range. We also present an evaluation of existing cloud-based and offline facial age prediction services, such as Amazon Rekognition, Microsoft Azure Cognitive Services, How-Old.net and DEX. Felix Anda, David Lillis, Aikaterini Kanta, Brett A. Becker, Elias Bou-Harb, Nhien-An Le-Khac, Mark Scanlon |
ARES | 6 |
| 2019 | Black Box Attacks on Deep Anomaly DetectorsabstractThe process of identifying the true anomalies from a given set of data instances is known as anomaly detection. It has been applied to address a diverse set of problems in multiple application domains including cybersecurity. Deep learning has recently demonstrated state-of-the-art performance on key anomaly detection applications, such as intrusion detection, Denial of Service (DoS) attack detection, security log analysis, and malware detection. Despite the great successes achieved by neural network architectures, models with very low test error have been shown to be consistently vulnerable to small, adversarially chosen perturbations of the input. The existence of evasion attacks during the test phase of machine learning algorithms represents a significant challenge to both their deployment and understanding. Aditya Kuppa, Slawomir Grzonkowski, Muhammad Rizwan Asghar, Nhien-An Le-Khac |
ARES | 4 |
| 2019 | Efficient LiDAR point cloud data encoding for scalable data management within the Hadoop eco-systemabstractThis paper introduces a novel LiDAR point cloud data encoding solution that is compact, flexible, and fully supports distributed data storage within the Hadoop distributed computing environment. The proposed data encoding solution is developed based on Sequence File and Google Protocol Buffers. Sequence File is a generic splittable binary file format built in the Hadoop framework for storage of arbitrary binary data. The key challenge in adopting the Sequence File format for LiDAR data is in the strategy for effectively encoding the LiDAR data as binary sequences in a way that the data can be represented compactly, while allowing necessary mutation. For that purpose, a data encoding solution, based on Google Protocol Buffers (a language-neutral, cross-platform, extensible data serialisation framework) was developed and evaluated. Since neither of the underlying technologies is sufficient to completely and efficiently represent all necessary point formats for distributed computing, an innovative fusion of them was required to provide a viable data storage solution. This paper presents the details of such a data encoding implementation and rigorously evaluates the efficiency of the proposed data encoding solution. Benchmarking was done against a straightforward, naive text encoding implementation using a high-density aerial LiDAR scan of a portion of Dublin, Ireland. The results demonstrated a 6-times reduction in data volume, a 4-times reduction in database ingestion time, and up to a 5 times reduction in querying time. Anh-Vu Vo, Chamin Nalinda Lokugam Hewage, Gianmarco Russo, Neel Chauhan, Debra F. Laefer, Michela Bertolotto, Nhien-An Le-Khac, Ulrich Ofterdinger |
IEEE BigData | 7 |
| 2019 | A Graph Database-Based Approach to Analyze Network Log Files
Lars Diederichsen, Kim-Kwang Raymond Choo, Nhien-An Le-Khac |
NSS | 3 |
| 2018 | Accuracy Enhancement of Electromagnetic Side-Channel Attacks on Computer MonitorsabstractElectromagnetic noise emitted from running computer displays modulates information about the picture frames being displayed on screen. Attacks have been demonstrated on eavesdropping computer displays by utilising these emissions as a side-channel vector. The accuracy of reconstructing a screen image depends on the emission sampling rate and bandwidth of the attackers signal acquisition hardware. The cost of radio frequency acquisition hardware increases with increased supported frequency range and bandwidth. A number of enthusiast-level, affordable software defined radio equipment solutions are currently available facilitating a number of radio-focused attacks at a more reasonable price point. This work investigates three accuracy influencing factors, other than the sample rate and bandwidth, namely noise removal, image blending, and image quality adjustments, that affect the accuracy of monitor image reconstruction through electromagnetic side-channel attacks. Asanka P. Sayakkara, Nhien-An Le-Khac, Mark Scanlon |
ARES | 2 |
| 2018 | Solid State Drive Forensics: Where Do We Stand?
John Vieyra, Mark Scanlon, Nhien-An Le-Khac |
ICDF2C | 3 |
| 2018 | Internet of Things Forensics - Challenges and a Case Study
Saad Alabdulsalam, M. Tahar Kechadi, Nhien-An Le-Khac |
IFIP Int. Conf. Digital Forensics | 4 |
| 2018 | Enabling Non-Expert Analysis OF Large Volumes OF Intercepted Network Traffic
Erwin van de Wiel, Mark Scanlon, Nhien-An Le-Khac |
IFIP Int. Conf. Digital Forensics | 3 |
| 2016 | The End of Effective Law Enforcement in the Cloud? - To Encrypt, or Not to EncryptabstractWith an exponentially increasing usage of cloud services, the need for forensic investigations of virtual space is equally in constantly increasing demand, which includes as a very first approach, the gaining of access to it as well as the data stored. This is an aspect that faces a number of challenges, stemming not only from the technical difficulties and peculiarities, but equally covers the interaction with an emerging line of businesses offering cloud storage and services. Beyond the forensic aspects, it also covers to an ever increasing amount the non-forensic considerations, such as the availability of logs and archives, legal and data protection considerations from a global perspective and the clashes in between, as well as the ever competing interests between law enforcement to seize evidence which is non-physical, and businesses who need to be able to continue to operate and provide their hosted services, even if law enforcement seek to collect evidence. The trend post-Snowden has been unequivocally towards default encryption, and driven by market leaders such as Apple, motivated to a large extent by the perceived demands for privacy of the consumer. The central question to be explored in this paper is to what extent this trend towards default encryption will have a negative impact on law enforcement investigations and possibilities, and will at the end attempt to provide a solution, which takes into account the needs of both law enforcement, but also of the cloud service providers. It is hoped that the recommendations from this paper will be able to have an impact in the ability for law enforcement to continue with their investigations in an efficient manner, whilst also safeguarding the ability for business to thrive and continue to develop and offer new and innovative solutions, which do not put law enforcement at risk. Steven Ryder, Nhien-An Le-Khac |
CLOUD | 2 |
| 2016 | Efficient Large Scale Clustering Based on Data PartitioningabstractClustering techniques are very attractive for extracting and identifying patterns in datasets. However, their application to very large spatial datasets presents numerous challenges such as high-dimensionality data, heterogeneity, and high complexity of some algorithms. For instance, some algorithms may have linear complexity but they require the domain knowledge in order to determine their input parameters. Distributed clustering techniques constitute a very good alternative to the big data challenges (e.g.,Volume, Variety, Veracity, and Velocity). Usually these techniques consist of two phases. The first phase generates local models or patterns and the second one tends to aggregate the local results to obtain global models. While the first phase can be executed in parallel on each site and, therefore, efficient, the aggregation phase is complex, time consuming and may produce incorrect and ambiguous global clusters and therefore incorrect models. In this paper we propose a new distributed clustering approach to deal efficiently with both phases, generation of local results and generation of global models by aggregation. For the first phase, our approach is capable of analysing the datasets located in each site using different clustering techniques. The aggregation phase is designed in such a way that the final clusters are compact and accurate while the overall process is efficient in time and memory allocation. For the evaluation, we use two well-known clustering algorithms, K-Means and DBSCAN. One of the key outputs of this distributed clustering technique is that the number of global clusters is dynamic, no need to be fixed in advance. Experimental results show that the approach is scalable and produces high quality results. Malika Bendechache, M. Tahar Kechadi, Nhien-An Le-Khac |
DSAA | 3 |
| 2016 | Improving Fitness Functions in Genetic Programming for Classification on Unbalanced Credit Card Data
Van Loi Cao, Nhien-An Le-Khac, Michael O'Neill 0001, Miguel Nicolau, James McDermott |
EvoApplications (1) | 2 |
| 2015 | Overview of the Forensic Investigation of Cloud ServicesabstractCloud Computing is a commonly used, yet ambiguous term, which can be used to refer to a multitude of differing dynamically allocated services. From a law enforcement and forensic investigation perspective, cloud computing can be thought of as a double edged sword. While on one hand, the gathering of digital evidence from cloud sources can bring with it complicated technical and cross-jurisdictional legal challenges. On the other, the employment of cloud storage and processing capabilities can expedite the forensics process and focus the investigation onto pertinent data earlier in an investigation. This paper examines the state-of-the-art in cloud-focused, digital forensic practises for the collection and analysis of evidence and an overview of the potential use of cloud technologies to provide Digital Forensics as a Service. Jason Farina, Mark Scanlon, Nhien-An Le-Khac, M. Tahar Kechadi |
ARES | 3 |
| 2015 | Towards the Forensic Identification and Investigation of Cloud Hosted Servers through Non-Invasive WiretapsabstractWhen conducting modern cybercrime investigations, evidence has often to be gathered from computer systems located at cloud-based data centres of hosting providers. In cases where the investigation cannot rely on the cooperation of the hosting provider, or where documentation is not available, investigators can often find the identification of which distinct server among many is of interest difficult and extremely time consuming. To address the problem of identifying these servers, in this paper a new approach to rapidly and reliably identify these cloud hosting computer systems is presented. In the outlined approach, a handheld device composed of an embedded computer combined with a method of undetectable interception of Ethernet based communications is presented. This device is tested and evaluated, and a discussion is provided on its usefulness in identifying of server of interest to an investigation. Hessel Schut, Mark Scanlon, Jason Farina, Nhien-An Le-Khac |
ARES | 4 |
| 2014 | Forensic Analysis of the TomTom Navigation Application
Nhien-An Le-Khac, Mark Roeloffs, M. Tahar Kechadi |
IFIP Int. Conf. Digital Forensics | 1 |
| 2013 | Feature Selection Parallel Technique for Remotely Sensed Imagery Classification
Nhien-An Le-Khac, Bo Wu 0019, Chongcheng Chen, M. Tahar Kechadi |
ICCSA (2) | 1 |
| 2012 | An Open Framework for Smartphone Evidence Acquisition
Lamine M. Aouad, M. Tahar Kechadi, Justin Trentesaux, Nhien-An Le-Khac |
IFIP Int. Conf. Digital Forensics | 4 |
| 2011 | A New Hybrid Clustering Method for Reducing Very Large Spatio-temporal Dataset
Michael Whelan, Nhien-An Le-Khac, M. Tahar Kechadi |
ADMA (1) | 2 |
| 2010 | A Clustering-Based Data Reduction for Very Large Spatio-Temporal Datasets
Nhien-An Le-Khac, Martin Bue, Michael Whelan, M. Tahar Kechadi |
ADMA (2) | 1 |
| 2010 | Performance study of distributed Apriori-like frequent itemsets mining
Lamine M. Aouad, Nhien-An Le-Khac, M. Tahar Kechadi |
Knowl. Inf. Syst. | 2 |
| 2009 | Towards a New Data Mining-Based Approach for Anti-Money Laundering in an International Investment Bank
Nhien-An Le-Khac, Sammer Markos, M. Tahar Kechadi |
ICDF2C | 1 |
| 2008 | Persistent Workflow on the GridabstractThe huge data requirements of large nowadays applications, in science, engineering, and commerce, make efficient data placement an essential need. For this purpose, we propose a framework which can be considered as a grid workflow system with advanced data placement capabilities. Computational jobs and data placement are handled concurrently. This framework also includes a specialised scheduling for data placement, and presents a multi-level architecture which allows support for a range of grid systems and middleware. We present results of improving application performance through data placement optimisations. A basic data mining application and an astronomy application are used as the basis of this study. Lamine M. Aouad, Nhien-An Le-Khac, M. Tahar Kechadi |
APSCC | 2 |
| 2008 | Handling Large Volumes of Mined Knowledge with a Self-Reconfigurable Topology on Distributed SystemsabstractNowadays, massive amounts of data which are often geographically distributed and owned by different organisations, are being mined. As consequence, large volumes of knowledge is being generated. This causes the problem of efficient knowledge management in distributed data mining ({\it DDM}). The main aim of {\it DDM} is to exploit fully the benefit of distributed data analysis while minimising the communication overhead. Existing {\it DDM} techniques perform partial analysis of local data at individual sites and then generate global models by aggregating the local results. These two steps are not independent since naive approaches to local analysis may produce incorrect and ambiguous global data models. To overcome this problem, we introduce a distributed knowledge map based on an efficient self-reconfiguration network topology to represent easily and exploit efficiently the knowledge mined in large scale distributed platforms. This will also facilitate the integration/coordination of local mining processes and existing knowledge to build global models. In this paper, we implement this knowledge map and present some preliminary results about its performance. Nhien-An Le-Khac, Lamine M. Aouad, M. Tahar Kechadi |
ICMLA | 1 |