Khanh Luong

dblp:228/4094 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0001-6981-7367ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Ransomware Encryption Detection: Adaptive File System Analysis Against Evasive Encryption Tactics
Arash Mahboubi, Hamed Aboutorab, Seyit Ahmet Çamtepe, Hang Thanh Bui, Khanh Luong, Keyvan Ansari, Shenlu Wang, Bazara I. A. Barry
ACISP (3)5
2025 ConceptUML: Multiphase unsupervised threat detection via latent concept learning, Hidden Markov Models and topic modelling
abstract
Detecting lateral movement threats in large-scale system logs is a critical challenge due to the scarcity of labelled attack data, the presence of imbalanced datasets, and the sophisticated nature of modern adversaries. To address these issues, we propose ConceptUML , a semantic-driven, fully unsupervised threat detection framework designed to automatically identify anomalies related to lateral movement in heterogeneous log data. ConceptUML is structured around a three-phase architecture. In Phase 1 (Latent Semantic Learning) , contextualized embeddings generated by Sentence-BERT are combined with Non-negative Matrix Factorization to extract abstract concepts from system logs and external threat intelligence sources such as MITRE ATT&CK and CAPEC. In Phase 2 (Unsupervised Threat Detection) , a Hidden Markov Model is applied to cluster logs based on learned concepts, and each cluster is scored according to its semantic similarity to known adversarial techniques. Phase 3 (Decision Refinement) uses topic modelling to further isolate malicious event log subsets from within suspicious clusters, enabling high-precision triage. We evaluate ConceptUML using four real-world event log datasets, including Windows Event Logs and multiple subsets of the LMD-23 dataset, encompassing attacks such as exploitation of hashing techniques and remote services. The enhanced model with topic modelling achieves up to 92.54% detection quality and reduces detection error to as low as 8.14%, outperforming several baseline approaches including AutoEncoder, LogAnomaly, LOF, and DBScan. Our results confirm that ConceptUML delivers interpretable, scalable, and highly effective detection of lateral movement threats without requiring labelled training data or extensive manual feature engineering.
Khanh Luong, Arash Mahboubi, Geoff Jarrad, Seyit Ahmet Çamtepe, Michael Bewong, Mohammed Bahutair, Hamed Aboutorab, Hang Thanh Bui
J. Inf. Secur. Appl.1
2025 Lurking in the shadows: Unsupervised decoding of beaconing communication for enhanced cyber threat hunting
abstract
The escalating prevalence of Advanced Persistent Threats (APTs) necessitates the development of more robust solutions capable of effectively thwarting these attacks by monitoring system activities across individual hosts. Existing cloud-native security applications utilize a combination of rule-based and machine learning-based detection techniques to protect digital assets . However, these approaches have limitations. Rule-based detection depends on predefined rules to identify specific attack patterns. Persistent attackers can often evade detection by carefully ensuring that their behavior circumvents these rules. In contrast, machine learning-based detection techniques, which learn attack patterns from data, rely heavily on the availability of labeled data for training. However, labeled data is often unavailable and can be labor-intensive and costly to obtain. In this paper, we address the challenge of detecting APT attacks more holistically by leveraging attackers’ behavior during communication with Command and Control (C2) servers, a critical phase observed in most APT attacks. We aim to reduce false positive alerts for threat hunters by analyzing system network logs to detect potential network beaconing, a common attribute of various malware . We introduce a novel hybrid approach, called NetSpectra Sentinel , which employs a Continuous Time Hidden Markov Model (CT-HMM) to detect hidden states underlying observed patterns within the network logs and Time Series Decomposition (TSD) to model temporal patterns. We evaluate the effectiveness of our approach using 14 benchmark datasets and one synthetic dataset , comparing our method with other state-of-the-art statistical-based and botnet detection techniques. The results demonstrate that our technique achieves significantly higher accuracy in most cases, and even when existing techniques fail, our approach can still detect beaconing post-initial compromise with up to 90% accuracy. Additionally, we achieve up to four times better performance in terms of precision compared to existing statistical-based techniques.
Arash Mahboubi, Khanh Luong, Geoff Jarrad, Seyit Ahmet Çamtepe, Michael Bewong, Mohammed Bahutair, Ganna Pogrebna
J. Netw. Comput. Appl.2
2024 A Lightweight Detection of Sequential Patterns in File System Events During Ransomware Attacks
Arash Mahboubi, Hang Thanh Bui, Hamed Aboutorab, Khanh Luong, Seyit Ahmet Çamtepe, Keyvan Ansari
WISE (5)4
2024 Evolving techniques in cyber threat hunting: A systematic review
abstract
In the rapidly changing cybersecurity landscape, threat hunting has become a critical proactive defense against sophisticated cyber threats. While traditional security measures are essential, their reactive nature often falls short in countering malicious actors’ increasingly advanced tactics. This paper explores the crucial role of threat hunting, a systematic, analyst-driven process aimed at uncovering hidden threats lurking within an organization's digital infrastructure before they escalate into major incidents. Despite its importance, the cybersecurity community grapples with several challenges, including the lack of standardized methodologies, the need for specialized expertise, and the integration of cutting-edge technologies like artificial intelligence (AI) for predictive threat identification. To tackle these challenges, this survey paper offers a comprehensive overview of current threat hunting practices, emphasizing the integration of AI-driven models for proactive threat prediction. Our research explores critical questions regarding the effectiveness of various threat hunting processes and the incorporation of advanced techniques such as augmented methodologies and machine learning. Our approach involves a systematic review of existing practices, including frameworks from industry leaders like IBM and CrowdStrike. We also explore resources for intelligence ontologies and automation tools. The background section clarifies the distinction between threat hunting and anomaly detection, emphasizing systematic processes crucial for effective threat hunting. We formulate hypotheses based on hidden states and observations, examine the interplay between anomaly detection and threat hunting, and introduce iterative detection methodologies and playbooks for enhanced threat detection. Our review encompasses supervised and unsupervised machine learning approaches, reasoning techniques, graph-based and rule-based methods, as well as other innovative strategies. We identify key challenges in the field, including the scarcity of labeled data, imbalanced datasets, the need for integrating multiple data sources, the rapid evolution of adversarial techniques, and the limited availability of human expertise and data intelligence. The discussion highlights the transformative impact of artificial intelligence on both threat hunting and cybercrime, reinforcing the importance of robust hypothesis development. This paper contributes a detailed analysis of the current state and future directions of threat hunting, offering actionable insights for researchers and practitioners to enhance threat detection and mitigation strategies in the ever-evolving cybersecurity landscape.
Arash Mahboubi, Khanh Luong, Hamed Aboutorab, Hang Thanh Bui, Geoff Jarrad, Mohammed Bahutair, Seyit Ahmet Çamtepe, Ganna Pogrebna, Bazara I. A. Barry, Hannah Gately
J. Netw. Comput. Appl.2
2024 DCCNMF: Deep Complementary and Consensus Non-negative Matrix Factorization for multi-view clustering
Sohan Gunawardena, Khanh Luong, Balasubramaniam Thirunavukarasu, Richi Nayak
Knowl. Based Syst.2
2022 Multi-layer manifold learning for deep non-negative matrix factorization-based multi-view clustering
Khanh Luong, Richi Nayak, Balasubramaniam Thirunavukarasu, Md. Abul Bashar
Pattern Recognit.1
2022 Learning Inter- and Intra-Manifolds for Matrix Factorization-Based Multi-Aspect Data Clustering
abstract
Clustering on the data with multiple aspects, such as multi-view or multi-type relational data, has become popular in recent years due to their wide applicability. The approach using manifold learning with the Non-negative Matrix Factorization (NMF) framework, that learns the accurate low-rank representation of the multi-dimensional data, has shown effectiveness. We propose to include the inter-manifold in the NMF framework, utilizing the distance information of data points of different data types (or views) to learn the diverse manifold for data clustering. Empirical analysis reveals that the proposed method can find partial representations of various interrelated types and select useful features during clustering. Results on several datasets demonstrate that the proposed method outperforms the state-of-the-art multi-aspect data clustering methods in both accuracy and efficiency.
Khanh Luong, Richi Nayak
IEEE Trans. Knowl. Data Eng.1
2020 A Novel Approach to Learning Consensus and Complementary Information for Multi-View Data Clustering
abstract
Effective methods are required to be developed that can deal with the multi-faceted nature of the multi-view data. We design a factorization-based loss function-based method to simultaneously learn two components encoding the consensus and complementary information present in multi-view data by using the Coupled Matrix Factorization (CMF) and Non-negative Matrix Factorization (NMF). We propose a novel optimal manifold for multi-view data which is the most consensed manifold embedded in the high-dimensional multi-view data. A new complementary enhancing term is added in the loss function to enhance the complementary information inherent in each view. An extensive experiment with diverse datasets, benchmarking the state-of-the-art multi-view clustering methods, has demonstrated the effectiveness of the proposed method in obtaining accurate clustering solution.
Khanh Luong, Richi Nayak
ICDE1
2019 Multi-type Relational Data Clustering for Community Detection by Exploiting Content and Structure Information in Social Networks
T. M. G. Tennakoon, Khanh Luong, Wathsala Anupama Mohotti, Sharma Chakravarthy, Richi Nayak
PRICAI (2)2
2018 Learning Association Relationship and Accurate Geometric Structures for Multi-Type Relational Data
abstract
Non-negative Matrix Factorization (NMF) methods have been effectively used for clustering high dimensional data. Manifold learning is combined with the NMF framework to ensure the projected lower dimensional representations preserve the local geometric structure of data. In this paper, considering the context of multi-type relational data clustering, we develop a new formulation of manifold learning to be embedded in the factorization process such that the new low-dimensional space can maintain both local and global structures of original data. We also propose to include the interactions between clusters of different data types by enforcing a Normalize Cut-type constraint that leads to a comprehensive NMF-based framework. A theoretical analysis and extensive experiments are provided to validate the effectiveness of the proposed work.
Khanh Luong, Richi Nayak
ICDE1
2018 A Novel Technique of Using Coupled Matrix and Greedy Coordinate Descent for Multi-view Data Representation
Khanh Luong, Balasubramaniam Thirunavukarasu, Richi Nayak
WISE (2)1