VLDB 2026 Research / reviewers in the wild / expert
Bishal Lakha
dblp:339/8133
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0001-8234-4389ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAPTOR: Cyber Attack Protection via Temporal Online Graph Representation LearningabstractOnline intrusion detection systems (IDSs), i.e., tools for detecting live cyberattacks in the form of unauthorized access gained to a (networking) system, play a role of paramount importance in the cybersecurity landscape. Among the plethora of existing IDSs, the ones based on temporal graph anomaly detection (TGAD) process a temporal graph representing entities of interest of the underlying system and time-varying relationships among them, and identify anomalous temporal edges in such a graph as potential intrusions. TGAD-based IDSs are superior to various other existing types of IDS for their peculiarities of high generality and powerfulness in data representation and types of cyberattack identifiable. However, existing TGAD-based IDSs are still far from being suitable for real-world settings, due to their severe limitations in efficiently and effectively handling the underlying temporal graphs, which are typically really big. In this paper, we devise CAPTOR (“Cyber Attack Protection via Temporal Online graph Representation learning”), a novel TGAD-based IDS which addresses the limitations of the state of the art. CAPTOR consists of a careful selection and clever combination of graph representation learning (GRL), TGAD, and temporal aggregation of GRL representations (embeddings). These design choices make CAPTOR achieve the best tradeoff between accuracy and scalability in TGAD-based intrusion detection, as testified by extensive experiments on real cybersecurity datasets. As such, with this work we take a significant step forward towards rendering the important TGAD-based IDS technology actually applicable in real-world cybersecurity scenarios. Bishal Lakha, Janet Layne, Edoardo Serra, Francesco Gullo, Sushil Jajodia |
IEEE Trans. Big Data | 1 |
| 2025 | Inconsistent Reasoning Attacks to Identify Weaknesses in Automatic Scientific Claim Verification Tools
Md Athikul Islam, Noel Ellison, Bishal Lakha, Edoardo Serra |
ECML/PKDD (7) | 3 |
| 2023 | Prediction of Future Nation-initiated Cyberattacks from News-based Political Event GraphabstractIn the world of cyber defense, anticipating potential attacks or any increase in risk of attacks is one of the most advantageous pieces of knowledge one can have. However, little research has been done in examining the larger geopolitical environment and using data sources available at the geopolitical level to predict cyberattacks in advance. To this end, we combine the use of a geopolitical conflict dataset, ICEWS, in combination with a cyberattack dataset from the Council on Foreign Relations to determine if we can predict cyberattacks targeting a given nation. We present a novel approach to identify periods of increased likelihood of cyberattacks at the country, regional, and global levels. The approach involves creating a news-based political event graph, generating vectorial representations of the graph using the SIR-GN structural iterative representation learning approach, and applying novelty detection models to predict future nation-initiated attacks. The proposed approach outperforms existing baselines for majority of cases in terms of F1-score, demonstrating its effectiveness in predicting cyberattacks. Bishal Lakha, Jason Duran, Edoardo Serra, Francesca Spezzano |
DSAA | 1 |
| 2023 | Analysis of Software Engineering Practices in General Software and Machine Learning StartupsabstractContext: On top of the inherent challenges startup software companies face applying proper software engineering practices, the non-deterministic nature of machine learning techniques makes it even more difficult for machine learning (ML) startups. Objective: Therefore, the objective of our study is to understand the whole picture of software engineering practices followed by ML startups and identify additional needs. Method: To achieve our goal, we conducted a systematic literature review study on 37 papers published in the last 21 years. We selected papers on both general software startups and ML startups. We collected data to understand software engineering (SE) practices in five phases of the software development life-cycle: requirement engineering, design, development, quality assurance, and deployment. Results: We find some interesting differences in software engineering practices in ML startups and general software startups. The data management and model learning phases are the most prominent among them. Conclusion: While ML startups face many similar challenges to general software startups, the additional difficulties of using stochastic ML models require different strategies in using software engineering practices to produce high-quality products. Bishal Lakha, Kalyan Bhetwal, Nasir U. Eisty |
SERA | 1 |
| 2022 | Anomaly Detection in Cybersecurity Events Through Graph Neural Network and Transformer Based Model: A Case Study with BETH DatasetabstractWith the increasing prevalence of the internet, detecting malicious behavior is becoming a greater need. This problem can be formulated as an anomaly detection task on provenance data, where attacks are detectable as anomalies in the behavior of the system. While network data is quite prevalent, we focus on system logs and propose a novel approach with two main components. The first is to make use of the graph-like structure of the logs in which processes enact events and generate additional processes, using a graph neural network (GNN) to produce representations of each event which encode information about their neighboring events in an unsupervised manner. The second is to make use of the complex features such as command arguments which vary widely and cannot be used in the presented format as features in typical machine learning algorithms. If these features are instead encoded using transformer models, they can then be used in other algorithms such as a GNN or anomaly detector. These two approaches combined improve anomaly detection results for the BETH dataset by around 8 percent as compared to the manually engineered features alone. Bishal Lakha, Sara Lilly Mount, Edoardo Serra, Alfredo Cuzzocrea |
IEEE Big Data | 1 |