EDBT 2026 Demo / reviewers in the wild / expert
Panos Kostakos 0001
dblp:211/9048 · also Panagiotis Kostakos 0001, Panos A. Kostakos 0001
· DBLP profile ↗
11ranked-venue papers in the field
3as first author
6since 2021 · last 2025
0000-0002-8545-599XORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 7 (3 first)Big Data, Cloud & Distributed Data Systems · 3Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cognitive SOC: Evidence-Backed Narrative Generation for Security Operations with Multi-Agent LLM Architecture
Saeid Sheikhi, Panos Kostakos 0001, Lauri Lovén |
IEEE Big Data | 2 |
| 2024 | Hex2Sign: Automatic IDS Signature Generation from Hexadecimal Data using LLMsabstractDespite the growing utilization of large language models (LLMs) in cyber defense operations, their integration within intrusion detection systems (IDS) remains substantially underexplored. This paper proposes a novel approach to generating human-readable IDS signatures by fine-tuning LLMs on hexadecimal data. In our experimental framework, we deploy honeypots to capture malicious network traffic in real-world conditions, generating packet capture (PCAP) files accompanied by text-based alerts and Suricata signatures. The collected hexadecimal data, derived from actual attack vectors, serves as the training corpus for multiple generative and classification models, which are fine-tuned for optimal performance in generating human-readable IDS alerts. According to the results, generative model GPT-3-Davinci-002 excelled across metrics with BERTscore over 96%, while RoBERTa base achieved high accuracy of 96% among classifiers. These findings enhance our understanding that foundational models can improve hexadecimal data processing for cybersecurity. Our conclusions emphasize the potential of advanced generative-AI models in automating dynamic Suricata rule generation, thus enhancing IDS efficiency and accuracy. Moreover, this paper proposes an AI-powered IDS system for securing network environments that can significantly mitigate the risks associated with diverse and widespread devices. By integrating LLMs into security frameworks, this system offers a robust defense mechanism that dynamically adapts to emerging threats, thus enhancing IDS efficiency and accuracy in handling big data challenges. Prasasthy Balasubramanian, Tarek Ali, Mohammad Salmani, Danial Khosh Kholgh, Panos Kostakos 0001 |
IEEE Big Data | 5 |
| 2024 | Addressing Data Challenges to Drive the Transformation of Smart CitiesabstractCities serve as vital hubs of economic activity and knowledge generation and dissemination. As such, cities bear a significant responsibility to uphold environmental protection measures while promoting the welfare and living comfort of their residents. There are diverse views on the development of smart cities, from integrating Information and Communication Technologies into urban environments for better operational decisions to supporting sustainability, wealth, and comfort of people. However, for all these cases, data are the key ingredient and enabler for the vision and realization of smart cities. This article explores the challenges associated with smart city data. We start with gaining an understanding of the concept of a smart city, how to measure that the city is a smart one, and what architectures and platforms exist to develop one. Afterwards, we research the challenges associated with the data of the cities, including availability, heterogeneity, management, analysis, privacy, and security. Finally, we discuss ethical issues. This article aims to serve as a “one-stop shop” covering data-related issues of smart cities with references for diving deeper into particular topics of interest. Ekaterina Gilman, Francesca Bugiotti, Ahmed Khalid, Hassan Mehmood, Panos Kostakos 0001, Lauri Tuovinen, Johanna Ylipulli, Xiang Su 0001, Denzil Ferreira |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2023 | Transformer-based LLMs in Cybersecurity: An in-depth Study on Log Anomaly Detection and Conversational Defense MechanismsabstractWith the advancement of conversational AI and Large Language Models (LLMs), interactive chatbots are emerging as pivotal assets for connecting with users across various sectors, enabling various capabilities and functions. However, their potential in the cybersecurity domain remains largely untapped. This article introduces a novel method to enhance chatbot performance by incorporating anomaly detection features. Our chatbot uses advanced GPT-3 models and rule-based logic to identify and extract unusual patterns and deviations within logs, making it more proficient in detecting anomalies. We present the architecture and methodology behind our anomaly detection system, showcasing its effectiveness in real-world scenarios. Combining machine learning and domain expertise, our chatbot sets a new standard in interactive, anomaly-aware conversational agents. Our anomaly detection classifier was able to achieve more than 99% of accuracy by illustrating its robust performance in accurately identifying and flagging outliers or unusual patterns in log file data. We also compared the performance of GPT-3 models with other LLMs: BERT, DistilBERT, and ALBERT. Our findings concluded that GPT-3 models consistently outperform all the other LLM models and exhibit significantly higher performance. Prasasthy Balasubramanian, Justin Seby, Panos Kostakos 0001 |
IEEE Big Data | 3 |
| 2021 | Using smart glasses for monitoring cyber threat intelligence feedsabstractThe surge of COVID-19 has introduced a new threat surface as malevolent actors are trying to benefit from the pandemic. Because of this, new information sources and visualization tools about COVID-19 have been introduced into the workflow of frontline practitioners. As a result, analysts are increasingly required to shift their focus between different visual displays to monitor pandemic related data, security threats, and incidents. Augmented reality (AR) smart glasses can overlay digital data to the physical environment in a comprehensible manner. However, the real-life use situations are often complex and require fast knowledge acquisition from multiple sources. In this study we report results from an experiment with six subjects using an AR overlaid information interface coupled with traditional computer monitors. Our goal was to evaluate a multi tasking setup with traditional monitors and an AR headset where notifications from the new COVID-19 MISP instance were visualized. Our results indicate that better situational awareness does translate to increased task performance, but at the cost of a gender gap that requires further attention. Mikko Korkiakoski, Fatima Sadiq, Febrian Setianto, Ummi Khaira Latif, Paula Alavesa, Panos Kostakos 0001 |
ASONAM | 6 |
| 2021 | GPT-2C: a parser for honeypot logs using large pre-trained language modelsabstractDeception technologies like honeypots generate large volumes of log data, which include illegal Unix shell commands used by latent intruders. Several prior works have reported promising results in overcoming the weaknesses of network-level and program-level Intrusion Detection Systems (IDSs) by fussing network traffic with data from honeypots. However, because honeypots lack the plug-in infrastructure to enable real-time parsing of log outputs, it remains technically challenging to feed illegal Unix commands into downstream predictive analytics. As a result, advances on honeypot-based user-level IDSs remain greatly hindered. This article presents a run-time system (GPT-2C) that leverages a large pre-trained language model (GPT-2) to parse dynamic logs generated by a live Cowrie SSH honeypot instance. After fine-tuning the GPT-2 model on an existing corpus of illegal Unix commands, the model achieved 89% inference accuracy in parsing Unix commands with acceptable execution latency. Febrian Setianto, Erion Tsani, Fatima Sadiq, Georgios Domalis, Dimitris Tsakalidis, Panos Kostakos 0001 |
ASONAM | 6 |
| 2020 | Strings and Things: A Semantic Search Engine for news quotes using Named Entity RecognitionabstractEmerging methods for content delivery such as quote-searching and entity-searching, enable users to quickly identify novel and relevant information from unstructured texts, news articles, and media sources. These methods have widespread applications in web surveillance and crime informatics, and can help improve intention disambiguation, character evaluation, threat analysis, and bias detection. Furthermore, quote-based and entity-based searching is also an empowering information retrieval tool that can enable non-technical users to gauge the quality of public discourse, allowing for more fine-grained analysis of core sociological questions. The paper presents a prototype search engine that allows users to search a news database containing quotes using a combination of strings and things. The ingestion pipeline, which forms the backend of the service, comprises of the following modules i) a crawler that ingests data from the GDELT Global Quotation Graph ii) a named entity recognition (NER) filter that labels data on the fly iii) an indexing mechanism that serves the data to an Elasticsearch cluster and iv) a user interface that allows users to formulate queries. The paper presents the high-level configuration of the pipeline and reports basic metrics and aggregations. Panos Kostakos 0001 |
ASONAM | 1 |
| 2020 | MaTED: Metadata-Assisted Twitter Event Detection System
Abhinay Pandya, Mourad Oussalah 0002, Panos Kostakos 0001, Ummul Fatima |
IPMU (1) | 3 |
| 2018 | Meta-Terrorism: Identifying Linguistic Patterns in Public Discourse After an AttackabstractWhen a terror-related event occurs, there is a surge of traffic on social media comprising of informative messages, emotional outbursts, helpful safety tips, and rumors. It is important to understand the behavior manifested on social media sites to gain a better understanding of how to govern and manage in a time of crisis. We undertook a detailed study of Twitter during two recent terror-related events: the Manchester attacks and the Las Vegas shooting. We analyze the tweets during these periods using (a) sentiment analysis, (b) topic analysis, and (c) fake news detection. Our analysis demonstrates the spectrum of emotions evinced in reaction and the way those reactions spread over the event timeline. Also, with respect to topic analysis, we find “echo chambers”, groups of people interested in similar aspects of the event. Encouraged by our results on these two event datasets, the paper seeks to enable a holistic analysis of social media messages in a time of crisis. Panos Kostakos 0001, Markus Nykanen, Mikael Martinviita, Abhinay Pandya, Mourad Oussalah 0002 |
ASONAM | 1 |
| 2018 | Covert Online Ethnography and Machine Learning for Detecting Individuals at Risk of Being Drawn into Online Sex WorkabstractHow can we identify individuals at risk of being drawn into online sex work? The spread of online communication removes transaction costs and enables a greater number of people to be involved in illicit activities, including online sex trade. As a result, social media platforms often work as springboard for criminal careers posing a significant risk to the economy, public health and trust. Detecting deviant behaviors online is limited by the poor availability of ground-truth data and machine learning tools. Unlike prior work which focuses exclusively on either qualitative or quantitative methods, in this paper we combine covert online ethnography with semi-supervised learning methodologies, using data from a popular European adult forum. We obtained risk assessment results of 78 users using covert online ethnography, and set out to build a machine learning model that can predict the risk factor in other 28,832 users. Results show that a combination-based approach in which all features are used yields the most accurate results. Panos Kostakos 0001, Lucie Sprachalova, Abhinay Pandya, Mohamed Aboeleinen, Mourad Oussalah 0002 |
ASONAM | 1 |
| 2018 | SemanPhone: Combining Semantic and Phonetic Word Association in Verbal Learning ContextabstractThis paper proposes an effective way to discover and memorize new English vocabulary based on both semantic and phonetic associations. The method we proposed aims to automatically find out the most associated words of a given target word. The measurement of semantic association was achieved by calculating cosine similarity of two-word vectors, and the measurement of phonetic association was achieved by calculating the longest common subsequence of phonetic symbol strings of two words. Finally, the method was implemented as a web application. Jiyan Lu, Panos Kostakos 0001, Mourad Oussalah 0002, Susanna Pirttikangas |
ASONAM | 2 |