EDBT 2026 Demo / reviewers in the wild / expert
Moumita Das Purba
dblp:280/6342
· DBLP profile ↗
3ranked-venue papers
3as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 3 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Automated and Explainable Threat Hunting with Generative AIabstractThis paper describes an attempt to automate threat hunting, taking cyber threat intelligence messages as input and generating queries to search logs for attack evidence using a popular query language, Kibana, used in Security Operations Centers (SOC). Our prototype implementation, AIThreatTrack, uses GPT-4 to extract actionable threat intelligence from real-time messages like X and Slack. The core idea is to explain the extracted intelligence in terms of MITRE ATT&CK TTPs using a knowledge graph containing "is-a" and "part-of" relationships (extracted using GPT-4) with the following benefits:(a) Significantly reduced hallucinations from 47% (GPT-4) to 1.5% using two orthogonal ways to cross-check answers. (b) Gaining analysts’ trust with explained results. (c) Using chain-of-knowledge prompting to significantly improve query generation accuracy. This approach supports expanding the scope of the knowledge graph to further improve query generation. Our approach significantly outperforms the Retrieval-Augmented Generation (RAG) approach and chain-of-thought reasoning LLM in reducing hallucinations. Moumita Das Purba, Bill Chu, Will French |
DSN | 1 |
| 2023 | Extracting Actionable Cyber Threat Intelligence from Twitter StreamabstractActionable cyber threat intelligence is vital for effective defense. In practice, indicators of compromises (IP addresses or domains) are used to alert potential malicious activities. However, such alerts lack important context for defenders to take effective actions. For example, given an alert concerning an IP address, a defender wants to know what type of malicious act (e.g., phishing website vs. C2) was observed associated with it. What technical manifestations (e.g., modification of a particular registry key) have been observed from the attacker using this IP address? Such context information helps defenders prioritize alerts and quickly determine whether a system has been compromised. Much of such context information exists in real-time information-sharing systems such as messaging apps and forums. This paper describes an approach to extract technical manifestations of attacks from tweets. We have compared our results with the performance of the GPT-3.5- Turbo model and text-embedding-ada-002 model of OpenAI and also created an open-source project to make our tools and data available for the cybersecurity research community [9]. Moumita Das Purba, Bill Chu |
ISI | 1 |
| 2020 | From Word Embedding to Cyber-Phrase Embedding: Comparison of Processing Cybersecurity TextsabstractMuch of the vital information about emerging threats and the corresponding defensive measures are contained in large volumes of natural language texts online. Capturing such actionable intelligence in real-time is critical to prevent large scale attacks automatically. The ATT&CK framework is a widely recognized standard to catalog technical details of cyber threats and deploy mitigating measures. A technique in ATT&CK specifies a set of adversary actions to achieve a particular goal, such as Exfiltration over Command and Control channel. Details of the technique include encrypted traffic and encoded data. A key challenge in identifying such cyber intelligence from natural language texts is that for a given action, such as encrypted traffic, many alternative expressions are possible (e.g., send using a self-signed certificate, send using HTTPS requests). It is not practical to manually provide an exhaustive list of all such variants. We demonstrate that using cyber-phrase embedding on a cybersecurity text corpus is a promising approach to overcome such difficulties. Our evaluation demonstrates that our model outperforms existing models. We have created an open-source project to make our tools and data available for the cybersecurity research community. Moumita Das Purba, Bill Chu, Ehab Al-Shaer |
ISI | 1 |