EDBT 2026 Demo / reviewers in the wild / expert
Dana Warmsley
dblp:198/0515
· DBLP profile ↗
8ranked-venue papers in the field
3as first author
3since 2021 · last 2023
—ORCID · none
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 7 (3 first)Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Learning Explainable Multi-view Representations for Malware Authorship AttributionabstractMalware poses an ever-growing threat to organizations, governments, and institutions. To effectively combat its proliferation and development by threat actors, identifying the authors of malware is a crucial step. Existing efforts in malware authorship attribution often focus on extracting stylistic features from source code. However, malware source code is frequently limited in availability. There is a strong need to develop effective methods to directly extract a robust and salient representation from malware binaries. This representation should adeptly capture the inherent characteristics of the binaries to facilitate the identification of malware authors. In our work, we introduce an approach that leverages multi-view graph representation learning for malware authorship attribution. We extract various intermediate representations from the binary, including control flow graphs and function call graphs. These representations are encoded using GNN-based functions and fused to create a comprehensive malware representation. Additionally, we introduce GNN-explainability approaches to identify subgraphs indicative of class membership. Our experiments on a recent benchmark demonstrate significant performance enhancements in authorship attribution. Irsyad Adam, Alex Waagen, Dana Warmsley, Jiejun Xu |
IEEE Big Data | 3 |
| 2022 | A Survey of Explainable Graph Neural Networks for Cyber Malware AnalysisabstractMalicious cybersecurity activities have become increasingly worrisome for individuals and companies alike. While machine learning methods like Graph Neural Networks (GNNs) have proven successful on the malware detection task, their output is often difficult to understand. Explainable malware detection methods are needed to automatically identify malicious programs and present results to malware analysts in a way that is human interpretable. In this survey, we outline a number of GNN explainability methods and compare their performance on a real-world malware detection dataset. Specifically, we formulated the detection problem as a graph classification problem on the malware Control Flow Graphs (CFGs). We find that gradient-based methods outperform perturbation-based methods in terms of computational expense and performance on explainer-specific metrics (e.g., Fidelity and Sparsity). Our results provide insights into designing new GNN-based models for cyber malware detection and attribution. Dana Warmsley, Alex Waagen, Jiejun Xu, Zhining Liu 0002, Hanghang Tong |
IEEE Big Data | 1 |
| 2021 | Influence in Transient PopulationsabstractIn this work we explore a simple influence problem using a novel variation of the Bounded Confidence Model (BCM) of opinion dynamics that can be applied to varying types of social media platforms with transient user populations. Via simulation-based analyses, we set out to understand the extent to which the opinions of a transient population that divides their attention across multiple channels is either more or less susceptible to progressively more extreme opinions. Dana Warmsley, Samuel D. Johnson |
IEEE BigData | 1 |
| 2020 | Causal Maps for Multi-Document SummarizationabstractConcept maps are concise graphical representations of text data which have been shown to be applicable as a tool for text summarization and downstream tasks. Most prior work has either focused on the generation of concept maps for small corpora or require expensive training data to implement. In this work, we focus on generating causal maps, a subset of concept maps in which only semantically causal relationships are considered. We propose a map generation framework which utilizes a novel mixture model to simultaneously derive concepts and links. This method is computationally efficient and therefore scalable to large datasets, and is fully unsupervised, which makes it suitable for practical applications. We show that our method performs better than a commonly used unsupervised text summarization algorithm, and has results which are comparable to the state-of-the-art supervised method. Sasha Strelnikoff, Aruna Jammalamadaka, Dana Warmsley |
IEEE BigData | 3 |
| 2019 | MATRICS: A System for Human-Machine Hybrid Forecasting of Geopolitical EventsabstractIn this paper, we present MATRICS, a humanmachine hybrid system that accurately performs geopolitical forecasting by combining crowdsourcing with ensemble machine learning on online data. The system employs a pair of parallel, but highly-interconnected processing pipelines to perform “machine-aided human forecasting” and “human-aided machine forecasting”. This configuration allows the machine to provide information to the human population, saving research time and reducing fatigue, while simultaneously allowing the human population to provide feedback to the machine learning components, allowing them to filter data sources and quickly adapt to a task via online machine learning. The final forecast for each question was computed as an aggregate of the human and machine responses. The system was evaluated using data collected during the IARPA Hybrid Forecasting Competition, in which it answered 187 forecasting questions with a mean Brier score of 0.27 using volunteers and participants that were recruited via Amazon Mechanical Turk and open-source “big” data scraped from online sources such as social media, search engine results, and online historical data. David Huber 0004, Samuel D. Johnson, Nigel Stepp, Aruna Jammalamadaka, Dana Warmsley, Tiffany Kim, Tsai-Ching Lu |
IEEE BigData | 5 |
| 2019 | MATRICS: A System for Human-Machine Hybrid Forecasting of Geopolitical EventsabstractIn this paper, we present MATRICS, a human-machine hybrid system that accurately performs geopolitical forecasting by combining crowdsourcing with ensemble machine learning on online data. The system employs a pair of parallel, but highly-interconnected processing pipelines to perform “machine-aided human forecasting” and “human-aided machine forecasting”. This configuration allows the machine to provide information to the human population, saving research time and reducing fatigue, while simultaneously allowing the human population to provide feedback to the machine learning components, allowing them to filter data sources and quickly adapt to a task via online machine learning. The final forecast for each question was computed as an aggregate of the human and machine responses. The system was evaluated using data collected during the IARPA Hybrid Forecasting Competition, in which it answered 187 forecasting questions with a mean Brier score of 0.27 using volunteers and participants that were recruited via Amazon Mechanical Turk and open-source “big” data scraped from online sources such as social media, search engine results, and online historical data. David Huber 0004, Nigel Stepp, Aruna Jammalamadaka, Tiffany Kim, Sam Johnson, Dana Warmsley, Tsai-Ching Lu |
IEEE BigData | 6 |
| 2018 | From Gamergate to FIFA: Identifying Polarized Groups in Online Social MediaabstractPolarizing topics and events are often widely discussed and debated in social media, allowing researchers a unique view into the minds of large populations on everything from politics to entertainment. Previous work in identifying polarization in social media has largely used traditional community detection methods that are often confounded by the existence of neutral users and content. To address this problem, we propose a novel approach based on nonnegative matrix factorization (NMF) on a tripartite network to illuminate latent polarization patterns. The proposed methods are designed to work in contexts varying in the nature of the controversy, the level of polarization, the number of polarity groups involved and the presence of neutral entities. We use real-world Tumblr datasets to show that our algorithm exhibits superior performance in identifying polarization in online communities with respect to a range of real-world topics. To the best of our knowledge, our work is the first attempt to analyze polarization on the Tumblr platform. Dana Warmsley, Jiejun Xu, Tsai-Ching Lu |
IEEE BigData | 1 |
| 2017 | Automated Hate Speech Detection and the Problem of Offensive Language
Thomas Davidson, Dana Warmsley, Michael W. Macy, Ingmar Weber |
ICWSM | 2 |