Arthur Drichel

dblp:224/2371 · DBLP profile ↗
← Back
15ranked-venue papers
11as first author
10since 2021 · last 2025
0000-0001-7326-7273ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 12 · 11 first-author · 9 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 The Persistent Threat of DGA-Domains Used by Botnets
abstract
Botnets often employ Domain Generation Algorithms (DGAs) to evade detection and maintain communication with their Command and Control (C2) servers. Despite extensive efforts to contain individual botnets and take down their C2 infrastructure, a significant knowledge gap remains regarding the extent to which their associated DGA-generated domains continue to be registered by malicious actors, posing a latent threat. In this paper, we close this gap through a comprehensive measurement study in which we quantify the threats posed by botnets, including both active botnets and those that have been subject to previous takedown operations, by analyzing the daily registered domain names included in 1165 DNS zone files, covering $80.62 \%$ of all 1445 currently valid Top-Level Domains (TLDs), over a period of $\mathbf{1 3}$ months. During our study, we observe a decade-old botnet being reactivated by new actors, allowing them to receive incoming connections from previous dormant infections and take over a number of machines. In total, we uncover malicious activities associated with 7058 domains generated by 58 different known DGAs, at least 17 of which are used by botnets that have been the target of previous takedown operations. To improve the status quo, we discuss approaches that could have prevented the malicious acts and highlight the potential of recently proposed Machine Learning (ML) techniques to uncover yet unknown DGAs, enabling a more proactive approach to threat detection.
Arthur Drichel, Ulrike Meyer
RAID1
2024 Towards Robust Domain Generation Algorithm Classification
abstract
In this work, we conduct a comprehensive study on the robustness of domain generation algorithm (DGA) classifiers. We implement 32 white-box attacks, 19 of which are very effective and induce a false-negative rate (FNR) of ≈ 100% on unhardened classifiers. To defend the classifiers, we evaluate different hardening approaches and propose a novel training scheme that leverages adversarial latent space vectors and discretized adversarial domains to significantly improve robustness. In our study, we highlight a pitfall to avoid when hardening classifiers and uncover training biases that can be easily exploited by attackers to bypass detection, but which can be mitigated by adversarial training (AT). In our study, we do not observe any trade-off between robustness and performance, on the contrary, hardening improves a classifier's detection performance for known and unknown DGAs. We implement all attacks and defenses discussed in this paper as a standalone library, which we make publicly available1 to facilitate hardening of DGA classifiers.
Arthur Drichel, Marc Meyer, Ulrike Meyer
AsiaCCS1
2024 Extended Abstract: A Transfer Learning-Based Training Approach for DGA Classification
Arthur Drichel, Benedikt von Querfurth, Ulrike Meyer
DIMVA1
2024 A Comprehensive Study on Multi-Task Learning for Domain Generation Algorithm (DGA) Detection
abstract
In this work, we perform a comparative evaluation of 21 approaches to multi-task learning (MTL) for the detection of domain generation algorithms (DGAs). To this end, we train and evaluate 2300 classifiers using a combination of 14 different optimization strategies and 6 MTL architectures and compare them statistically with the state of the art. In this context, we propose a novel ResNet backbone, which already surpasses the state of the art on its own, but shines especially in combination with MTL. We evaluate the novel DGA classifiers in a real-world study that avoids temporal and spatial experimental biases to assess whether they generalize well between different networks and are robust over time. Moreover, we analyze the classifiers' capability to detect yet unknown DGAs and discuss their practical application. Our best-performing classifier surpasses the state of the art by over 5.7% in area under the curve (AUC) for practically relevant false-positive rates (FPRs) and exceeds the state of the art by over 7.3% in true-positive rate (TPR) at the same fixed FPR of 0.001 in a real-world setting.
Arthur Drichel, Ulrike Meyer
PST1
2023 False Sense of Security: Leveraging XAI to Analyze the Reasoning and True Performance of Context-less DGA Classifiers
abstract
The problem of revealing botnet activity through Domain Generation Algorithm (DGA) detection seems to be solved, considering that available deep learning classifiers achieve accuracies of over 99.9%. However, these classifiers provide a false sense of security as they are heavily biased and allow for trivial detection bypass. In this work, we leverage explainable artificial intelligence (XAI) methods to analyze the reasoning of deep learning classifiers and to systematically reveal such biases. We show that eliminating these biases from DGA classifiers considerably deteriorates their performance. Nevertheless we are able to design a context-aware detection system that is free of the identified biases and maintains the detection rate of state-of-the art deep learning classifiers. In this context, we propose a visual analysis system that helps to better understand a classifier’s reasoning, thereby increasing trust in and transparency of detection methods and facilitating decision-making.
Arthur Drichel, Ulrike Meyer
RAID1
2022 Detecting Unknown DGAs without Context Information
abstract
New malware emerges at a rapid pace and often incorporates Domain Generation Algorithms (DGAs) to avoid blocking the malware’s connection to the command and control (C2) server. Current state-of-the-art classifiers are able to separate benign from malicious domains (binary classification) and attribute them with high probability to the DGAs that generated them (multiclass classification). While binary classifiers can label domains of yet unknown DGAs as malicious, multiclass classifiers can only assign domains to DGAs that are known at the time of training, limiting the ability to uncover new malware families. In this work, we perform a comprehensive study on the detection of new DGAs, which includes an evaluation of 59,690 classifiers. We examine four different approaches in 15 different configurations and propose a simple yet effective approach based on the combination of a softmax classifier and regular expressions (regexes) to detect multiple unknown DGAs with high probability. At the same time, our approach retains state-of-the-art classification performance for known DGAs. Our evaluation is based on a leave-one-group-out cross-validation with a total of 94 DGA families. By using the maximum number of known DGAs, our evaluation scenario is particularly difficult and close to the real world. All of the approaches examined are privacy-preserving, since they operate without context and exclusively on a single domain to be classified. We round up our study with a thorough discussion of class-incremental learning strategies that can adapt an existing classifier to newly discovered classes.
Arthur Drichel, Justus von Brandt, Ulrike Meyer
ARES1
2021 Finding Phish in a Haystack: A Pipeline for Phishing Classification on Certificate Transparency Logs
abstract
Current popular phishing prevention techniques mainly utilize reactive blocklists, which leave a “window of opportunity” for attackers during which victims are unprotected. One possible approach to shorten this window aims to detect phishing attacks earlier, during website preparation, by monitoring Certificate Transparency (CT) logs. Previous attempts to work with CT log data for phishing classification exist, however they lack evaluations on actual CT log data. In this paper, we present a pipeline that facilitates such evaluations by addressing a number of problems when working with CT log data. The pipeline includes dataset creation, training, and past or live classification of CT logs. Its modular structure makes it possible to easily exchange classifiers or verification sources to support ground truth labeling efforts and classifier comparisons. We test the pipeline on a number of new and existing classifiers, and find a general potential to improve classifiers for this scenario in the future. We publish the source code of the pipeline and the used datasets along with this paper [12], thus making future research in this direction more accessible.
Arthur Drichel, Vincent Drury, Justus von Brandt, Ulrike Meyer
ARES1
2021 First Step Towards EXPLAINable DGA Multiclass Classification
abstract
Numerous malware families rely on domain generation algorithms (DGAs) to establish a connection to their command and control (C2) server. Counteracting DGAs, several machine learning classifiers have been proposed enabling the identification of the DGA that generated a specific domain name and thus triggering targeted remediation measures. However, the proposed state-of-the-art classifiers are based on deep learning models. The black box nature of these makes it difficult to evaluate their reasoning. The resulting lack of confidence makes the utilization of such models impracticable. In this paper, we propose EXPLAIN, a feature-based and contextless DGA multiclass classifier. We comparatively evaluate several combinations of feature sets and hyperparameters for our approach against several state-of-the-art classifiers in a unified setting on the same real-world data. Our classifier achieves competitive results, is real-time capable, and its predictions are easier to trace back to features than the predictions made by the DGA multiclass classifiers proposed in related work.
Arthur Drichel, Nils Faerber, Ulrike Meyer
ARES1
2021 Towards Privacy-Preserving Classification-as-a-Service for DGA Detection
abstract
Domain generation algorithm (DGA) classifiers can be used to detect and block the establishment of a connection between bots and their command-and-control server. Classification-as-a-service (CaaS) can separate the classification of domain names from the need for real-world training data, which are difficult to obtain but mandatory for well performing classifiers. However, domain names as well as trained models may contain privacy-critical information which should not be leaked to either the model provider or the data provider. Several generic frameworks for privacy-preserving machine learning (ML) have been proposed in the past that can preserve data and model privacy. Thus, it seems high time to combine state-of-the-art DGA classifiers and privacy-preservation frameworks to enable privacy-preserving CaaS, preserving both, data and model privacy for the DGA detection use case. In this work, we examine the real-world applicability of four generic frameworks for privacy-preserving ML using different state-of-the-art DGA detection models. Our results show that out-of-the-box DGA detection models are computationally infeasible for privacy-preserving inference in a real-world setting. We propose model simplifications that achieve a reduction in inference latency of up to 95%, and up to 97% in communication complexity while causing an accuracy penalty of less than 0.17%. Despite this significant improvement, real-time classification is still not feasible in a traditional two-party setting. Thus, more efficient secure multi-party computation (SMPC) or homomorphic encryption (HE) schemes are required to enable real-world feasibility of privacy-preserving CaaS for DGA detection.
Arthur Drichel, Mehdi Akbari Gurabi, Tim Amelung, Ulrike Meyer
PST1
2021 CoinPrune: Shrinking Bitcoin's Blockchain Retrospectively
abstract
Popular cryptocurrencies continue to face serious scalability issues due to their ever-growing blockchains. Thus, modern blockchain designs began to prune old blocks and rely on recent snapshots for their bootstrapping processes instead. Unfortunately, established systems are often considered incapable of adopting these improvements. In this work, we present CoinPrune, our block-pruning scheme with full Bitcoin compatibility, to revise this popular belief. CoinPrune bootstraps joining nodes via snapshots that are periodically created from Bitcoin's set of unspent transaction outputs (UTXO set). Our scheme establishes trust in these snapshots by relying on CoinPrunesupporting miners to mutually reaffirm a snapshot's correctness on the blockchain. This way, snapshots remain trustworthy even if adversaries attempt to tamper with them. Our scheme maintains its retrospective deployability by relying on positive feedback only, i.e., blocks containing invalid reaffirmations are not rejected, but invalid reaffirmations are outpaced by the benign ones created by an honest majority among CoinPrunesupporting miners. Already today, CoinPrune reduces the storage requirements for Bitcoin nodes by two orders of magnitude, as joining nodes need to fetch and process only 6 GiB instead of 271 GiB of data in our evaluation, reducing the synchronization time of powerful devices from currently 7 h to 51 min, with even larger potential drops for less powerful devices. CoinPrune is further aware of higher-level application data, i.e., it conserves otherwise pruned application data and allows nodes to obfuscate objectionable and potentially illegal blockchain content from their UTXO set and the snapshots they distribute.
Roman Matzutt, Benedikt Kalde, Jan Pennekamp, Arthur Drichel, Martin Henze, Klaus Wehrle
IEEE Trans. Netw. Serv. Manag.4
2020 Analyzing the real-world applicability of DGA classifiers
abstract
Separating benign domains from domains generated by DGAs with the help of a binary classifier is a well-studied problem for which promising performance results have been published. The corresponding multiclass task of determining the exact DGA that generated a domain enabling targeted remediation measures is less well studied. Selecting the most promising classifier for these tasks in practice raises a number of questions that have not been addressed in prior work so far. These include the questions on which traffic to train in which network and when, just as well as how to assess robustness against adversarial attacks. Moreover, it is unclear which features lead a classifier to a decision and whether the classifiers are real-time capable. In this paper, we address these issues and thus contribute to bringing DGA detection classifiers closer to practical use. In this context, we propose one novel classifier based on residual neural networks for each of the two tasks and extensively evaluate them as well as previously proposed classifiers in a unified setting. We not only evaluate their classification performance but also compare them with respect to explainability, robustness, and training and classification speed. Finally, we show that our newly proposed binary classifier generalizes well to other networks, is time-robust, and able to identify previously unknown DGAs.
Arthur Drichel, Ulrike Meyer, Samuel Schüppen, Dominik Teubert
ARES1
2020 Making use of NXt to nothing: the effect of class imbalances on DGA detection classifiers
abstract
Numerous machine learning classifiers have been proposed for binary classification of domain names as either benign or malicious, and even for multiclass classification to identify the domain generation algorithm (DGA) that generated a specific domain name. Both classification tasks have to deal with the class imbalance problem of strongly varying amounts of training samples per DGA. Currently, it is unclear whether the inclusion of DGAs for which only a few samples are known to the training sets is beneficial or harmful to the overall performance of the classifiers. In this paper, we perform a comprehensive analysis of various contextless DGA classifiers, which reveals the high value of a few training samples per class for both classification tasks. We demonstrate that the classifiers are able to detect various DGAs with high probability by including the underrepresented classes which were previously hardly recognizable. Simultaneously, we show that the classifiers' detection capabilities of well represented classes do not decrease.
Arthur Drichel, Ulrike Meyer, Samuel Schüppen, Dominik Teubert
ARES1
2020 How to Securely Prune Bitcoin's Blockchain
Roman Matzutt, Benedikt Kalde, Jan Pennekamp, Arthur Drichel, Martin Henze, Klaus Wehrle
Networking4
2020 Interpretable Visualizations of Deep Neural Networks for Domain Generation Algorithm Detection
abstract
Due to their success in many application areas, deep learning models have found wide adoption for many problems. However, their black-box nature makes it hard to trust their decisions and to evaluate their line of reasoning. In the field of cybersecurity, this lack of trust and understanding poses a significant challenge for the utilization of deep learning models. Thus, we present a visual analytics system that provides designers of deep learning models for the classification of domain generation algorithms with understandable interpretations of their model. We cluster the activations of the model's nodes and leverage decision trees to explain these clusters. In combination with a 2D projection, the user can explore how the model views the data at different layers. In a preliminary evaluation of our system, we show how it can be employed to better understand misclassifications, identify potential biases and reason about the role different layers in a model may play.
Franziska Becker, Arthur Drichel, Christoph Müller 0001, Thomas Ertl
VizSec2
2017 CloudAnalyzer: Uncovering the Cloud Usage of Mobile Apps
abstract
Developers of smartphone apps increasingly rely on cloud services for ready-made functionalities, e.g., to track app usage, to store data, or to integrate social networks. At the same time, mobile apps have access to various private information, ranging from users' contact lists to their precise locations. As a result, app deployment models and data flows have become too complex and entangled for users to understand. We present CloudAnalyzer, a transparency technology that reveals the cloud usage of smartphone apps and hence provides users with the means to reclaim informational self-determination. We apply CloudAnalyzer to study the cloud exposure of 29 volunteers over the course of 19 days. In addition, we analyze the cloud usage of the 5000 most accessed mobile websites as well as 500 popular apps from five different countries. Our results reveal an excessive exposure to cloud services: 90 % of apps use cloud services and 36 % of apps used by volunteers solely communicate with cloud services. Given the information provided by CloudAnalyzer, users can critically review the cloud usage of their apps.
Martin Henze, Jan Pennekamp, David Hellmanns, Erik Mühmer, Jan Henrik Ziegeldorf, Arthur Drichel, Klaus Wehrle
MobiQuitous6