VLDB 2026 Research / reviewers in the wild / expert
Richard Zak
dblp:215/0176
· DBLP profile ↗
4ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0003-4272-2565ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
3 papers |
Malware analysis · 58% Systems and software security · 42% | |
| Software engineering, system software, and programming languages
1 paper |
Software testing · 100% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Malware analysis
malware detection |
1.4 | 2 | 2025 | EMBER2024 - A Benchmark Dataset for Holistic Evaluation of Malware Classifiers · KDD (2) 2025 Classifying Sequences of Extreme Length with Constant Memory Applied to Malware Detection · AAAI 2021 |
Malware analysis › malware detection evasion
evasive malware |
0.9 | 1 | 2025 | EMBER2024 - A Benchmark Dataset for Holistic Evaluation of Malware Classifiers · KDD (2) 2025 |
Malware analysis
malware classification |
0.9 | 1 | 2025 | EMBER2024 - A Benchmark Dataset for Holistic Evaluation of Malware Classifiers · KDD (2) 2025 |
Systems and software security
binary analysis |
0.8 | 1 | 2024 | Is Function Similarity Over-Engineered? Building a Benchmark · NeurIPS 2024 |
Systems and software security › binary analysis
binary function similarity |
0.8 | 1 | 2024 | Is Function Similarity Over-Engineered? Building a Benchmark · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
efficient training |
0.1 | 1 | 2021 | Classifying Sequences of Extreme Length with Constant Memory Applied to Malware Detection · AAAI 2021 |
Methods — techniques the papers use, named apart from their topics
machine learning · 2.4disassembly · 1.5decompilation · 1.5temporal max pooling · 1.0convolutional neural network · 1.0attention · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EMBER2024 - A Benchmark Dataset for Holistic Evaluation of Malware ClassifiersabstractA lack of accessible data has historically restricted malware analysis research, and practitioners have relied heavily on datasets provided by industry sources to advance. Existing public datasets are limited by narrow scope - most include files targeting a single platform, have labels supporting just one type of malware classification task, and make no effort to capture the evasive files that make malware detection difficult in practice. We present EMBER2024, a new dataset that enables holistic evaluation of malware classifiers. Created in collaboration with the authors of EMBER2017 and EMBER2018, the EMBER2024 dataset includes hashes, metadata, feature vectors, and labels for more than 3.2 million files from six file formats. Our dataset supports the training and evaluation of machine learning models on seven malware classification tasks, including malware detection, malware family classification, and malware behavior identification. EMBER2024 is the first to include a collection of malicious files that initially went undetected by a set of antivirus products, creating a ''challenge'' set to assess classifier performance against evasive malware. This work also introduces EMBER feature version 3, with added support for several new feature types. We are releasing the EMBER2024 dataset to promote reproducibility and empower researchers in the pursuit of new malware research topics. Robert J. Joyce, Gideon Miller, Phil Roth 0002, Richard Zak, Elliott Zaresky-Williams, Hyrum S. Anderson, Edward Raff, James Holt |
KDD (2) | 4 |
| 2024 | Is Function Similarity Over-Engineered? Building a BenchmarkabstractBinary analysis is a core component of many critical security tasks, including reverse engineering, malware analysis, and vulnerability detection. Manual analysis is often time-consuming, but identifying commonly-used or previously-seen functions can reduce the time it takes to understand a new file. However, given the complexity of assembly, and the NP-hard nature of determining function equivalence, this task is extremely difficult. Common approaches often use sophisticated disassembly and decompilation tools, graph analysis, and other expensive pre-processing steps to perform function similarity searches over some corpus. In this work, we identify a number of discrepancies between the current research environment and the underlying application need. To remedy this, we build a new benchmark, REFuSe-Bench, for binary function similarity detection consisting of high-quality datasets and tests that better reflect real-world use cases. In doing so, we address issues like data duplication and accurate labeling, experiment with real malware, and perform the first serious evaluation of ML binary function similarity models on Windows data. Our benchmark reveals that a new, simple baseline — one which looks at only the raw bytes of a function, and requires no disassembly or other pre-processing --- is able to achieve state-of-the-art performance in multiple settings. Our findings challenge conventional assumptions that complex models with highly-engineered features are being used to their full potential, and demonstrate that simpler approaches can provide significant value. Rebecca Saul, Chang Liu 0188, Noah Fleischmann, Richard Zak, Kristopher K. Micinski, Edward Raff, James Holt |
NeurIPS | 4 |
| 2021 | Classifying Sequences of Extreme Length with Constant Memory Applied to Malware DetectionabstractRecent works within machine learning have been tackling inputs of ever increasing size, with cyber security presenting sequence classification problems of particularly extreme lengths. In the case of Windows executable malware detection, an input executable could be >=100 MB, which would translate to a time series with T=100,000,000 steps. To date, the closest approach to handling such task is MalConv --- a convolutional neural network capable of processing T=2,000,000 steps. Because the memory used by CNNs is O(T), this has prevented many from processing all executables or further extending the MalConv approach. In this work, we develop a new approach to temporal max pooling that makes the required memory invariant to the sequence length T. This makes MalConv 116x more memory efficient, and up to 25.8x faster to train, while removing the input length restrictions to MalConv. We re-invest these gains into improving the MalConv architecture by developing a new Global Channel Gating design, giving us an attention mechanism capable of learning feature interactions across 100 million time steps in an efficient manner, a capability lacked by the original MalConv approach. Edward Raff, William Fleshman, Richard Zak, Hyrum S. Anderson, Bobby Filar, Mark McLean |
AAAI | 3 |
| 2019 | RelExt: relation extraction using deep learning approaches for cybersecurity knowledge graph improvementabstractSecurity Analysts that work in a 'Security Operations Center' (SoC) play a major role in ensuring the security of the organization. The amount of background knowledge they have about the evolving and new attacks makes a significant difference in their ability to detect attacks. Open source threat intelligence sources, like text descriptions about cyber-attacks, can be stored in a structured fashion in a cybersecurity knowledge graph. A cybersecurity knowledge graph can be paramount in aiding a security analyst to detect cyber threats because it stores a vast range of cyber threat information in the form of semantic triples which can be queried. A semantic triple contains two cybersecurity entities with a relationship between them. In this work, we propose a system to create semantic triples over cybersecurity text, using deep learning approaches to extract possible relationships. We use the set of semantic triples generated through our system to assert in a cybersecurity knowledge graph. Security Analysts can retrieve this data from the knowledge graph, and use this information to form a decision about a cyber-attack. Aditya Pingle, Aritran Piplai, Sudip Mittal, Anupam Joshi, James Holt, Richard Zak |
ASONAM | 6 |