VLDB 2026 Research / reviewers in the wild / expert
Hyrum S. Anderson
dblp:05/8364
· DBLP profile ↗
13ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0009-4720-6907ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EMBER2024 - A Benchmark Dataset for Holistic Evaluation of Malware ClassifiersabstractA lack of accessible data has historically restricted malware analysis research, and practitioners have relied heavily on datasets provided by industry sources to advance. Existing public datasets are limited by narrow scope - most include files targeting a single platform, have labels supporting just one type of malware classification task, and make no effort to capture the evasive files that make malware detection difficult in practice. We present EMBER2024, a new dataset that enables holistic evaluation of malware classifiers. Created in collaboration with the authors of EMBER2017 and EMBER2018, the EMBER2024 dataset includes hashes, metadata, feature vectors, and labels for more than 3.2 million files from six file formats. Our dataset supports the training and evaluation of machine learning models on seven malware classification tasks, including malware detection, malware family classification, and malware behavior identification. EMBER2024 is the first to include a collection of malicious files that initially went undetected by a set of antivirus products, creating a ''challenge'' set to assess classifier performance against evasive malware. This work also introduces EMBER feature version 3, with added support for several new feature types. We are releasing the EMBER2024 dataset to promote reproducibility and empower researchers in the pursuit of new malware research topics. Robert J. Joyce, Gideon Miller, Phil Roth 0002, Richard Zak, Elliott Zaresky-Williams, Hyrum S. Anderson, Edward Raff, James Holt |
KDD (2) | 6 |
| 2024 | Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyabstractWhile Large Language Models (LLMs) display versatile functionality, they continue to generate harmful, biased, and toxic content, as demonstrated by the prevalence of human-designed *jailbreaks*. In this work, we present *Tree of Attacks with Pruning* (TAP), an automated method for generating jailbreaks that only requires black-box access to the target LLM. TAP utilizes an attacker LLM to iteratively refine candidate (attack) prompts until one of the refined prompts jailbreaks the target. In addition, before sending prompts to the target, TAP assesses them and prunes the ones unlikely to result in jailbreaks, reducing the number of queries sent to the target LLM. In empirical evaluations, we observe that TAP generates prompts that jailbreak state-of-the-art LLMs (including GPT4-Turbo and GPT4o) for more than 80% of the prompts. This significantly improves upon the previous state-of-the-art black-box methods for generating jailbreaks while using a smaller number of queries than them. Furthermore, TAP is also capable of jailbreaking LLMs protected by state-of-the-art *guardrails*, e.g., LlamaGuard. Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum S. Anderson, Yaron Singer, Amin Karbasi |
NeurIPS | 5 |
| 2024 | Poisoning Web-Scale Training Datasets is PracticalabstractDeep learning models are often trained on distributed, web-scale datasets crawled from the internet. In this paper, we introduce two new dataset poisoning attacks that intentionally introduce malicious examples to a model’s performance. Our attacks are immediately practical and could, today, poison 10 popular datasets. Our first attack, split-view poisoning, exploits the mutable nature of internet content to ensure a dataset annotator’s initial view of the dataset differs from the view downloaded by subsequent clients. By exploiting specific invalid trust assumptions, we show how we could have poisoned 0.01% of the LAION-400M or COYO-700M datasets for just $60 USD. Our second attack, frontrunning poisoning, targets web-scale datasets that periodically snapshot crowd-sourced content—such as Wikipedia—where an attacker only needs a time-limited window to inject malicious examples. In light of both attacks, we notify the maintainers of each affected dataset and recommended several low-overhead defenses. Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum S. Anderson, Andreas Terzis, Kurt Thomas, Florian Tramèr |
SP | 6 |
| 2024 | Metadata-Based Detection of Child Sexual Abuse MaterialabstractChild Sexual Abuse Media (CSAM) is any visual record of a sexually explicit activity involving minors. Machine learning-based solutions can help law enforcement identify CSAM and block distribution. Yet, collecting CSAM imagery to train machine learning models has ethical and legal constraints. CSAM detection systems based on file metadata offer several opportunities. Metadata is not a record of a crime and, therefore, clear of legal restrictions. This paper proposes a CSAM detection framework consisting of machine learning models trained on file paths extracted from a real-world data set of over 1 million file paths obtained in criminal investigations. Our framework includes guidelines for model evaluation that account for data changes caused by adversarial data modification and variations in data distribution caused by limited access to training data, as well as an assessment of false positive rates against file paths from common crawl data. We achieve accuracies as high as 0.97 while presenting stable behavior under adversarial attacks previously used in natural language tasks. When evaluating the model on publicly available file paths from common crawl data, we observed a false positive rate of 0.002, showing that the model operating in distinct data distributions maintains low false positive rates. Mayana Pereira, Rahul Dodhia, Hyrum S. Anderson |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2021 | Classifying Sequences of Extreme Length with Constant Memory Applied to Malware DetectionabstractRecent works within machine learning have been tackling inputs of ever increasing size, with cyber security presenting sequence classification problems of particularly extreme lengths. In the case of Windows executable malware detection, an input executable could be >=100 MB, which would translate to a time series with T=100,000,000 steps. To date, the closest approach to handling such task is MalConv --- a convolutional neural network capable of processing T=2,000,000 steps. Because the memory used by CNNs is O(T), this has prevented many from processing all executables or further extending the MalConv approach. In this work, we develop a new approach to temporal max pooling that makes the required memory invariant to the sequence length T. This makes MalConv 116x more memory efficient, and up to 25.8x faster to train, while removing the input length restrictions to MalConv. We re-invest these gains into improving the MalConv architecture by developing a new Global Channel Gating design, giving us an attention mechanism capable of learning feature interactions across 100 million time steps in an efficient manner, a capability lacked by the original MalConv approach. Edward Raff, William Fleshman, Richard Zak, Hyrum S. Anderson, Bobby Filar, Mark McLean |
AAAI | 4 |
| 2014 | Sparse reconstruction of equivalence classes of moving targets using single-channel synthetic aperture radarabstractSimultaneously estimating position and velocity of moving targets using only phase information from single-channel SAR data is impossible. This paper defines classes of equivalent target motion and solves the GMTI problem up to membership in an equivalence class using single-channel SAR phase data. We present a definitions for endo- and exo-clutter that is consistent with the equivalence classes, and show that most target motion can be detected, i.e. the set of endo-clutter targets is very small. We exploit the sparsity of moving targets in the scene to develop an algorithm to resolve target motion up to membership in an equivalence class, and demonstrate the effectiveness of the proposed technique using simulated data. Jacob H. Gunther, Josh Hunsaker, Hyrum S. Anderson, Todd K. Moon |
ICASSP | 3 |
| 2013 | Classifying with confidence from incomplete information
Nathan Parrish, Hyrum S. Anderson, Maya R. Gupta, Dun-Yu Hsiao |
J. Mach. Learn. Res. | 2 |
| 2012 | Reliable early classification of time seriesabstractEarly classification of time series is important in time-sensitive applications. An approach is presented for early classification using generative classifiers with the dual objectives of providing a class label as early as possible while guaranteeing with high probability that the early class matches the class that would be assigned to a longer time series. We give a specific algorithm for early quadratic discriminant analysis (QDA), and demonstrate that this classifier meets the requirement of reliable early classification. Hyrum S. Anderson, Nathan Parrish, Kristi Tsukida, Maya R. Gupta |
ICASSP | 1 |
| 2010 | Robust sequential classification of tracks
Nathan Parrish, Hyrum S. Anderson, Maya R. Gupta |
FUSION | 2 |
| 2010 | Training a support vector machine to classify signals in a real environment given clean training dataabstractWhen building a classifier from clean training data for a particular test environment, knowledge about the environmental noise and channel should be taken into account. We propose training a support vector machine (SVM) classifier using a modified kernel that is the expected kernel with respect to a probability distribution over channels and noise that might affect the test signal. We compare the proposed expected SVM to an SVM that ignores the environment, to an SVM that trains with multiple random samples of the environment, and to a quadratic discriminant analysis classifier that takes advantage of environment statistics (Joint QDA). Simulations classifying narrowband signals in a noisy acoustic reverberation environment indicate that the expected SVM can improve performance over a range of noise levels. Kevin Jamieson 0001, Maya R. Gupta, Eric Swanson, Hyrum S. Anderson |
ICASSP | 4 |
| 2007 | Joint Deconvolution and Classification for Signals with MultipathabstractFor many sensing modalities such as sonar, received signals are corrupted by multipath and can be challenging for automatic classification systems. An approach to jointly deconvolve and classify such signals is proposed. Specifically, a filter is estimated that minimizes the distortion between the received signal and a set of training signals, then the received signal is assigned to the class that corresponds to the training signal whose estimated filter is most sparse. Simulations compare the new method with blind deconvolution using Cabrelli's algorithm followed by a correlation-based nearest neighbor classifier. Results indicate that joint deconvolution and classification performs similarly to blind deconvolution in the presence of severe noise, and outperforms blind deconvolution at low and moderate noise levels. Maya R. Gupta, Hyrum S. Anderson, Yihua Chen 0003 |
ICASSP (3) | 2 |
| 2005 | Sea ice mapping method for SeaWindsabstractA sea ice mapping algorithm for SeaWinds is developed that incorporates statistical and spatial a priori information in a modified maximum a posteriori (MAP) framework. Spatial a priori data are incorporated in the loss terms of a Bayes risk formulation. Conditional distributions and priors for sea ice and ocean statistics are represented as empirical histograms that are forced to conform to a set of expected histograms via principal component filtering. Tuning parameters for the algorithm allow adjustments in the algorithm's performance. Results of the algorithm exhibit high correlation with the Remund-Long sea ice mapping algorithm for SeaWinds and the Special Sensor Microwave/Imager National Aeronautics and Space Administration Team 30% ice edge, and are verified with RADARSAT-1 ScanSAR imagery. The resulting sea ice maps exhibit high edge detail, preserve polynyas and ice bodies disjoint from the primary ice sheet, and thus are suitable for use with wind retrieval and sea ice studies. Principles employed in the algorithm may be of interest in other classification studies. Hyrum S. Anderson, David G. Long |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2003 | Polar sea ice mapping using SeaWinds dataabstractMicrowave remote sensing provides an excellent means for mapping polar ice extent. In this study, a new algorithm for polar sea ice mapping is developed for use with the SeaWinds instrument. The approach utilizes a priori information within the framework of Bayes detection to produce sea ice extent maps. Statistical models for sea ice and ocean are represented in histograms which are filtered using a principal component (PC) based filtering technique. Spatial a priori information is incorporated through the loss terms associated with Bayes risk. Sea ice extent maps produced by the algorithm correlate well with the Remund-Long algorithm. Hyrum S. Anderson, David G. Long |
IGARSS | 1 |