Benjamin Feuer

dblp:322/5063 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-7938-542XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 50% Image recognition and object detection · 16% Efficient and distributed learning · 10%
Network and information security
1 paper
Digital forensics and information hiding · 100%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%

Topics — the 26 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
0.912025
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking · ICLR 2025
Natural language and speech › Language models and text generation › large language model evaluation
benchmark contamination
0.912025
LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025
Natural language and speech › Language models and text generation › large language model evaluation
LLM judge
0.912025
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking · ICLR 2025
Natural language and speech › Language models and text generation › large language model training
post-training
0.912025
WildChat-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training · ICML 2025
Natural language and speech › Language models and text generation
preference optimization
0.912025
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking · ICLR 2025
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.912025
WildChat-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training · ICML 2025
Natural language and speech › Language models and text generation
synthetic data
0.912025
WildChat-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training · ICML 2025
Digital forensics and information hiding › watermarking
image watermarking
0.912025
Hidden in the Noise: Two-Stage Robust Watermarking for Images · ICLR 2025
Digital forensics and information hiding
watermarking
0.912025
Hidden in the Noise: Two-Stage Robust Watermarking for Images · ICLR 2025
Machine learning › Efficient and distributed learning
data curation
0.812024
SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification · NeurIPS 2024
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.812024
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024
Computer vision › Image recognition and object detection
image classification
0.812024
SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification · NeurIPS 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks · NeurIPS 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › neural processes
prior-data fitted networks
0.812024
TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.812024
SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification · NeurIPS 2024
Computer vision › Image recognition and object detection › image classification › fine-grained image classification
species recognition
0.812024
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024
Environmental and earth informatics
biodiversity informatics
0.812024
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024
Data integration and cleaning › table understanding › table annotation
column type annotation
0.812024
ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models · Proc. VLDB Endow. 2024
Machine learning › Reinforcement learning
benchmark design
0.312025
LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.312025
Hidden in the Noise: Two-Stage Robust Watermarking for Images · ICLR 2025
Natural language and speech › Language models and text generation
instruction tuning
0.312025
WildChat-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training · ICML 2025
Machine learning › Trustworthy machine learning
fairness
0.212024
TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks · NeurIPS 2024
Computer vision › Vision and language
vision-language pretraining
0.212024
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024
Data integration and cleaning › table understanding › table annotation
semantic type detection
0.212024
ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models · Proc. VLDB Endow. 2024

Methods — techniques the papers use, named apart from their topics

fourier patterns · 1.7diffusion model · 1.7supervised fine-tuning · 1.7large language model · 1.6crowdsourcing · 0.9automated scoring · 0.9LLM-as-judge · 0.9zero-shot learning · 0.8prompt serialization · 0.8label remapping · 0.8in-context learning · 0.8context sampling · 0.8context optimization · 0.8CLIP embedding · 0.8CLIP · 0.8
YearPublicationVenuePosition
2025 Hidden in the Noise: Two-Stage Robust Watermarking for Images
abstract
As the quality of image generators continues to improve, deepfakes become a topic of considerable societal debate. Image watermarking allows responsible model owners to detect and label their AI-generated content, which can mitigate the harm. Yet, current state-of-the-art methods in image watermarking remain vulnerable to forgery and removal attacks. This vulnerability occurs in part because watermarks distort the distribution of generated images, unintentionally revealing information about the watermarking techniques. In this work, we first demonstrate a distortion-free watermarking method for images, based on a diffusion model's initial noise. However, detecting the watermark requires comparing the initial noise reconstructed for an image to all previously used initial noises. To mitigate these issues, we propose a two-stage watermarking framework for efficient detection. During generation, we augment the initial noise with generated Fourier patterns to embed information about the group of initial noises we used. For detection, we (i) retrieve the relevant group of noises, and (ii) search within the given group for an initial noise that might match our image. This watermarking approach achieves state-of-the-art robustness to forgery and removal against a large battery of attacks. The project code is available at https://github.com/Kasraarabi/Hidden-in-the-Noise.
Kasra Arabi, Benjamin Feuer, R. Teal Witter, Chinmay Hegde, Niv Cohen
ICLR2
2025 Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
abstract
The release of ChatGPT in November 2022 sparked an explosion of interest in post-training and an avalanche of new preference optimization (PO) methods. These methods claim superior alignment by virtue of better correspondence with human pairwise preferences, often measured by LLM-judges. In this work, we attempt to answer the following question -- do LLM-judge preferences translate to progress on other, more concrete metrics for alignment, and if not, why not? We define a concrete metric for alignment, and introduce SOS-Bench (Substance Outweighs Style Benchmark), the largest standardized, reproducible LLM meta-benchmark to date. We find that (1) LLM-judge preferences do not correlate with concrete measures of safety, world knowledge, and instruction following; (2) LLM-judges have powerful implicit biases, prioritizing style over factuality and safety; and (3) the supervised fine-tuning (SFT) stage of post-training has a large impact on alignment, with data scaling and prompt diversity as the driving factors.
Benjamin Feuer, Micah Goldblum, Teresa Datta, Sanjana Nambiar, Raz Besaleli, Samuel Dooley, Max Cembalest, John Dickerson 0001
ICLR1
2025 LiveBench: A Challenging, Contamination-Limited LLM Benchmark
abstract
Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM judges; however, these can introduce significant biases, and break down when scoring hard questions. In this work, we introduce a new benchmark for LLMs designed to be resistant to both test set contamination and the pitfalls of LLM judging and human crowdsourcing. We release LiveBench, the first benchmark that (1) contains frequently-updated questions from recent information sources, (2) scores answers automatically according to objective ground-truth values, and (3) contains a wide variety of challenging tasks, spanning math, coding, reasoning, language, instruction following, and data analysis. To achieve this, LiveBench contains questions that are based on recently-released math competitions, arXiv papers, news articles, and datasets, and it contains harder, contamination-limited versions of tasks from previous benchmarks such as Big-Bench Hard, AMPS, and IFEval. We evaluate many prominent closed-source models, as well as dozens of open-source models ranging from 0.5B to 405B in size. LiveBench is difficult, with top models achieving below 70% accuracy. We release all questions, code, and model answers. Questions are added and updated on a monthly basis, and we release new tasks and harder versions of tasks over time so that LiveBench can distinguish between the capabilities of LLMs as they improve in the future. We welcome community engagement and collaboration for expanding the benchmark tasks and models.
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain 0001, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, Micah Goldblum
ICLR5
2025 WildChat-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training
abstract
Language model (LLM) post-training can refine behaviors and unlock new skills, but the open science supporting these post-training techniques is still in its infancy. One limiting factor has been the difficulty of conducting large-scale comparative analyses of synthetic data generating models and LLM judges. To close this gap, we introduce WildChat-50M, the largest public chat dataset to date. We extend the existing WildChat dataset to include responses not only from GPT, but from over 50 different open-weight models, ranging in size from 0.5B to 104B parameters. We conduct an extensive comparative analysis and demonstrate the potential of this dataset by creating Re-Wild, our own public SFT mix, which outperforms the recent Tulu-3 SFT mixture from Allen AI with only 40% as many samples.
Benjamin Feuer, Chinmay Hegde
ICML1
2024 TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks
abstract
While tabular classification has traditionally relied on from-scratch training, a recent breakthrough called prior-data fitted networks (PFNs) challenges this approach. Similar to large language models, PFNs make use of pretraining and in-context learning to achieve strong performance on new tasks in a single forward pass. However, current PFNs have limitations that prohibit their widespread adoption. Notably, TabPFN achieves very strong performance on small tabular datasets but is not designed to make predictions for datasets of size larger than 1000. In this work, we overcome these limitations and substantially improve the performance of PFNs via context optimization. We introduce TuneTables, a parameter-efficient fine-tuning strategy for PFNs that compresses large datasets into a smaller learned context. We conduct extensive experiments on nineteen algorithms over 98 datasets and find that TuneTables achieves the best performance on average, outperforming boosted trees such as CatBoost, while optimizing fewer than 5\% of TabPFN's parameters. Furthermore, we show that TuneTables can be used as an interpretability tool and can even be used to mitigate biases by optimizing a fairness objective.
Benjamin Feuer, Robin Schirrmeister, Valeriia Cherepanova, Chinmay Hegde, Frank Hutter, Micah Goldblum, Niv Cohen, Colin White
NeurIPS1
2024 SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification
abstract
Data curation is the problem of how to collect and organize samples into a dataset that supports efficient learning. Despite the centrality of the task, little work has been devoted towards a large-scale, systematic comparison of various curation methods. In this work, we take steps towards a formal evaluation of data curation strategies and introduce SELECT, the first large-scale benchmark of curation strategies for image classification.In order to generate baseline methods for the SELECT benchmark, we create a new dataset, ImageNet++, which constitutes the largest superset of ImageNet-1K to date. Our dataset extends ImageNet with 5 new training-data shifts, each approximately the size of ImageNet-1K, and each assembled using a distinct curation strategy. We evaluate our data curation baselines in two ways: (i) using each training-data shift to train identical image classification models from scratch (ii) using it to inspect a fixed pretrained self-supervised representation.Our findings show interesting trends, particularly pertaining to recent methods for data curation such as synthetic data generation and lookup based on CLIP embeddings. We show that although these strategies are highly competitive for certain tasks, the curation strategy used to assemble the original ImageNet-1K dataset remains the gold standard. We anticipate that our benchmark can illuminate the path for new methods to further reduce the gap. We release our checkpoints, code, documentation, and a link to our dataset at https://github.com/jimmyxu123/SELECT.
Benjamin Feuer, Niv Cohen, Patrick Yubeaton, Govind Mittal, Chinmay Hegde
NeurIPS1
2024 BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity
abstract
We introduce BioTrove, the largest publicly accessible dataset designed to advance AI applications in biodiversity. Curated from the iNaturalist platform and vetted to include only research-grade data, BioTrove contains 161.9 million images, offering unprecedented scale and diversity from three primary kingdoms: Animalia ("animals"), Fungi ("fungi"), and Plantae ("plants"), spanning approximately 366.6K species. Each image is annotated with scientific names, taxonomic hierarchies, and common names, providing rich metadata to support accurate AI model development across diverse species and ecosystems.We demonstrate the value of BioTrove by releasing a suite of CLIP models trained using a subset of 40 million captioned images, known as BioTrove-Train. This subset focuses on seven categories within the dataset that are underrepresented in standard image recognition models, selected for their critical role in biodiversity and agriculture: Aves ("birds"), Arachnida} ("spiders/ticks/mites"), Insecta ("insects"), Plantae ("plants"), Fungi ("fungi"), Mollusca ("snails"), and Reptilia ("snakes/lizards"). To support rigorous assessment, we introduce several new benchmarks and report model accuracy for zero-shot learning across life stages, rare species, confounding species, and multiple taxonomic levels.We anticipate that BioTrove will spur the development of AI models capable of supporting digital tools for pest control, crop monitoring, biodiversity assessment, and environmental conservation. These advancements are crucial for ensuring food security, preserving ecosystems, and mitigating the impacts of climate change. BioTrove is publicly available, easily accessible, and ready for immediate use.
Chih-Hsuan Yang, Benjamin Feuer, Talukder Z. Jubery, Zi K. Deng, Andre Nakkab, Md. Zahid Hasan, Shivani Chiranjeevi, Kelly O. Marshall, Nirmal Baishnab, Asheesh Kumar Singh, Arti Singh, Soumik Sarkar, Nirav C. Merchant, Chinmay Hegde, Baskar Ganapathysubramanian
NeurIPS2
2024 ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models
abstract
Existing deep-learning approaches to semantic column type annotation (CTA) have important shortcomings: they rely on semantic types which are fixed at training time; require a large number of training samples per type; incur high run-time inference costs; and their performance can degrade when evaluated on novel datasets, even when types remain constant. Large language models have exhibited strong zero-shot classification performance on a wide range of tasks and in this paper we explore their use for CTA. We introduce ArcheType, a simple, practical method for context sampling, prompt serialization, model querying, and label remapping, which enables large language models to solve CTA problems in a fully zero-shot manner. We ablate each component of our method separately, and establish that improvements to context sampling and label remapping provide the most consistent gains. ArcheType establishes a new state-of-the-art performance on zero-shot CTA benchmarks (including three new domain-specific benchmarks which we release along with this paper), and when used in conjunction with classical CTA techniques, it outperforms a SOTA DoDuo model on the fine-tuned SOTAB benchmark.
Benjamin Feuer, Yurong Liu, Chinmay Hegde, Juliana Freire
Proc. VLDB Endow.1