VLDB 2026 Research / reviewers in the wild / expert
Benjamin Feuer
dblp:322/5063
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-7938-542XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 50% Image recognition and object detection · 16% Efficient and distributed learning · 10% | |
| Network and information security
1 paper |
Digital forensics and information hiding · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Environmental and earth informatics · 100% |
Topics — the 26 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
benchmark contamination |
0.9 | 1 | 2025 | LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM judge |
0.9 | 1 | 2025 | Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model training
post-training |
0.9 | 1 | 2025 | WildChat-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training · ICML 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning |
0.9 | 1 | 2025 | WildChat-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training · ICML 2025 |
Natural language and speech › Language models and text generation
synthetic data |
0.9 | 1 | 2025 | WildChat-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training · ICML 2025 |
Digital forensics and information hiding › watermarking
image watermarking |
0.9 | 1 | 2025 | Hidden in the Noise: Two-Stage Robust Watermarking for Images · ICLR 2025 |
Digital forensics and information hiding
watermarking |
0.9 | 1 | 2025 | Hidden in the Noise: Two-Stage Robust Watermarking for Images · ICLR 2025 |
Machine learning › Efficient and distributed learning
data curation |
0.8 | 1 | 2024 | SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification · NeurIPS 2024 |
Computer vision › Image recognition and object detection › image classification
fine-grained image classification |
0.8 | 1 | 2024 | BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024 |
Computer vision › Image recognition and object detection
image classification |
0.8 | 1 | 2024 | SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.8 | 1 | 2024 | TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › neural processes
prior-data fitted networks |
0.8 | 1 | 2024 | TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.8 | 1 | 2024 | SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification · NeurIPS 2024 |
Computer vision › Image recognition and object detection › image classification › fine-grained image classification
species recognition |
0.8 | 1 | 2024 | BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024 |
Environmental and earth informatics
biodiversity informatics |
0.8 | 1 | 2024 | BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024 |
Data integration and cleaning › table understanding › table annotation
column type annotation |
0.8 | 1 | 2024 | ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models · Proc. VLDB Endow. 2024 |
Machine learning › Reinforcement learning
benchmark design |
0.3 | 1 | 2025 | LiveBench: A Challenging, Contamination-Limited LLM Benchmark · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
0.3 | 1 | 2025 | Hidden in the Noise: Two-Stage Robust Watermarking for Images · ICLR 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.3 | 1 | 2025 | WildChat-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training · ICML 2025 |
Machine learning › Trustworthy machine learning
fairness |
0.2 | 1 | 2024 | TuneTables: Context Optimization for Scalable Prior-Data Fitted Networks · NeurIPS 2024 |
Computer vision › Vision and language
vision-language pretraining |
0.2 | 1 | 2024 | BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity · NeurIPS 2024 |
Data integration and cleaning › table understanding › table annotation
semantic type detection |
0.2 | 1 | 2024 | ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language Models · Proc. VLDB Endow. 2024 |
Methods — techniques the papers use, named apart from their topics
fourier patterns · 1.7diffusion model · 1.7supervised fine-tuning · 1.7large language model · 1.6crowdsourcing · 0.9automated scoring · 0.9LLM-as-judge · 0.9zero-shot learning · 0.8prompt serialization · 0.8label remapping · 0.8in-context learning · 0.8context sampling · 0.8context optimization · 0.8CLIP embedding · 0.8CLIP · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hidden in the Noise: Two-Stage Robust Watermarking for ImagesabstractAs the quality of image generators continues to improve, deepfakes become a topic of considerable societal debate. Image watermarking allows responsible model owners to detect and label their AI-generated content, which can mitigate the harm. Yet, current state-of-the-art methods in image watermarking remain vulnerable to forgery and removal attacks.
This vulnerability occurs in part because watermarks distort the distribution of generated images, unintentionally revealing information about the watermarking techniques.
In this work, we first demonstrate a distortion-free watermarking method for images, based on a diffusion model's initial noise.
However, detecting the watermark requires comparing the initial noise reconstructed for an image to all previously used initial noises.
To mitigate these issues, we propose a two-stage watermarking framework for efficient detection. During generation, we augment the initial noise with generated Fourier patterns to embed information about the group of initial noises we used. For detection, we (i) retrieve the relevant group of noises, and (ii) search within the given group for an initial noise that might match our image. This watermarking approach achieves state-of-the-art robustness to forgery and removal against a large battery of attacks. The project code is available at https://github.com/Kasraarabi/Hidden-in-the-Noise. Kasra Arabi, Benjamin Feuer, R. Teal Witter, Chinmay Hegde, Niv Cohen |
ICLR | 2 |
| 2025 | Style Outweighs Substance: Failure Modes of LLM Judges in Alignment BenchmarkingabstractThe release of ChatGPT in November 2022 sparked an explosion of interest in post-training and an avalanche of new preference optimization (PO) methods. These methods claim superior alignment by virtue of better correspondence with human pairwise preferences, often measured by LLM-judges. In this work, we attempt to answer the following question -- do LLM-judge preferences translate to progress on other, more concrete metrics for alignment, and if not, why not? We define a concrete metric for alignment, and introduce SOS-Bench (Substance Outweighs Style Benchmark), the largest standardized, reproducible LLM meta-benchmark to date. We find that (1) LLM-judge preferences do not correlate with concrete measures of safety, world knowledge, and instruction following; (2) LLM-judges have powerful implicit biases, prioritizing style over factuality and safety; and (3) the supervised fine-tuning (SFT) stage of post-training has a large impact on alignment, with data scaling and prompt diversity as the driving factors. Benjamin Feuer, Micah Goldblum, Teresa Datta, Sanjana Nambiar, Raz Besaleli, Samuel Dooley, Max Cembalest, John Dickerson 0001 |
ICLR | 1 |
| 2025 | LiveBench: A Challenging, Contamination-Limited LLM BenchmarkabstractTest set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM judges; however, these can introduce significant biases, and break down when scoring hard questions. In this work, we introduce a new benchmark for LLMs designed to be resistant to both test set contamination and the pitfalls of LLM judging and human crowdsourcing. We release LiveBench, the first benchmark that (1) contains frequently-updated questions from recent information sources, (2) scores answers automatically according to objective ground-truth values, and (3) contains a wide variety of challenging tasks, spanning math, coding, reasoning, language, instruction following, and data analysis. To achieve this, LiveBench contains questions that are based on recently-released math competitions, arXiv papers, news articles, and datasets, and it contains harder, contamination-limited versions of tasks from previous benchmarks such as Big-Bench Hard, AMPS, and IFEval. We evaluate many prominent closed-source models, as well as dozens of open-source models ranging from 0.5B to 405B in size. LiveBench is difficult, with top models achieving below 70% accuracy. We release all questions, code, and model answers. Questions are added and updated on a monthly basis, and we release new tasks and harder versions of tasks over time so that LiveBench can distinguish between the capabilities of LLMs as they improve in the future. We welcome community engagement and collaboration for expanding the benchmark tasks and models. Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain 0001, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, Micah Goldblum |
ICLR | 5 |
| 2025 | WildChat-50M: A Deep Dive Into the Role of Synthetic Data in Post-TrainingabstractLanguage model (LLM) post-training can refine behaviors and unlock new skills, but the open science supporting these post-training techniques is still in its infancy. One limiting factor has been the difficulty of conducting large-scale comparative analyses of synthetic data generating models and LLM judges. To close this gap, we introduce WildChat-50M, the largest public chat dataset to date. We extend the existing WildChat dataset to include responses not only from GPT, but from over 50 different open-weight models, ranging in size from 0.5B to 104B parameters. We conduct an extensive comparative analysis and demonstrate the potential of this dataset by creating Re-Wild, our own public SFT mix, which outperforms the recent Tulu-3 SFT mixture from Allen AI with only 40% as many samples. Benjamin Feuer, Chinmay Hegde |
ICML | 1 |
| 2024 | TuneTables: Context Optimization for Scalable Prior-Data Fitted NetworksabstractWhile tabular classification has traditionally relied on from-scratch training, a recent breakthrough called prior-data fitted networks (PFNs) challenges this approach. Similar to large language models, PFNs make use of pretraining and in-context learning to achieve strong performance on new tasks in a single forward pass. However, current PFNs have limitations that prohibit their widespread adoption. Notably, TabPFN achieves very strong performance on small tabular datasets but is not designed to make predictions for datasets of size larger than 1000. In this work, we overcome these limitations and substantially improve the performance of PFNs via context optimization. We introduce TuneTables, a parameter-efficient fine-tuning strategy for PFNs that compresses large datasets into a smaller learned context. We conduct extensive experiments on nineteen algorithms over 98 datasets and find that TuneTables achieves the best performance on average, outperforming boosted trees such as CatBoost, while optimizing fewer than 5\% of TabPFN's parameters. Furthermore, we show that TuneTables can be used as an interpretability tool and can even be used to mitigate biases by optimizing a fairness objective. Benjamin Feuer, Robin Schirrmeister, Valeriia Cherepanova, Chinmay Hegde, Frank Hutter, Micah Goldblum, Niv Cohen, Colin White |
NeurIPS | 1 |
| 2024 | SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image ClassificationabstractData curation is the problem of how to collect and organize samples into a dataset that supports efficient learning. Despite the centrality of the task, little work has been devoted towards a large-scale, systematic comparison of various curation methods. In this work, we take steps towards a formal evaluation of data curation strategies and introduce SELECT, the first large-scale benchmark of curation strategies for image classification.In order to generate baseline methods for the SELECT benchmark, we create a new dataset, ImageNet++, which constitutes the largest superset of ImageNet-1K to date. Our dataset extends ImageNet with 5 new training-data shifts, each approximately the size of ImageNet-1K, and each assembled using a distinct curation strategy. We evaluate our data curation baselines in two ways: (i) using each training-data shift to train identical image classification models from scratch (ii) using it to inspect a fixed pretrained self-supervised representation.Our findings show interesting trends, particularly pertaining to recent methods for data curation such as synthetic data generation and lookup based on CLIP embeddings. We show that although these strategies are highly competitive for certain tasks, the curation strategy used to assemble the original ImageNet-1K dataset remains the gold standard. We anticipate that our benchmark can illuminate the path for new methods to further reduce the gap. We release our checkpoints, code, documentation, and a link to our dataset at https://github.com/jimmyxu123/SELECT. Benjamin Feuer, Niv Cohen, Patrick Yubeaton, Govind Mittal, Chinmay Hegde |
NeurIPS | 1 |
| 2024 | BioTrove: A Large Curated Image Dataset Enabling AI for BiodiversityabstractWe introduce BioTrove, the largest publicly accessible dataset designed to advance AI applications in biodiversity. Curated from the iNaturalist platform and vetted to include only research-grade data, BioTrove contains 161.9 million images, offering unprecedented scale and diversity from three primary kingdoms: Animalia ("animals"), Fungi ("fungi"), and Plantae ("plants"), spanning approximately 366.6K species. Each image is annotated with scientific names, taxonomic hierarchies, and common names, providing rich metadata to support accurate AI model development across diverse species and ecosystems.We demonstrate the value of BioTrove by releasing a suite of CLIP models trained using a subset of 40 million captioned images, known as BioTrove-Train. This subset focuses on seven categories within the dataset that are underrepresented in standard image recognition models, selected for their critical role in biodiversity and agriculture: Aves ("birds"), Arachnida} ("spiders/ticks/mites"), Insecta ("insects"), Plantae ("plants"), Fungi ("fungi"), Mollusca ("snails"), and Reptilia ("snakes/lizards"). To support rigorous assessment, we introduce several new benchmarks and report model accuracy for zero-shot learning across life stages, rare species, confounding species, and multiple taxonomic levels.We anticipate that BioTrove will spur the development of AI models capable of supporting digital tools for pest control, crop monitoring, biodiversity assessment, and environmental conservation. These advancements are crucial for ensuring food security, preserving ecosystems, and mitigating the impacts of climate change. BioTrove is publicly available, easily accessible, and ready for immediate use. Chih-Hsuan Yang, Benjamin Feuer, Talukder Z. Jubery, Zi K. Deng, Andre Nakkab, Md. Zahid Hasan, Shivani Chiranjeevi, Kelly O. Marshall, Nirmal Baishnab, Asheesh Kumar Singh, Arti Singh, Soumik Sarkar, Nirav C. Merchant, Chinmay Hegde, Baskar Ganapathysubramanian |
NeurIPS | 2 |
| 2024 | ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language ModelsabstractExisting deep-learning approaches to semantic column type annotation (CTA) have important shortcomings: they rely on semantic types which are fixed at training time; require a large number of training samples per type; incur high run-time inference costs; and their performance can degrade when evaluated on novel datasets, even when types remain constant. Large language models have exhibited strong zero-shot classification performance on a wide range of tasks and in this paper we explore their use for CTA. We introduce ArcheType, a simple, practical method for context sampling, prompt serialization, model querying, and label remapping, which enables large language models to solve CTA problems in a fully zero-shot manner. We ablate each component of our method separately, and establish that improvements to context sampling and label remapping provide the most consistent gains. ArcheType establishes a new state-of-the-art performance on zero-shot CTA benchmarks (including three new domain-specific benchmarks which we release along with this paper), and when used in conjunction with classical CTA techniques, it outperforms a SOTA DoDuo model on the fine-tuned SOTAB benchmark. Benjamin Feuer, Yurong Liu, Chinmay Hegde, Juliana Freire |
Proc. VLDB Endow. | 1 |