EDBT 2026 Demo / reviewers in the wild / expert
Clifford Brunk
dblp:25/4500 · also Cliff Brunk
· DBLP profile ↗
12ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 60% Generative modeling · 28% Trustworthy machine learning · 10% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 96% Data mining · 4% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
machine-generated text detection |
0.5 | 1 | 2021 | Generative Models are Unsupervised Predictors of Page Quality: A Colossal-Scale Study · WSDM 2021 |
Information retrieval
web search |
0.5 | 1 | 2021 | Generative Models are Unsupervised Predictors of Page Quality: A Colossal-Scale Study · WSDM 2021 |
Machine learning › Generative modeling › trustworthy generative modeling
model attribution |
0.4 | 1 | 2020 | Reverse Engineering Configurations of Neural Text Generation Models · ACL 2020 |
Natural language and speech › Language models and text generation › text generation
neural text generation |
0.4 | 1 | 2020 | Reverse Engineering Configurations of Neural Text Generation Models · ACL 2020 |
Machine learning › Trustworthy machine learning
interpretability |
0.1 | 1 | 2021 | Generative Models are Unsupervised Predictors of Page Quality: A Colossal-Scale Study · WSDM 2021 |
Data mining
data mining system |
0.0 | 1 | 1997 | MineSet: An Integrated System for Data Mining · KDD 1997 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › belief revision
theory revision |
0.0 | 1 | 1995 | A Lexical Based Semantic Bias for Theory Revision · ICML 1995 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
relational learning |
0.0 | 1 | 1993 | Finding Accurate Frontiers: A Knowledge-Intensive Approach to Relational Learning · AAAI 1993 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › concept learning
relational concept learning |
0.0 | 1 | 1991 | An Investigation of Noise-Tolerant Relational Concept Learning Algorithms · ML 1991 |
Visualization and visual analytics
data visualization |
0.0 | 1 | 1997 | MineSet: An Integrated System for Data Mining · KDD 1997 |
Natural language and speech › Information extraction and text analysis
lexical semantics |
0.0 | 1 | 1995 | A Lexical Based Semantic Bias for Theory Revision · ICML 1995 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge-intensive learning |
0.0 | 1 | 1991 | A Knowledge-intensive Approach to Learning Relational Concepts · ML 1991 |
Methods — techniques the papers use, named apart from their topics
qualitative analysis · 1.0human evaluation · 1.0generative language model · 1.0diagnostic tests · 0.4theory revision · 0.0lexical semantics · 0.0relational learning · 0.0noise tolerance · 0.0inductive logic programming · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Generative Models are Unsupervised Predictors of Page Quality: A Colossal-Scale StudyabstractLarge generative language models such as GPT-2 are well-known for their ability to generate text as well as their utility in supervised downstream tasks via fine-tuning. Its prevalence on the web, however, is still not well understood - if we run GPT-2 detectors across the web, what will we find? Our work is twofold: firstly we demonstrate via human evaluation that classifiers trained to discriminate between human and machine-generated text emerge as unsupervised predictors of "page quality", able to detect low quality content without any training. This enables fast bootstrapping of quality indicators in a low-resource setting. Secondly, curious to understand the prevalence and nature of low quality pages in the wild, we conduct extensive qualitative and quantitative analysis over 500 million web articles, making this the largest-scale study ever conducted on the topic. Dara Bahri, Yi Tay, Che Zheng, Clifford Brunk, Donald Metzler, Andrew Tomkins |
WSDM | 4 |
| 2020 | Reverse Engineering Configurations of Neural Text Generation ModelsabstractThis paper seeks to develop a deeper understanding of the fundamental properties of neural text generations models.The study of artifacts that emerge in machine generated text as a result of modeling choices is a nascent research area.Previously, the extent and degree to which these artifacts surface in generated text has not been well studied.In the spirit of better understanding generative text models and their artifacts, we propose the new task of distinguishing which of several variants of a given model generated a piece of text, and we conduct an extensive suite of diagnostic tests to observe whether modeling choices (e.g., sampling methods, top-k probabilities, model architectures, etc.) leave detectable artifacts in the text they generate.Our key finding, which is backed by a rigorous set of experiments, is that such artifacts are present and that different modeling choices can be inferred by observing the generated text alone.This suggests that neural text generators may be more sensitive to various modeling choices than previously thought. Yi Tay, Dara Bahri, Che Zheng, Clifford Brunk, Donald Metzler, Andrew Tomkins |
ACL | 4 |
| 2016 | Multilingual Language Processing From BytesabstractDan Gillick, Cliff Brunk, Oriol Vinyals, Amarnag Subramanya. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Daniel Gillick, Clifford Brunk, Oriol Vinyals, Amarnag Subramanya |
HLT-NAACL | 2 |
| 2008 | Towards Click-Based Models of Geographic Interests in Web SearchabstractWith the recent surge in the volume of search queries that explicitly or implicitly express users' geographical interests, to accurately infer users' locality preference becomes an increasingly important yet challenging issue. We study two click-based models of the distribution of such geographical interests by mining the user click stream data in the search engine logs, addressing three important issues in spatial Web search. First, search queries and documents can be classified by the models according to their spatial specificity. Second, the geographic center(s) of interests for queries and documents can be inferred. Finally, the model can be applied to generate relevance features for search ranking. We evaluated our proposals on a large dataset with about 10,000 unique queries sampled from the Yahoo! Search query logs, and about 450 million user clicks on 1.4 million unique Web pages over a six-months period. We report about 90% accuracy and about 3% false positive rate in identifying search queries with or without specific geographical interests, as well as statistically significant improvement in relevance ranking over a strong baseline. Ziming Zhuang, Clifford Brunk, Prasenjit Mitra 0001, C. Lee Giles |
Web Intelligence | 2 |
| 1998 | Pruning Decision Trees with Misclassification Costs
Jeffrey P. Bradford, Clayton Kunz, Ron Kohavi, Clifford Brunk, Carla E. Brodley |
ECML | 4 |
| 1997 | MineSet: An Integrated System for Data Mining
Clifford Brunk, James Kelly, Ron Kohavi |
KDD | 1 |
| 1995 | A Lexical Based Semantic Bias for Theory Revision
Clifford Brunk, Michael J. Pazzani |
ICML | 1 |
| 1994 | Reducing Misclassification Costs
Michael J. Pazzani, Christopher J. Merz, Patrick M. Murphy, Kamal M. Ali, Timothy Hume, Clifford Brunk |
ICML | 6 |
| 1994 | On Learning Multiple Descriptions of a ConceptabstractIn sparse data environments, greater classification accuracy can be achieved by learning several concept descriptions of the data and combining their classifications. Stochastic searching can be used to generate many concept descriptions (rule sets) for each class in the data. We use a tractable approximation to the optimal Bayesian method for combining classifications from such descriptions. The primary result of this paper is that multiple concept descriptions are particularly helpful in "flat" hypothesis spaces in which there are many equally good ways to grow a rule, each having similar gain. Another result is experimental evidence that learning multiple rule sets yields more accurate classifications than learning multiple rules for some domains.> Kamal M. Ali, Clifford Brunk, Michael J. Pazzani |
ICTAI | 2 |
| 1993 | Finding Accurate Frontiers: A Knowledge-Intensive Approach to Relational Learning
Michael J. Pazzani, Clifford Brunk |
AAAI | 2 |
| 1991 | An Investigation of Noise-Tolerant Relational Concept Learning Algorithms
Clifford Brunk, Michael J. Pazzani |
ML | 1 |
| 1991 | A Knowledge-intensive Approach to Learning Relational Concepts
Michael J. Pazzani, Clifford Brunk, Glenn Silverstein |
ML | 2 |