VLDB 2026 Research / reviewers in the wild / expert
Jannis Bulian
dblp:09/10967
· DBLP profile ↗
11ranked-venue papers
7as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Theory of computation · 5 · 5 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 39% Efficient and distributed learning · 24% Question answering and dialogue systems · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Environmental and earth informatics · 50% Computing education · 50% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 2 | 2024 | Assessing Large Language Models on Climate Information · ICML 2024 Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation · EMNLP 2022 |
Machine learning › Trustworthy machine learning
AI safety |
0.8 | 1 | 2024 | On scalable oversight with weak LLMs judging strong LLMs · NeurIPS 2024 |
Natural language and speech › Language models and text generation › alignment
scalable oversight |
0.8 | 1 | 2024 | On scalable oversight with weak LLMs judging strong LLMs · NeurIPS 2024 |
Environmental and earth informatics › climate science › climate change
climate change communication |
0.8 | 1 | 2024 | Assessing Large Language Models on Climate Information · ICML 2024 |
Computing education
large language model evaluation |
0.8 | 1 | 2024 | Assessing Large Language Models on Climate Information · ICML 2024 |
Machine learning › Efficient and distributed learning › distributed training
distributed training systems |
0.7 | 1 | 2023 | Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023 |
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training |
0.7 | 1 | 2023 | Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023 |
Natural language and speech › Question answering and dialogue systems
question answering evaluation |
0.6 | 1 | 2022 | Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation · EMNLP 2022 |
Information retrieval
evaluation |
0.6 | 1 | 2022 | Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation · EMNLP 2022 |
Information retrieval
retrieval models |
0.6 | 1 | 2022 | Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation · EMNLP 2022 |
Natural language and speech › Language models and text generation
large language model |
0.4 | 2 | 2024 | On scalable oversight with weak LLMs judging strong LLMs · NeurIPS 2024 Scaling Up Models and Data with t5x and seqio · J. Mach. Learn. Res. 2023 |
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering |
0.3 | 1 | 2018 | Ask the Right Questions: Active Question Reformulation with Reinforcement Learning · ICLR 2018 |
Machine learning › Reinforcement learning
policy learning |
0.3 | 1 | 2018 | Ask the Right Questions: Active Question Reformulation with Reinforcement Learning · ICLR 2018 |
Human-AI interaction › AI-assisted decision-making
AI-assisted evaluation |
0.2 | 1 | 2024 | Assessing Large Language Models on Climate Information · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
human rating protocol · 2.3AI assistance · 2.3human annotation · 1.1BERT matching · 1.1weak-to-strong supervision · 0.8debate · 0.8consultancy · 0.8reinforcement learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Assessing Large Language Models on Climate InformationabstractAs Large Language Models (LLMs) rise in popularity, it is necessary to assess their capability in critically relevant domains. We present a comprehensive evaluation framework, grounded in science communication research, to assess LLM responses to questions about climate change. Our framework emphasizes both presentational and epistemological adequacy, offering a fine-grained analysis of LLM generations spanning 8 dimensions and 30 issues. Our evaluation task is a real-world example of a growing number of challenging problems where AI can complement and lift human performance. We introduce a novel protocol for scalable oversight that relies on AI Assistance and raters with relevant education. We evaluate several recent LLMs on a set of diverse climate questions. Our results point to a significant gap between surface and epistemological qualities of LLMs in the realm of climate communication. Jannis Bulian, Mike S. Schäfer, Afra Amini, Heidi Lam, Massimiliano Ciaramita, Ben Gaiarin, Michelle Chen Huebscher, Christian Buck, Niels Mede, Markus Leippold, Nadine Strauß |
ICML | 1 |
| 2024 | On scalable oversight with weak LLMs judging strong LLMsabstractScalable oversight protocols aim to enable humans to accurately supervise superhuman AI.
In this paper we study debate, where two AI's compete to convince a judge; consultancy,
where a single AI tries to convince a judge that asks questions;
and compare to a baseline of direct question-answering, where the judge just answers outright without the AI.
We use large language models (LLMs) as both AI agents and as stand-ins for human judges, taking the judge models to be weaker than agent models.
We benchmark on a diverse range of asymmetries between judges and agents, extending previous work on a single extractive QA task with information asymmetry, to also include mathematics, coding, logic and multimodal reasoning asymmetries.
We find that debate outperforms consultancy across all tasks when the consultant is randomly assigned to argue for the correct/incorrect answer. Comparing debate to direct question answering, the results depend on the type of task: in extractive QA tasks with information asymmetry debate outperforms direct question answering, but in other tasks without information asymmetry the results are mixed.
Previous work assigned debaters/consultants an answer to argue for. When we allow them to instead choose which answer to argue for, we find judges are less frequently convinced by the wrong answer in debate than in consultancy.
Further, we find that stronger debater models increase judge accuracy, though more modestly than in previous studies. Zachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah D. Goodman, Rohin Shah |
NeurIPS | 6 |
| 2023 | Scaling Up Models and Data with t5x and seqioabstractScaling up training datasets and model parameters have benefited neural network-based language models, but also present challenges like distributed compute, input data bottlenecks and reproducibility of results. We introduce two simple and scalable software libraries that simplify these issues: t5x enables training large language models at scale, while seqio enables reproducible input and evaluation pipelines. These open-source libraries have been used to train models with hundreds of billions of parameters on multi-terabyte datasets. Configurations and instructions for T5-like and GPT-like models are also provided. The libraries can be found at https://github.com/google-research/t5x and https://github.com/google/seqio. Adam Roberts, Hyung Won Chung, Anselm Levskaya, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio B. Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Kathleen Kenealy, Kehang Han, Michelle Casbon, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Tachard Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, Andrea Gesmundo |
J. Mach. Learn. Res. | 22 |
| 2022 | Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering EvaluationabstractThe predictions of question answering (QA) systems are typically evaluated against manually annotated finite sets of one or more answers.This leads to a coverage limitation that results in underestimating the true performance of systems, and is typically addressed by extending over exact match (EM) with predefined rules or with the token-level F 1 measure.In this paper, we present the first systematic conceptual and data-driven analysis to examine the shortcomings of token-level equivalence measures.To this end, we define the asymmetric notion of answer equivalence (AE), accepting answers that are equivalent to or improve over the reference, and publish over 23k human judgments for candidates produced by multiple QA systems on SQuAD. 1 Through a careful analysis of this data, we reveal and quantify several concrete limitations of the F 1 measure, such as a false impression of graduality, or missing dependence on the question.Since collecting AE annotations for each evaluated model is expensive, we learn a BERT matching (BEM) measure to approximate this task.Being a simpler task than QA, we find BEM to provide significantly better AE approximations than F 1 , and to more accurately reflect the performance of systems.Finally, we demonstrate the practical utility of AE and BEM on the concrete application of minimal accurate prediction sets, reducing the number of required answers by up to ×2.6. Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, Tal Schuster |
EMNLP | 1 |
| 2021 | Fool Me Twice: Entailment from Wikipedia GamificationabstractJulian Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin Börschinger, Jordan Boyd-Graber. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Julian Martin Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin Börschinger, Jordan L. Boyd-Graber |
NAACL-HLT | 3 |
| 2018 | Ask the Right Questions: Active Question Reformulation with Reinforcement Learning
Christian Buck, Jannis Bulian, Massimiliano Ciaramita, Wojciech Gajewski, Andrea Gesmundo, Neil Houlsby, Wei Wang 0236 |
ICLR | 2 |
| 2017 | Fixed-Parameter Tractable Distances to Sparse Graph ClassesabstractWe show that for various classes $$\mathcal {C}$$ of sparse graphs, and several measures of distance to such classes (such as edit distance and elimination distance), the problem of determining the distance of a given graph G to $$\mathcal {C}$$ is fixed-parameter tractable. The results are based on two general techniques. The first of these, building on recent work of Grohe et al. establishes that any class of graphs that is slicewise nowhere dense and slicewise first-order definable is $$\mathrm {FPT} $$ . The second shows that determining the elimination distance of a graph G to a minor-closed class $$\mathcal {C}$$ is $$\mathrm {FPT} $$ . We demonstrate that several prior results (of Golovach, Moser and Thilikos and Mathieson) on the fixed-parameter tractability of distance measures are special cases of our first method. Jannis Bulian, Anuj Dawar |
Algorithmica | 1 |
| 2016 | Graph Isomorphism Parameterized by Elimination Distance to Bounded DegreeabstractA commonly studied means of parameterizing graph problems is the deletion distance from triviality (Guo et al., Parameterized and exact computation, Springer, Berlin, pp. 162–173, 2004), which counts vertices that need to be deleted from a graph to place it in some class for which efficient algorithms are known. In the context of graph isomorphism, we define triviality to mean a graph with maximum degree bounded by a constant, as such graph classes admit polynomial-time isomorphism tests. We generalise deletion distance to a measure we call elimination distance to triviality, based on elimination trees or tree-depth decompositions. We establish that graph canonisation, and thus graph isomorphism, is $$\mathsf {FPT}$$ when parameterized by elimination distance to bounded degree, extending results of Bouland et al. (Parameterized and exact computation, Springer, Berlin, pp. 218–230, 2012). Jannis Bulian, Anuj Dawar |
Algorithmica | 1 |
| 2015 | Fixed-parameter Tractable Distances to Sparse Graph ClassesabstractWe show that for various classes C of sparse graphs, and several measures of distance to such classes (such as edit distance and elimination distance), the problem of determining the distance of a given graph G to C is fixed-parameter tractable. The results are based on two general techniques. The first of these, building on recent work of Grohe et al. establishes that any class of graphs that is slicewise nowhere dense and slicewise first-order definable is FPT. The second shows that determining the elimination distance of a graph G to a minor-closed class C is FPT. Jannis Bulian, Anuj Dawar |
IPEC | 1 |
| 2014 | Graph Isomorphism Parameterized by Elimination Distance to Bounded Degree
Jannis Bulian, Anuj Dawar |
IPEC | 1 |
| 2013 | Bare canonicity of representable cylindric and polyadic algebrasabstractWe show that for finite n⩾3, every first-order axiomatisation of the varieties of representable n-dimensional cylindric algebras, diagonal-free cylindric algebras, polyadic algebras, and polyadic equality algebras contains an infinite number of non-canonical formulas. We also show that the class of structures for each of these varieties is non-elementary. The proofs employ algebras derived from random graphs. Jannis Bulian, Ian M. Hodkinson |
Ann. Pure Appl. Log. | 1 |