EDBT 2026 Demo / reviewers in the wild / expert
János Kramár
dblp:49/9013
· DBLP profile ↗
8ranked-venue papers
0as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Trustworthy machine learning · 52% Reinforcement learning · 14% Language models and text generation · 14% | |
| Theoretical computer science
3 papers |
Algorithmic game theory and mechanism design · 98% Logic in computer science · 2% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 100% |
Topics — the 21 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
1.4 | 2 | 2024 | Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024 Tracr: Compiled Transformers as a Laboratory for Interpretability · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
1.4 | 2 | 2024 | Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024 Tracr: Compiled Transformers as a Laboratory for Interpretability · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
AI safety |
0.8 | 1 | 2024 | On scalable oversight with weak LLMs judging strong LLMs · NeurIPS 2024 |
Natural language and speech › Language models and text generation › alignment
scalable oversight |
0.8 | 1 | 2024 | On scalable oversight with weak LLMs judging strong LLMs · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder |
0.8 | 1 | 2024 | Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
sparse feature learning |
0.8 | 1 | 2024 | Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024 |
Collaborative and social computing › cooperative work
group decision-making |
0.5 | 1 | 2021 | A Neural Network Auction For Group Decision Making Over a Continuous Space · IJCAI 2021 |
Algorithmic game theory and mechanism design › mechanism design
auction design |
0.5 | 1 | 2021 | A Neural Network Auction For Group Decision Making Over a Continuous Space · IJCAI 2021 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.4 | 1 | 2020 | Learning to Play No-Press Diplomacy with Best Response Policy Iteration · NeurIPS 2020 |
Machine learning › Reinforcement learning › dynamic programming
policy iteration |
0.4 | 1 | 2020 | Learning to Play No-Press Diplomacy with Best Response Policy Iteration · NeurIPS 2020 |
Algorithmic game theory and mechanism design
equilibrium computation |
0.4 | 1 | 2020 | Learning to Play No-Press Diplomacy with Best Response Policy Iteration · NeurIPS 2020 |
Algorithmic game theory and mechanism design › learning in games
fictitious play |
0.4 | 1 | 2020 | Learning to Play No-Press Diplomacy with Best Response Policy Iteration · NeurIPS 2020 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.4 | 1 | 2019 | Relational Forward Models for Multi-Agent Learning · ICLR (Poster) 2019 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning |
0.4 | 1 | 2019 | Relational Forward Models for Multi-Agent Learning · ICLR (Poster) 2019 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2017 | Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations · ICLR (Poster) 2017 |
Machine learning › Deep learning architectures and training
regularization |
0.3 | 1 | 2017 | Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations · ICLR (Poster) 2017 |
Natural language and speech › Language models and text generation
large language model |
0.2 | 1 | 2024 | On scalable oversight with weak LLMs judging strong LLMs · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
transformer interpretability |
0.2 | 1 | 2023 | Tracr: Compiled Transformers as a Laboratory for Interpretability · NeurIPS 2023 |
Information retrieval
compact coding |
0.1 | 1 | 2010 | A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices · ACL 2010 |
Data mining › pattern mining › formal concept analysis
concept lattice |
0.1 | 1 | 2010 | A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices · ACL 2010 |
Logic in computer science › knowledge representation and reasoning
formal concept analysis |
0.0 | 1 | 2010 | A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices · ACL 2010 |
Methods — techniques the papers use, named apart from their topics
neural network · 1.0mechanism design · 1.0gradient ascent · 1.0deep reinforcement learning · 0.9approximate best response operator · 0.9weak-to-strong supervision · 0.8l1 penalty · 0.8gated sparse autoencoder · 0.8debate · 0.8consultancy · 0.8superposition analysis · 0.7program compilation · 0.7graph neural network · 0.4zero-preserving encoding · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | On scalable oversight with weak LLMs judging strong LLMsabstractScalable oversight protocols aim to enable humans to accurately supervise superhuman AI.
In this paper we study debate, where two AI's compete to convince a judge; consultancy,
where a single AI tries to convince a judge that asks questions;
and compare to a baseline of direct question-answering, where the judge just answers outright without the AI.
We use large language models (LLMs) as both AI agents and as stand-ins for human judges, taking the judge models to be weaker than agent models.
We benchmark on a diverse range of asymmetries between judges and agents, extending previous work on a single extractive QA task with information asymmetry, to also include mathematics, coding, logic and multimodal reasoning asymmetries.
We find that debate outperforms consultancy across all tasks when the consultant is randomly assigned to argue for the correct/incorrect answer. Comparing debate to direct question answering, the results depend on the type of task: in extractive QA tasks with information asymmetry debate outperforms direct question answering, but in other tasks without information asymmetry the results are mixed.
Previous work assigned debaters/consultants an answer to argue for. When we allow them to instead choose which answer to argue for, we find judges are less frequently convinced by the wrong answer in debate than in consultancy.
Further, we find that stronger debater models increase judge accuracy, though more modestly than in previous studies. Zachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah D. Goodman, Rohin Shah |
NeurIPS | 3 |
| 2024 | Improving Sparse Decomposition of Language Model Activations with Gated Sparse AutoencodersabstractRecent work has found that sparse autoencoders (SAEs) are an effective technique for unsupervised discovery of interpretable features in language models' (LMs) activations, by finding sparse, linear reconstructions of those activations. We introduce the Gated Sparse Autoencoder (Gated SAE), which achieves a Pareto improvement over training with prevailing methods. In SAEs, the L1 penalty used to encourage sparsity introduces many undesirable biases, such as shrinkage -- systematic underestimation of feature activations. The key insight of Gated SAEs is to separate the functionality of (a) determining which directions to use and (b) estimating the magnitudes of those directions: this enables us to apply the L1 penalty only to the former, limiting the scope of undesirable side effects. Through training SAEs on LMs of up to 7B parameters we find that, in typical hyper-parameter ranges, Gated SAEs solve shrinkage, are similarly interpretable, and require half as many firing features to achieve comparable reconstruction fidelity. Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah, Neel Nanda |
NeurIPS | 6 |
| 2023 | Tracr: Compiled Transformers as a Laboratory for InterpretabilityabstractWe show how to "compile" human-readable programs into standard decoder-only transformer models. Our compiler, Tracr, generates models with known structure. This structure can be used to design experiments. For example, we use it to study "superposition" in transformers that execute multi-step algorithms. Additionally, the known structure of Tracr-compiled models can serve as _ground-truth_ for evaluating interpretability methods. Commonly, because the "programs" learned by transformers are unknown it is unclear whether an interpretation succeeded. We demonstrate our approach by implementing and examining programs including computing token frequencies, sorting, and parenthesis checking. We provide an open-source implementation of Tracr at https://github.com/google-deepmind/tracr. David Lindner, János Kramár, Sebastian Farquhar, Matthew Rahtz, Thomas McGrath 0001, Vladimir Mikulik |
NeurIPS | 2 |
| 2021 | A Neural Network Auction For Group Decision Making Over a Continuous SpaceabstractWe propose a system for conducting an auction over locations in a continuous space. It enables participants to express their preferences over possible choices of location in the space, selecting the location that maximizes the total utility of all agents. We prevent agents from tricking the system into selecting a location that improves their individual utility at the expense of others by using a pricing rule that gives agents no incentive to misreport their true preferences. The system queries participants for their utility in many random locations, then trains a neural network to approximate the preference function of each participant. The parameters of these neural network models are transmitted and processed by the auction mechanism, which composes these into differentiable models that are optimized through gradient ascent to compute the final chosen location and charged prices. Yoram Bachrach, Ian Gemp, Marta Garnelo, János Kramár, Tom Eccles, Dan Rosenbaum, Thore Graepel |
IJCAI | 4 |
| 2020 | Learning to Play No-Press Diplomacy with Best Response Policy IterationabstractRecent advances in deep reinforcement learning (RL) have led to considerable progress in many 2-player zero-sum games, such as Go, Poker and Starcraft. The purely adversarial nature of such games allows for conceptually simple and principled application of RL methods. However real-world settings are many-agent, and agent interactions are complex mixtures of common-interest and competitive aspects. We consider Diplomacy, a 7-player board game designed to accentuate dilemmas resulting from many-agent interactions. It also features a large combinatorial action space and simultaneous moves, which are challenging for RL algorithms. We propose a simple yet effective approximate best response operator, designed to handle large combinatorial action spaces and simultaneous moves. We also introduce a family of policy iteration methods that approximate fictitious play. With these methods, we successfully apply RL to Diplomacy: we show that our agents convincingly outperform the previous state-of-the-art, and game theoretic equilibrium analysis shows that the new process yields consistent improvements. Thomas W. Anthony 0001, Tom Eccles, Andrea Tacchetti, János Kramár, Ian Gemp, Thomas C. Hudson, Nicolas Porcel, Marc Lanctot, Julien Pérolat, Richard Everett 0001, Satinder Singh 0001, Thore Graepel, Yoram Bachrach |
NeurIPS | 4 |
| 2019 | Relational Forward Models for Multi-Agent Learning
Andrea Tacchetti, H. Francis Song, Pedro A. M. Mediano, Vinícius Flores Zambaldi, János Kramár, Neil C. Rabinowitz, Thore Graepel, Matt M. Botvinick, Peter W. Battaglia |
ICLR (Poster) | 5 |
| 2017 | Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
David Krueger 0001, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Aaron C. Courville, Christopher Joseph Pal |
ICLR (Poster) | 3 |
| 2010 | A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices
Matthew Skala, Victoria Krakovna, János Kramár, Gerald Penn |
ACL | 3 |