János Kramár

dblp:49/9013 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 52% Reinforcement learning · 14% Language models and text generation · 14%
Theoretical computer science
3 papers
Algorithmic game theory and mechanism design · 98% Logic in computer science · 2%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%

Topics — the 21 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.422024
Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024
Tracr: Compiled Transformers as a Laboratory for Interpretability · NeurIPS 2023
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
1.422024
Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024
Tracr: Compiled Transformers as a Laboratory for Interpretability · NeurIPS 2023
Machine learning › Trustworthy machine learning
AI safety
0.812024
On scalable oversight with weak LLMs judging strong LLMs · NeurIPS 2024
Natural language and speech › Language models and text generation › alignment
scalable oversight
0.812024
On scalable oversight with weak LLMs judging strong LLMs · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder
0.812024
Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
sparse feature learning
0.812024
Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024
Collaborative and social computing › cooperative work
group decision-making
0.512021
A Neural Network Auction For Group Decision Making Over a Continuous Space · IJCAI 2021
Algorithmic game theory and mechanism design › mechanism design
auction design
0.512021
A Neural Network Auction For Group Decision Making Over a Continuous Space · IJCAI 2021
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.412020
Learning to Play No-Press Diplomacy with Best Response Policy Iteration · NeurIPS 2020
Machine learning › Reinforcement learning › dynamic programming
policy iteration
0.412020
Learning to Play No-Press Diplomacy with Best Response Policy Iteration · NeurIPS 2020
Algorithmic game theory and mechanism design
equilibrium computation
0.412020
Learning to Play No-Press Diplomacy with Best Response Policy Iteration · NeurIPS 2020
Algorithmic game theory and mechanism design › learning in games
fictitious play
0.412020
Learning to Play No-Press Diplomacy with Best Response Policy Iteration · NeurIPS 2020
Machine learning › Reinforcement learning
model-based reinforcement learning
0.412019
Relational Forward Models for Multi-Agent Learning · ICLR (Poster) 2019
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning
0.412019
Relational Forward Models for Multi-Agent Learning · ICLR (Poster) 2019
Machine learning › Deep learning architectures and training
recurrent neural network
0.312017
Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations · ICLR (Poster) 2017
Machine learning › Deep learning architectures and training
regularization
0.312017
Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations · ICLR (Poster) 2017
Natural language and speech › Language models and text generation
large language model
0.212024
On scalable oversight with weak LLMs judging strong LLMs · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
transformer interpretability
0.212023
Tracr: Compiled Transformers as a Laboratory for Interpretability · NeurIPS 2023
Information retrieval
compact coding
0.112010
A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices · ACL 2010
Data mining › pattern mining › formal concept analysis
concept lattice
0.112010
A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices · ACL 2010
Logic in computer science › knowledge representation and reasoning
formal concept analysis
0.012010
A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices · ACL 2010

Methods — techniques the papers use, named apart from their topics

neural network · 1.0mechanism design · 1.0gradient ascent · 1.0deep reinforcement learning · 0.9approximate best response operator · 0.9weak-to-strong supervision · 0.8l1 penalty · 0.8gated sparse autoencoder · 0.8debate · 0.8consultancy · 0.8superposition analysis · 0.7program compilation · 0.7graph neural network · 0.4zero-preserving encoding · 0.2
YearPublicationVenuePosition
2024 On scalable oversight with weak LLMs judging strong LLMs
abstract
Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI. In this paper we study debate, where two AI's compete to convince a judge; consultancy, where a single AI tries to convince a judge that asks questions; and compare to a baseline of direct question-answering, where the judge just answers outright without the AI. We use large language models (LLMs) as both AI agents and as stand-ins for human judges, taking the judge models to be weaker than agent models. We benchmark on a diverse range of asymmetries between judges and agents, extending previous work on a single extractive QA task with information asymmetry, to also include mathematics, coding, logic and multimodal reasoning asymmetries. We find that debate outperforms consultancy across all tasks when the consultant is randomly assigned to argue for the correct/incorrect answer. Comparing debate to direct question answering, the results depend on the type of task: in extractive QA tasks with information asymmetry debate outperforms direct question answering, but in other tasks without information asymmetry the results are mixed. Previous work assigned debaters/consultants an answer to argue for. When we allow them to instead choose which answer to argue for, we find judges are less frequently convinced by the wrong answer in debate than in consultancy. Further, we find that stronger debater models increase judge accuracy, though more modestly than in previous studies.
Zachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah D. Goodman, Rohin Shah
NeurIPS3
2024 Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders
abstract
Recent work has found that sparse autoencoders (SAEs) are an effective technique for unsupervised discovery of interpretable features in language models' (LMs) activations, by finding sparse, linear reconstructions of those activations. We introduce the Gated Sparse Autoencoder (Gated SAE), which achieves a Pareto improvement over training with prevailing methods. In SAEs, the L1 penalty used to encourage sparsity introduces many undesirable biases, such as shrinkage -- systematic underestimation of feature activations. The key insight of Gated SAEs is to separate the functionality of (a) determining which directions to use and (b) estimating the magnitudes of those directions: this enables us to apply the L1 penalty only to the former, limiting the scope of undesirable side effects. Through training SAEs on LMs of up to 7B parameters we find that, in typical hyper-parameter ranges, Gated SAEs solve shrinkage, are similarly interpretable, and require half as many firing features to achieve comparable reconstruction fidelity.
Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah, Neel Nanda
NeurIPS6
2023 Tracr: Compiled Transformers as a Laboratory for Interpretability
abstract
We show how to "compile" human-readable programs into standard decoder-only transformer models. Our compiler, Tracr, generates models with known structure. This structure can be used to design experiments. For example, we use it to study "superposition" in transformers that execute multi-step algorithms. Additionally, the known structure of Tracr-compiled models can serve as _ground-truth_ for evaluating interpretability methods. Commonly, because the "programs" learned by transformers are unknown it is unclear whether an interpretation succeeded. We demonstrate our approach by implementing and examining programs including computing token frequencies, sorting, and parenthesis checking. We provide an open-source implementation of Tracr at https://github.com/google-deepmind/tracr.
David Lindner, János Kramár, Sebastian Farquhar, Matthew Rahtz, Thomas McGrath 0001, Vladimir Mikulik
NeurIPS2
2021 A Neural Network Auction For Group Decision Making Over a Continuous Space
abstract
We propose a system for conducting an auction over locations in a continuous space. It enables participants to express their preferences over possible choices of location in the space, selecting the location that maximizes the total utility of all agents. We prevent agents from tricking the system into selecting a location that improves their individual utility at the expense of others by using a pricing rule that gives agents no incentive to misreport their true preferences. The system queries participants for their utility in many random locations, then trains a neural network to approximate the preference function of each participant. The parameters of these neural network models are transmitted and processed by the auction mechanism, which composes these into differentiable models that are optimized through gradient ascent to compute the final chosen location and charged prices.
Yoram Bachrach, Ian Gemp, Marta Garnelo, János Kramár, Tom Eccles, Dan Rosenbaum, Thore Graepel
IJCAI4
2020 Learning to Play No-Press Diplomacy with Best Response Policy Iteration
abstract
Recent advances in deep reinforcement learning (RL) have led to considerable progress in many 2-player zero-sum games, such as Go, Poker and Starcraft. The purely adversarial nature of such games allows for conceptually simple and principled application of RL methods. However real-world settings are many-agent, and agent interactions are complex mixtures of common-interest and competitive aspects. We consider Diplomacy, a 7-player board game designed to accentuate dilemmas resulting from many-agent interactions. It also features a large combinatorial action space and simultaneous moves, which are challenging for RL algorithms. We propose a simple yet effective approximate best response operator, designed to handle large combinatorial action spaces and simultaneous moves. We also introduce a family of policy iteration methods that approximate fictitious play. With these methods, we successfully apply RL to Diplomacy: we show that our agents convincingly outperform the previous state-of-the-art, and game theoretic equilibrium analysis shows that the new process yields consistent improvements.
Thomas W. Anthony 0001, Tom Eccles, Andrea Tacchetti, János Kramár, Ian Gemp, Thomas C. Hudson, Nicolas Porcel, Marc Lanctot, Julien Pérolat, Richard Everett 0001, Satinder Singh 0001, Thore Graepel, Yoram Bachrach
NeurIPS4
2019 Relational Forward Models for Multi-Agent Learning
Andrea Tacchetti, H. Francis Song, Pedro A. M. Mediano, Vinícius Flores Zambaldi, János Kramár, Neil C. Rabinowitz, Thore Graepel, Matt M. Botvinick, Peter W. Battaglia
ICLR (Poster)5
2017 Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
David Krueger 0001, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Aaron C. Courville, Christopher Joseph Pal
ICLR (Poster)3
2010 A Generalized-Zero-Preserving Method for Compact Encoding of Concept Lattices
Matthew Skala, Victoria Krakovna, János Kramár, Gerald Penn
ACL3