Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Akbir Khan

dblp:232/1874 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Language models and text generation · 37% Reinforcement learning · 26% Multi-agent systems · 11%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 19 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.622025
Language Models Learn to Mislead Humans via RLHF · ICLR 2025
Debating with More Persuasive LLMs Leads to More Truthful Answers · ICML 2024
Natural language and speech › Language models and text generation
large language model
1.622025
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games · ICLR 2025
Debating with More Persuasive LLMs Leads to More Truthful Answers · ICML 2024
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.522024
Melting Pot Contest: Charting the Future of Generalized Cooperative Intelligence · NeurIPS 2024
JaxMARL: Multi-Agent RL Environments and Algorithms in JAX · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems › agentic AI
agentic reasoning
0.912025
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games · ICLR 2025
Natural language and speech › Language models and text generation
LLM agents
0.912025
Factorio Learning Environment · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
long-term planning
0.912025
Factorio Learning Environment · NeurIPS 2025
Security and privacy of machine learning › adversarial attack
backdoor attack
0.912025
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats · ICLR 2025
Machine learning › Reinforcement learning › multi-agent reinforcement learning › markov games
mixed-motive games
0.812024
Melting Pot Contest: Charting the Future of Generalized Cooperative Intelligence · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent environments
0.812024
JaxMARL: Multi-Agent RL Environments and Algorithms in JAX · NeurIPS 2024
Human-AI interaction
AI-assisted decision-making
0.812024
Debating with More Persuasive LLMs Leads to More Truthful Answers · ICML 2024
Machine learning › Reinforcement learning › reinforcement learning environment
environment design
0.712023
MAESTRO: Open-Ended Environment Design for Multi-Agent Reinforcement Learning · ICLR 2023
Natural language and speech › Language models and text generation
instruction tuning
0.712023
The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs · NeurIPS 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › pragmatics
pragmatic language understanding
0.712023
The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs · NeurIPS 2023
Natural language and speech › Language models and text generation
code generation
0.312025
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats · ICLR 2025
Natural language and speech › Language models and text generation › alignment
deceptive alignment
0.312025
Language Models Learn to Mislead Humans via RLHF · ICLR 2025
Machine learning › Trustworthy machine learning
robustness
0.312025
Language Models Learn to Mislead Humans via RLHF · ICLR 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
0.312025
Factorio Learning Environment · NeurIPS 2025
Computer vision › Vision and language
vision-language model
0.312025
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games · ICLR 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
open-world learning
0.212023
MAESTRO: Open-Ended Environment Design for Multi-Agent Reinforcement Learning · ICLR 2023

Methods — techniques the papers use, named apart from their topics

reinforcement learning from human feedback · 1.7probing · 1.7large language model · 1.7control evaluation · 1.7adaptive risk estimation · 1.7persuasion optimization · 1.5debate · 1.5multi-agent reinforcement learning · 1.4red-teaming · 0.9red teaming · 0.9benchmark evaluation · 0.9
YearPublicationVenuePosition
2025 BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
abstract
Large Language Models (LLMs) and Vision Language Models (VLMs) possess extensive knowledge and exhibit promising reasoning abilities, however, they still struggle to perform well in complex, dynamic environments. Real-world tasks require handling intricate interactions, advanced spatial reasoning, long-term planning, and continuous exploration of new strategies—areas in which we lack effective methodologies for comprehensively evaluating these capabilities. To address this gap, we introduce BALROG, a novel benchmark designed to assess the agentic capabilities of LLMs and VLMs through a diverse set of challenging games. Our benchmark incorporates a range of existing reinforcement learning environments with varying levels of difficulty, including tasks that are solvable by non-expert humans in seconds to extremely challenging ones that may take years to master (e.g., the NetHack Learning Environment). We devise fine-grained metrics to measure performance and conduct an extensive evaluation of several popular open-source and closed-source LLMs and VLMs. Our findings indicate that while current models achieve partial success in the easier games, they struggle significantly with more challenging tasks. Notably, we observe severe deficiencies in vision-based decision-making, as several models perform worse when visual representations of the environments are provided. We release BALROG as an open and user-friendly benchmark to facilitate future research and development in the agentic community. Code and Leaderboard at balrogai.com
Davide Paglieri, Bartlomiej Cupial, Samuel Coward, Ulyana Piterbarg, Maciej Wolczyk, Akbir Khan, Eduardo Pignatelli, Lukasz Kucinski, Lerrel Pinto, Rob Fergus, Jakob N. Foerster, Jack Parker-Holder, Tim Rocktäschel
ICLR6
2025 Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
abstract
As large language models (LLMs) grow more powerful, they also become more difficult to trust. They could be either aligned with human intentions, or exhibit "subversive misalignment" -- introducing subtle errors that bypass safety checks. Although individual errors may not immediately cause harm, each increases the risk of an eventual safety failure. With this uncertainty, model deployment often grapples with the tradeoff between ensuring safety and harnessing the capabilities of untrusted models. In this work, we introduce the ``Diffuse Risk Management'' problem, aiming to balance the average-case safety and usefulness in the deployment of untrusted models over a large sequence of tasks. We approach this problem by developing a two-level framework: the single-task level (micro-protocol) and the whole-scenario level (macro-protocol). At the single-task level, we develop various \textit{micro}-protocols that use a less capable, but extensively tested (trusted) model to harness and monitor the untrusted model. At the whole-scenario level, we find an optimal \textit{macro}-protocol that uses an adaptive estimate of the untrusted model's risk to choose between micro-protocols. To evaluate the robustness of our method, we follow \textit{control evaluations} in a code generation testbed, which involves a red team attempting to generate subtly backdoored code with an LLM whose deployment is safeguarded by a blue team. Experiment results show that our approach retains 99.6\% usefulness of the untrusted model while ensuring near-perfect safety, significantly outperforming existing deployment methods. Our approach also demonstrates robustness when the trusted and untrusted models have a large capability gap. Our findings demonstrate the promise of managing diffuse risks in the deployment of increasingly capable but untrusted LLMs.
Jiaxin Wen, Vivek Hebbar, Caleb Larson, Aryan Bhatt, Ansh Radhakrishnan, Mrinank Sharma, Henry Sleight, Shi Feng 0005, He He 0001, Ethan Perez, Buck Shlegeris, Akbir Khan
ICLR12
2025 Language Models Learn to Mislead Humans via RLHF
abstract
Language models (LMs) can produce errors that are hard to detect for humans, especially when the task is complex. RLHF, the most popular post-training method, may exacerbate this problem: to achieve higher rewards, LMs might get better at convincing humans that they are right even when they are wrong. We study this phenomenon under a standard RLHF pipeline, calling it ``U-Sophistry'' since it is \textbf{U}nintended by model developers. Specifically, we ask time-constrained (e.g., 3-10 minutes) human subjects to evaluate the correctness of model outputs and calculate humans' accuracy against gold labels. On a question-answering task (QuALITY) and programming task (APPS), RLHF makes LMs better at convincing our subjects but not at completing the task correctly. RLHF also makes the model harder to evaluate: our subjects' false positive rate increases by 24.1% on QuALITY and 18.3% on APPS. Finally, we show that probing, a state-of-the-art approach for detecting \textbf{I}ntended Sophistry (e.g.~backdoored LMs), does not generalize to U-Sophistry. Our results highlight an important failure mode of RLHF and call for more research in assisting humans to align them.
Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, He He 0001, Shi Feng 0005
ICLR3
2025 Factorio Learning Environment
abstract
Large Language Models (LLMs) are rapidly saturating existing benchmarks, necessitating new open-ended evaluations. We introduce the Factorio Learning Environment (FLE), based on the game of Factorio, that tests agents in long-term planning, spatial reasoning, program synthesis, and resource optimization. FLE provides exponentially scaling challenges -- from basic automation to complex factories processing millions of resource units per second. We provide two settings: (1) open-play with the open-ended task of building the largest factory on an procedurally generated map and (2) lab-play consisting of 33 bounded tasks accross three settings with fixed resources. We demonstrate across both settings that models still lack strong spatial reasoning. In lab-play, we find that LLMs exhibit promising short-horizon skills, yet are unable to operate effectively in constrained environments, reflecting limitations in error analysis. In open-play, while LLMs discover automation strategies that improve growth (e.g electric-powered drilling), they fail to achieve complex automation (e.g electronic-circuit manufacturing)
Jack Hopkins, Mart Bakler, Akbir Khan
NeurIPS3
2024 Debating with More Persuasive LLMs Leads to More Truthful Answers
abstract
Common methods for aligning large language models (LLMs) with desired behaviour heavily rely on human-labelled data. However, as models grow increasingly sophisticated, they will surpass human expertise, and the role of human evaluation will evolve into non-experts overseeing experts. In anticipation of this, we ask: can weaker models assess the correctness of stronger models? We investigate this question in an analogous setting, where stronger models (experts) possess the necessary information to answer questions and weaker models (non-experts) lack this information. The method we evaluate is debate, where two LLM experts each argue for a different answer, and a non-expert selects the answer. We find that debate consistently helps both non-expert models and humans answer questions, achieving 76% and 88% accuracy respectively (naive baselines obtain 48% and 60%). Furthermore, optimising expert debaters for persuasiveness in an unsupervised manner improves non-expert ability to identify the truth in debates. Our results provide encouraging empirical evidence for the viability of aligning models with debate in the absence of ground truth.
Akbir Khan, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, Ethan Perez
ICML1
2024 JaxMARL: Multi-Agent RL Environments and Algorithms in JAX
abstract
Benchmarks are crucial in the development of machine learning algorithms, significantly influencing reinforcement learning (RL) research through the available environments. Traditionally, RL environments run on the CPU, which limits their scalability with the computational resources typically available in academia. However, recent advancements in JAX have enabled the wider use of hardware acceleration, enabling massively parallel RL training pipelines and environments. While this has been successfully applied to single-agent RL, it has not yet been widely adopted for multi-agent scenarios. In this paper, we present JaxMARL, the first open-source, easy-to-use code base that combines GPU-enabled efficiency with support for a large number of commonly used MARL environments and popular baseline algorithms. Our experiments show that, in terms of wall clock time, our JAX-based training pipeline is up to 12,500 times faster than existing approaches. This enables efficient and thorough evaluations, potentially alleviating the evaluation crisis in the field. We also introduce and benchmark SMAX, a vectorised, simplified version of the popular StarCraft Multi-Agent Challenge, which removes the need to run the StarCraft II game engine. This not only enables GPU acceleration, but also provides a more flexible MARL environment, unlocking the potential for self-play, meta-learning, and other future applications in MARL. The code is available at https://github.com/flairox/jaxmarl.
Alex Rutherford, Benjamin Ellis, Matteo Gallici, Jonathan Cook 0004, Andrei Lupu, Garðar Ingvarsson, Timon Willi, Ravi Hammond, Akbir Khan, Christian Schröder de Witt, Alexandra Souly, Saptarashmi Bandyopadhyay, Mikayel Samvelyan, Minqi Jiang, Robert T. Lange, Shimon Whiteson, Bruno Lacerda, Nick Hawes, Tim Rocktäschel, Chris Lu 0001, Jakob N. Foerster
NeurIPS9
2024 Melting Pot Contest: Charting the Future of Generalized Cooperative Intelligence
abstract
Multi-agent AI research promises a path to develop human-like and human-compatible intelligent technologies that complement the solipsistic view of other approaches, which mostly do not consider interactions between agents. Aiming to make progress in this direction, the Melting Pot contest 2023 focused on the problem of cooperation among interacting agents and challenged researchers to push the boundaries of multi-agent reinforcement learning (MARL) for mixed-motive games. The contest leveraged the Melting Pot environment suite to rigorously evaluate how well agents can adapt their cooperative skills to interact with novel partners in unforeseen situations. Unlike other reinforcement learning challenges, this challenge focused on social rather than environmental generalization. In particular, a population of agents performs well in Melting Pot when its component individuals are adept at finding ways to cooperate both with others in their population and with strangers. Thus Melting Pot measures cooperative intelligence.The contest attracted over 600 participants across 100+ teams globally and was a success on multiple fronts: (i) it contributed to our goal of pushing the frontiers of MARL towards building more cooperatively intelligent agents, evidenced by several submissions that outperformed established baselines; (ii) it attracted a diverse range of participants, from independent researchers to industry affiliates and academic labs, both with strong background and new interest in the area alike, broadening the field’s demographic and intellectual diversity; and (iii) analyzing the submitted agents provided important insights, highlighting areas for improvement in evaluating agents' cooperative intelligence. This paper summarizes the design aspects and results of the contest and explores the potential of Melting Pot as a benchmark for studying Cooperative AI. We further analyze the top solutions and conclude with a discussion on promising directions for future research.
Rakshit S. Trivedi, Akbir Khan, Jesse Clifton, Lewis Hammond, Edgar A. Duéñez-Guzmán, Dipam Chakraborty, John P. Agapiou, Jayd Matyas, Alexander Vezhnevets, Barna Pásztor, Yunke Ao, Omar G. Younis, Benjamin Swain, Haoyuan Qin, Mian Deng, Ziwei Deng, Utku Erdoganaras, Yue Zhao 0023, Marko Tesic, Natasha Jaques, Jakob N. Foerster, Vincent Conitzer, José Hernández-Orallo, Dylan Hadfield-Menell, Joel Z. Leibo
NeurIPS2
2023 MAESTRO: Open-Ended Environment Design for Multi-Agent Reinforcement Learning
Mikayel Samvelyan, Akbir Khan, Michael Dennis 0001, Minqi Jiang, Jack Parker-Holder, Jakob N. Foerster, Roberta Raileanu, Tim Rocktäschel
ICLR2
2023 The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs
abstract
Despite widespread use of LLMs as conversational agents, evaluations of performance fail to capture a crucial aspect of communication: interpreting language in context---incorporating its pragmatics. Humans interpret language using beliefs and prior knowledge about the world. For example, we intuitively understand the response "I wore gloves" to the question "Did you leave fingerprints?" as meaning "No". To investigate whether LLMs have the ability to make this type of inference, known as an implicature, we design a simple task and evaluate four categories of widely used state-of-the-art models. We find that, despite only evaluating on utterances that require a binary inference (yes or no), models in three of these categories perform close to random. However, LLMs instruction-tuned at the example-level perform significantly better. These results suggest that certain fine-tuning strategies are far better at inducing pragmatic understanding in models. We present our findings as the starting point for further research into evaluating how LLMs interpret language in context and to drive the development of more pragmatic and useful models of human discourse.
Laura Ruis, Akbir Khan, Stella Biderman, Sara Hooker, Tim Rocktäschel, Edward Grefenstette
NeurIPS2
2021 Multi-dimensional Affect in Poetry (POCA) Dataset: Acquisition, Annotation and Baseline Results
abstract
Detecting emotions and affect in text has received enormous attention in recent years, and yet majority of the works in this area reduce the nuanced emotional responses into ‘positive’, ‘negative’ and ‘neutral’. In this paper, we introduce a novel multi-dimensional affect in poetry (POCA) dataset for sentiment analysis annotated using the Geneva Emotion Wheel (GEW), to capture and analyse the multi-dimensional affect evoked in listeners. The POCA dataset is based on poems and their corresponding recitals from an online poetry database where recitals are curated by the website, and performed by the poet or an approved artist. The POCA dataset contains 330 poems (text and audio), from the English language, each of which is annotated across 20 different emotion classes, by 5 listeners. A subset of the dataset (50 poems) have also been annotated by an inlab study by 3 listeners each while their Electrodermal activity (EDA) was being recorded. As a proof of concept, we (i) introduce representative problem formulations to be addressed by machine learning approaches using the POCA dataset, from single emotion recognition (e.g., does this poem evoke joy?) to continuous affect prediction (e.g., what level arousal and valence does this poem evoke?), (ii) provide baseline results for text-based affect recognition using several classification and regression models, and (iii) provide baseline results for EDA-based affect prediction. Our results show that (i) for text-based affect recognition, classical approaches can provide as accurate results as their fine-tuned neural network counterparts, and (ii) in EDA-based affect prediction, in general there is a strong relation between the EDA signals and the self-reported valence and arousal quadrants, while predictions are better for arousal than valence.
Akbir Khan, Jack Hopkins, Hatice Gunes
ACII1