Hamish Ivison

dblp:288/1956 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-0069-7659ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 37% Reinforcement learning · 33% Generative modeling · 16%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
preference learning
1.522024
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.522024
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024
Machine learning › Reinforcement learning › reward learning
reward modeling
1.522024
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024
Natural language and speech › Language models and text generation
instruction tuning
1.322023
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources · NeurIPS 2023
HINT: Hypernetwork Instruction Tuning for Efficient Zero- and Few-Shot Generalisation · ACL (1) 2023
Machine learning › Generative modeling › diffusion model › discrete diffusion model
diffusion language model
0.912025
TESS 2: A Large-Scale Generalist Diffusion Language Model · ACL (1) 2025
Machine learning › Generative modeling
diffusion model
0.912025
TESS 2: A Large-Scale Generalist Diffusion Language Model · ACL (1) 2025
Machine learning › Generative modeling › diffusion model › guided diffusion
inference-time guidance
0.912025
TESS 2: A Large-Scale Generalist Diffusion Language Model · ACL (1) 2025
Natural language and speech › Language models and text generation
instruction following
0.912025
Generalizing Verifiable Instruction Following · NeurIPS 2025
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards
0.912025
Generalizing Verifiable Instruction Following · NeurIPS 2025
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.812024
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.812024
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model
open language model development
0.812024
OLMo: Accelerating the Science of Language Models · ACL (1) 2024
Natural language and speech › Language models and text generation › alignment
pluralistic alignment
0.812024
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.812024
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
instruction-following evaluation
0.712023
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources · NeurIPS 2023
Natural language and speech › Language models and text generation
large language model evaluation
0.712023
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources · NeurIPS 2023
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.712023
HINT: Hypernetwork Instruction Tuning for Efficient Zero- and Few-Shot Generalisation · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

instruction tuning · 1.5reinforcement learning with verifiable rewards · 0.9cross-entropy diffusion loss · 0.9continued pretraining · 0.9constraint verification · 0.9variational preference learning · 0.8active learning · 0.8PPO · 0.8DPO · 0.8hypernetwork · 0.7
YearPublicationVenuePosition
2025 TESS 2: A Large-Scale Generalist Diffusion Language Model
abstract
We introduce TESS 2, a general instructionfollowing diffusion language model that outperforms contemporary instruction-tuned diffusion models, as well as matches and sometimes exceeds strong autoregressive (AR) models.We train TESS 2 by first adapting an AR model via continued pretraining with the usual cross-entropy as diffusion loss, and then performing further instruction tuning.We find that adaptation training as well as the choice of the base model is crucial for training good instruction-following diffusion models.Furthermore, we propose reward guidance, a novel and modular inference-time guidance procedure to align model outputs without needing to train the underlying model.Finally, we show that TESS 2 further improves with increased inference-time compute, highlighting the utility of diffusion LMs in having fine-grained controllability over the amount of compute used at inference time.Code and models are available at https://github.com/hamishivi/tess-2.
Jaesung Tae, Hamish Ivison, Sachin Kumar 0009, Arman Cohan
ACL (1)2
2025 Generalizing Verifiable Instruction Following
abstract
A crucial factor for successful human and AI interaction is the ability of language models or chatbots to follow human instructions precisely. A common feature of instructions are output constraints like only answer with yes or no" ormention the word `abracadabra' at least 3 times" that the user adds to craft a more useful answer.Even today's strongest models struggle with fulfilling such constraints. We find that most models strongly overfit on a small set of verifiable constraints from the benchmarks that test these abilities, a skill called precise instruction following, and are not able to generalize well to unseen output constraints. We introduce a new benchmark, IFBench, to evaluate precise instruction following generalization on 58 new, diverse, and challenging verifiable out-of-domain constraints. In addition, we perform an extensive analysis of how and on what data models can be trained to improve precise instruction following generalization. Specifically, we carefully design constraint verification modules and show that reinforcement learning with verifiable rewards (RLVR) significantly improves instruction following. In addition to IFBench, we release 29 additional new hand-annotated training constraints and verification functions, RLVR training prompts, and code.
Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison, Shengyi Huang, Pradeep Dasigi, Nathan Lambert 0001, Hannaneh Hajishirzi
NeurIPS4
2024 OLMo: Accelerating the Science of Language Models
abstract
Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi
ACL (1)8
2024 TESS: Text-to-Text Self-Conditioned Simplex Diffusion
abstract
Rabeeh Karimi Mahabadi, Hamish Ivison, Jaesung Tae, James Henderson, Iz Beltagy, Matthew Peters, Arman Cohan. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Rabeeh Karimi Mahabadi, Hamish Ivison, Jaesung Tae, James Henderson 0001, Iz Beltagy, Matthew E. Peters, Arman Cohan
EACL (1)2
2024 Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
abstract
Learning from preference feedback has emerged as an essential step for improving the generation quality and performance of modern language models (LMs). Despite its widespread use, the way preference-based learning is applied varies wildly, with differing data, learning algorithms, and evaluations used, making disentangling the impact of each aspect difficult. In this work, we identify four core aspects of preference-based learning: preference data, learning algorithm, reward model, and policy training prompts, systematically investigate the impact of these components on downstream model performance, and suggest a recipe for strong learning for preference feedback. Our findings indicate that all aspects are important for performance, with better preference data leading to the largest improvements, followed by the choice of learning algorithm, the use of improved reward models, and finally the use of additional unlabeled prompts for policy training. Notably, PPO outperforms DPO by up to 2.5% in math and 1.2% in general domains. High-quality preference data leads to improvements of up to 8% in instruction following and truthfulness. Despite significant gains of up to 5% in mathematical evaluation when scaling up reward models, we surprisingly observe marginal improvements in other categories.
Hamish Ivison, Yizhong Wang, Jiacheng Liu 0010, Zeqiu Wu, Valentina Pyatkin, Nathan Lambert 0001, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi
NeurIPS1
2024 Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
abstract
Reinforcement Learning from Human Feedback (RLHF) is a powerful paradigm for aligning foundation models to human values and preferences. However, current RLHF techniques cannot account for the naturally occurring differences in individual human preferences across a diverse population. When these differences arise, traditional RLHF frameworks simply average over them, leading to inaccurate rewards and poor performance for individual subgroups. To address the need for pluralistic alignment, we develop a class of multimodal RLHF methods. Our proposed techniques are based on a latent variable formulation - inferring a novel user-specific latent and learning reward models and policies conditioned on this latent without additional user-specific data. While conceptually simple, we show that in practice, this reward modeling requires careful algorithmic considerations around model architecture and reward scaling. To empirically validate our proposed technique, we first show that it can provide a way to combat underspecification in simulated control problems, inferring and optimizing user-specific reward functions. Next, we conduct experiments on pluralistic language datasets representing diverse user preferences and demonstrate improved reward function accuracy. We additionally show the benefits of this probabilistic framework in terms of measuring uncertainty, and actively learning user preferences. This work enables learning from diverse populations of users with divergent preferences, an important challenge that naturally occurs in problems from robot learning to foundation model alignment.
Sriyash Poddar, Yanming Wan, Hamish Ivison, Natasha Jaques
NeurIPS3
2023 HINT: Hypernetwork Instruction Tuning for Efficient Zero- and Few-Shot Generalisation
abstract
Recent NLP models have shown the remarkable ability to effectively generalise 'zero-shot' to new tasks using only natural language instructions as guidance.However, many of these approaches suffer from high computational costs due to their reliance on concatenating lengthy instructions with every input example, resulting in costly reprocessing of the instruction.To avoid this, we introduce Hypernetworks for INstruction Tuning (HINT), which convert task instructions and examples into parameter-efficient modules inserted into an underlying model using a pretrained text encoder, eliminating the need to include instructions in the model input.The hypernetwork in HINT also produces an encoded instruction, which we concatenate with encoded inputs during decoding to further improve performance.HINT models outperform strong state-of-theart baselines by over 10% when controlling for compute (measured in FLOPs).By converting instructions into modules, HINT models can effectively disregard the length of instructions and few-shot example inputs in terms of compute usage.As a result, HINT can enhance its performance by up to 25% by incorporating additional few-shot data, while utilizing only up to 5% more compute.This combines the strengths of parameter-efficient fine-tuning and in-context learning.We release our code publicly 1 .
Hamish Ivison, Akshita Bhagia, Yizhong Wang, Hannaneh Hajishirzi, Matthew E. Peters
ACL (1)1
2023 How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources
abstract
In this work we explore recent advances in instruction-tuning language models on a range of open instruction-following datasets. Despite recent claims that open models can be on par with state-of-the-art proprietary models, these claims are often accompanied by limited evaluation, making it difficult to compare models across the board and determine the utility of various resources. We provide a large set of instruction-tuned models from 6.7B to 65B parameters in size, trained on 12 instruction datasets ranging from manually curated (e.g., OpenAssistant) to synthetic and distilled (e.g., Alpaca) and systematically evaluate them on their factual knowledge, reasoning, multilinguality, coding, safety, and open-ended instruction following abilities through a collection of automatic, model-based, and human-based metrics. We further introduce Tülu, our best performing instruction-tuned model suite finetuned on a combination of high-quality open resources.Our experiments show that different instruction-tuning datasets can uncover or enhance specific skills, while no single dataset (or combination) provides the best performance across all evaluations. Interestingly, we find that model and human preference-based evaluations fail to reflect differences in model capabilities exposed by benchmark-based evaluations, suggesting the need for the type of systemic evaluation performed in this work. Our evaluations show that the best model in any given evaluation reaches on average 87% of ChatGPT performance, and 73% of GPT-4 performance, suggesting that further investment in building better base models and instruction-tuning data is required to close the gap. We release our instruction-tuned models, including a fully finetuned 65B Tülu, along with our code, data, and evaluation framework to facilitate future research.
Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, Dave Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, Hannaneh Hajishirzi
NeurIPS2