VLDB 2026 Research / reviewers in the wild / expert
Hamish Ivison
dblp:288/1956
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-0069-7659ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 37% Reinforcement learning · 33% Generative modeling · 16% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
preference learning |
1.5 | 2 | 2024 | Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024 Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
1.5 | 2 | 2024 | Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024 Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
1.5 | 2 | 2024 | Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024 Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024 |
Natural language and speech › Language models and text generation
instruction tuning |
1.3 | 2 | 2023 | How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources · NeurIPS 2023 HINT: Hypernetwork Instruction Tuning for Efficient Zero- and Few-Shot Generalisation · ACL (1) 2023 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
diffusion language model |
0.9 | 1 | 2025 | TESS 2: A Large-Scale Generalist Diffusion Language Model · ACL (1) 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | TESS 2: A Large-Scale Generalist Diffusion Language Model · ACL (1) 2025 |
Machine learning › Generative modeling › diffusion model › guided diffusion
inference-time guidance |
0.9 | 1 | 2025 | TESS 2: A Large-Scale Generalist Diffusion Language Model · ACL (1) 2025 |
Natural language and speech › Language models and text generation
instruction following |
0.9 | 1 | 2025 | Generalizing Verifiable Instruction Following · NeurIPS 2025 |
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards |
0.9 | 1 | 2025 | Generalizing Verifiable Instruction Following · NeurIPS 2025 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.8 | 1 | 2024 | Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.8 | 1 | 2024 | Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model
open language model development |
0.8 | 1 | 2024 | OLMo: Accelerating the Science of Language Models · ACL (1) 2024 |
Natural language and speech › Language models and text generation › alignment
pluralistic alignment |
0.8 | 1 | 2024 | Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.8 | 1 | 2024 | Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
instruction-following evaluation |
0.7 | 1 | 2023 | How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources · NeurIPS 2023 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.7 | 1 | 2023 | How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources · NeurIPS 2023 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.7 | 1 | 2023 | HINT: Hypernetwork Instruction Tuning for Efficient Zero- and Few-Shot Generalisation · ACL (1) 2023 |
Methods — techniques the papers use, named apart from their topics
instruction tuning · 1.5reinforcement learning with verifiable rewards · 0.9cross-entropy diffusion loss · 0.9continued pretraining · 0.9constraint verification · 0.9variational preference learning · 0.8active learning · 0.8PPO · 0.8DPO · 0.8hypernetwork · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TESS 2: A Large-Scale Generalist Diffusion Language ModelabstractWe introduce TESS 2, a general instructionfollowing diffusion language model that outperforms contemporary instruction-tuned diffusion models, as well as matches and sometimes exceeds strong autoregressive (AR) models.We train TESS 2 by first adapting an AR model via continued pretraining with the usual cross-entropy as diffusion loss, and then performing further instruction tuning.We find that adaptation training as well as the choice of the base model is crucial for training good instruction-following diffusion models.Furthermore, we propose reward guidance, a novel and modular inference-time guidance procedure to align model outputs without needing to train the underlying model.Finally, we show that TESS 2 further improves with increased inference-time compute, highlighting the utility of diffusion LMs in having fine-grained controllability over the amount of compute used at inference time.Code and models are available at https://github.com/hamishivi/tess-2. Jaesung Tae, Hamish Ivison, Sachin Kumar 0009, Arman Cohan |
ACL (1) | 2 |
| 2025 | Generalizing Verifiable Instruction FollowingabstractA crucial factor for successful human and AI interaction is the ability of language models or chatbots to follow human instructions precisely. A common feature of instructions are output constraints like only answer with yes or no" ormention the word `abracadabra' at least 3 times" that the user adds to craft a more useful answer.Even today's strongest models struggle with fulfilling such constraints. We find that most models strongly overfit on a small set of verifiable constraints from the benchmarks that test these abilities, a skill called precise instruction following, and are not able to generalize well to unseen output constraints. We introduce a new benchmark, IFBench, to evaluate precise instruction following generalization on 58 new, diverse, and challenging verifiable out-of-domain constraints. In addition, we perform an extensive analysis of how and on what data models can be trained to improve precise instruction following generalization. Specifically, we carefully design constraint verification modules and show that reinforcement learning with verifiable rewards (RLVR) significantly improves instruction following. In addition to IFBench, we release 29 additional new hand-annotated training constraints and verification functions, RLVR training prompts, and code. Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison, Shengyi Huang, Pradeep Dasigi, Nathan Lambert 0001, Hannaneh Hajishirzi |
NeurIPS | 4 |
| 2024 | OLMo: Accelerating the Science of Language ModelsabstractDirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi |
ACL (1) | 8 |
| 2024 | TESS: Text-to-Text Self-Conditioned Simplex DiffusionabstractRabeeh Karimi Mahabadi, Hamish Ivison, Jaesung Tae, James Henderson, Iz Beltagy, Matthew Peters, Arman Cohan. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Rabeeh Karimi Mahabadi, Hamish Ivison, Jaesung Tae, James Henderson 0001, Iz Beltagy, Matthew E. Peters, Arman Cohan |
EACL (1) | 2 |
| 2024 | Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference FeedbackabstractLearning from preference feedback has emerged as an essential step for improving the generation quality and performance of modern language models (LMs). Despite its widespread use, the way preference-based learning is applied varies wildly, with differing data, learning algorithms, and evaluations used, making disentangling the impact of each aspect difficult. In this work, we identify four core aspects of preference-based learning: preference data, learning algorithm, reward model, and policy training prompts, systematically investigate the impact of these components on downstream model performance, and suggest a recipe for strong learning for preference feedback. Our findings indicate that all aspects are important for performance, with better preference data leading to the largest improvements, followed by the choice of learning algorithm, the use of improved reward models, and finally the use of additional unlabeled prompts for policy training. Notably, PPO outperforms DPO by up to 2.5% in math and 1.2% in general domains. High-quality preference data leads to improvements of up to 8% in instruction following and truthfulness. Despite significant gains of up to 5% in mathematical evaluation when scaling up reward models, we surprisingly observe marginal improvements in other categories. Hamish Ivison, Yizhong Wang, Jiacheng Liu 0010, Zeqiu Wu, Valentina Pyatkin, Nathan Lambert 0001, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi |
NeurIPS | 1 |
| 2024 | Personalizing Reinforcement Learning from Human Feedback with Variational Preference LearningabstractReinforcement Learning from Human Feedback (RLHF) is a powerful paradigm for aligning foundation models to human values and preferences. However, current RLHF techniques cannot account for the naturally occurring differences in individual human preferences across a diverse population. When these differences arise, traditional RLHF frameworks simply average over them, leading to inaccurate rewards and poor performance for individual subgroups. To address the need for pluralistic alignment, we develop a class of multimodal RLHF methods. Our proposed techniques are based on a latent variable formulation - inferring a novel user-specific latent and learning reward models and policies conditioned on this latent without additional user-specific data. While conceptually simple, we show that in practice, this reward modeling requires careful algorithmic considerations around model architecture and reward scaling. To empirically validate our proposed technique, we first show that it can provide a way to combat underspecification in simulated control problems, inferring and optimizing user-specific reward functions. Next, we conduct experiments on pluralistic language datasets representing diverse user preferences and demonstrate improved reward function accuracy. We additionally show the benefits of this probabilistic framework in terms of measuring uncertainty, and actively learning user preferences. This work enables learning from diverse populations of users with divergent preferences, an important challenge that naturally occurs in problems from robot learning to foundation model alignment. Sriyash Poddar, Yanming Wan, Hamish Ivison, Natasha Jaques |
NeurIPS | 3 |
| 2023 | HINT: Hypernetwork Instruction Tuning for Efficient Zero- and Few-Shot GeneralisationabstractRecent NLP models have shown the remarkable ability to effectively generalise 'zero-shot' to new tasks using only natural language instructions as guidance.However, many of these approaches suffer from high computational costs due to their reliance on concatenating lengthy instructions with every input example, resulting in costly reprocessing of the instruction.To avoid this, we introduce Hypernetworks for INstruction Tuning (HINT), which convert task instructions and examples into parameter-efficient modules inserted into an underlying model using a pretrained text encoder, eliminating the need to include instructions in the model input.The hypernetwork in HINT also produces an encoded instruction, which we concatenate with encoded inputs during decoding to further improve performance.HINT models outperform strong state-of-theart baselines by over 10% when controlling for compute (measured in FLOPs).By converting instructions into modules, HINT models can effectively disregard the length of instructions and few-shot example inputs in terms of compute usage.As a result, HINT can enhance its performance by up to 25% by incorporating additional few-shot data, while utilizing only up to 5% more compute.This combines the strengths of parameter-efficient fine-tuning and in-context learning.We release our code publicly 1 . Hamish Ivison, Akshita Bhagia, Yizhong Wang, Hannaneh Hajishirzi, Matthew E. Peters |
ACL (1) | 1 |
| 2023 | How Far Can Camels Go? Exploring the State of Instruction Tuning on Open ResourcesabstractIn this work we explore recent advances in instruction-tuning language models on a range of open instruction-following datasets. Despite recent claims that open models can be on par with state-of-the-art proprietary models, these claims are often accompanied by limited evaluation, making it difficult to compare models across the board and determine the utility of various resources. We provide a large set of instruction-tuned models from 6.7B to 65B parameters in size, trained on 12 instruction datasets ranging from manually curated (e.g., OpenAssistant) to synthetic and distilled (e.g., Alpaca) and systematically evaluate them on their factual knowledge, reasoning, multilinguality, coding, safety, and open-ended instruction following abilities through a collection of automatic, model-based, and human-based metrics. We further introduce Tülu, our best performing instruction-tuned model suite finetuned on a combination of high-quality open resources.Our experiments show that different instruction-tuning datasets can uncover or enhance specific skills, while no single dataset (or combination) provides the best performance across all evaluations. Interestingly, we find that model and human preference-based evaluations fail to reflect differences in model capabilities exposed by benchmark-based evaluations, suggesting the need for the type of systemic evaluation performed in this work. Our evaluations show that the best model in any given evaluation reaches on average 87% of ChatGPT performance, and 73% of GPT-4 performance, suggesting that further investment in building better base models and instruction-tuning data is required to close the gap. We release our instruction-tuned models, including a fully finetuned 65B Tülu, along with our code, data, and evaluation framework to facilitate future research. Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, Dave Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, Hannaneh Hajishirzi |
NeurIPS | 2 |