Andreas Köpf

dblp:255/4763 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 57% Reinforcement learning · 35% Deep learning architectures and training · 8%
Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 100%
Human-computer interaction and pervasive computing
1 paper
Learning and educational technologies · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
procedural task generation
0.912025
Reasoning Gym: Reasoning Environments for Reinforcement Learning with Verifiable Rewards · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model
reasoning model
0.912025
Reasoning Gym: Reasoning Environments for Reinforcement Learning with Verifiable Rewards · NeurIPS 2025
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards
0.912025
Reasoning Gym: Reasoning Environments for Reinforcement Learning with Verifiable Rewards · NeurIPS 2025
Natural language and speech › Language models and text generation
alignment
0.712023
OpenAssistant Conversations - Democratizing Large Language Model Alignment · NeurIPS 2023
Natural language and speech › Language models and text generation › alignment › preference alignment
human feedback alignment
0.712023
OpenAssistant Conversations - Democratizing Large Language Model Alignment · NeurIPS 2023
Natural language and speech › Language models and text generation
instruction tuning
0.712023
OpenAssistant Conversations - Democratizing Large Language Model Alignment · NeurIPS 2023
Machine learning › Deep learning architectures and training › deep learning systems
deep learning framework
0.412019
PyTorch: An Imperative Style, High-Performance Deep Learning Library · NeurIPS 2019
GPUs and heterogeneous computing
GPU computing
0.112019
PyTorch: An Imperative Style, High-Performance Deep Learning Library · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.7procedural generation · 1.7supervised fine-tuning · 0.7reinforcement learning from human feedback · 0.7
YearPublicationVenuePosition
2025 Reasoning Gym: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
abstract
We introduce Reasoning Gym, a library of reasoning environments for reinforcement learning with verifiable rewards (RLVR). It provides over 100 tasks spanning multiple domains including algebra, arithmetic, computation, cognition, geometry, graph theory, logic, and various common games. Its key innovation is the ability to generate virtually infinite training data with adjustable complexity, unlike most previous reasoning datasets, which are typically fixed. This procedural generation approach allows for continuous evaluation across varying difficulty levels and task configurations. Our experimental results demonstrate the efficacy of Reasoning Gym in both evaluating and reinforcement learning of reasoning models.
Zafir Stojanovski, Oliver Stanley, Joe Sharratt, Abdulhakeem Adefioye, Jean Kaddour, Andreas Köpf
NeurIPS7
2023 OpenAssistant Conversations - Democratizing Large Language Model Alignment
abstract
Aligning large language models (LLMs) with human preferences has proven to drastically improve usability and has driven rapid adoption as demonstrated by ChatGPT.Alignment techniques such as supervised fine-tuning (\textit{SFT}) and reinforcement learning from human feedback (\textit{RLHF}) greatly reduce the required skill and domain knowledge to effectively harness the capabilities of LLMs, increasing their accessibility and utility across various domains.However, state-of-the-art alignment techniques like \textit{RLHF} rely on high-quality human feedback data, which is expensive to create and often remains proprietary.In an effort to democratize research on large-scale alignment, we release OpenAssistant Conversations, a human-generated, human-annotated assistant-style conversation corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292 quality ratings, resulting in over 10,000 complete and fully annotated conversation trees.The corpus is a product of a worldwide crowd-sourcing effort involving over 13,500 volunteers.Models trained on OpenAssistant Conversations show consistent improvements on standard benchmarks over respective base models.We release our code\footnote{\git} and data\footnote{\data} under a fully permissive licence.
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, Alexander Mattick
NeurIPS1
2019 PyTorch: An Imperative Style, High-Performance Deep Learning Library
abstract
Deep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it was designed from first principles to support an imperative and Pythonic programming style that supports code as a model, makes debugging easy and is consistent with other popular scientific computing libraries, while remaining efficient and supporting hardware accelerators such as GPUs. In this paper, we detail the principles that drove the implementation of PyTorch and how they are reflected in its architecture. We emphasize that every aspect of PyTorch is a regular Python program under the full control of its user. We also explain how the careful and pragmatic implementation of the key components of its runtime enables them to work together to achieve compelling performance. We demonstrate the efficiency of individual subsystems, as well as the overall speed of PyTorch on several commonly used benchmarks.
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Soumith Chintala
NeurIPS12