Jörg K. H. Franke

dblp:251/8540 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-4390-4582ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Deep learning architectures and training · 62% Probabilistic and Bayesian machine learning · 18% Reinforcement learning · 8%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 50% Performance modeling and evaluation · 50%

Topics — the 23 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
transformer
1.422025
Learning in Compact Spaces with Approximately Normalized Transformer · NeurIPS 2025
Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule Design · NeurIPS 2022
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics › RNA structure prediction
RNA folding
1.022025
KinPFN: Bayesian Approximation of RNA Folding Kinetics using Prior-Data Fitted Networks · ICLR 2025
Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule Design · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian approximation
0.912025
KinPFN: Bayesian Approximation of RNA Folding Kinetics using Prior-Data Fitted Networks · ICLR 2025
Machine learning › Deep learning architectures and training
data augmentation
0.912025
Beyond Random Augmentations: Pretraining with Hard Views · ICLR 2025
Machine learning › Deep learning architectures and training › recurrent neural network
linear recurrent neural network
0.912025
Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues · ICLR 2025
Machine learning › Deep learning architectures and training
normalization
0.912025
Learning in Compact Spaces with Approximately Normalized Transformer · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › neural processes
prior-data fitted networks
0.912025
KinPFN: Bayesian Approximation of RNA Folding Kinetics using Prior-Data Fitted Networks · ICLR 2025
Machine learning › Deep learning architectures and training
recurrent neural network
0.912025
Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues · ICLR 2025
Machine learning › Deep learning architectures and training
scaling laws
0.912025
Learning in Compact Spaces with Approximately Normalized Transformer · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.912025
Beyond Random Augmentations: Pretraining with Hard Views · ICLR 2025
Machine learning › Deep learning architectures and training › sequence modeling
state tracking
0.912025
Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues · ICLR 2025
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA folding kinetics
0.912025
KinPFN: Bayesian Approximation of RNA Folding Kinetics using Prior-Data Fitted Networks · ICLR 2025
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA structure prediction
0.912025
KinPFN: Bayesian Approximation of RNA Folding Kinetics using Prior-Data Fitted Networks · ICLR 2025
Machine learning › Deep learning architectures and training
regularization
0.812024
Improving Deep Learning Optimization through Constrained Parameter Regularization · NeurIPS 2024
Machine learning › Deep learning architectures and training › regularization › norm-based regularization
weight decay
0.812024
Improving Deep Learning Optimization through Constrained Parameter Regularization · NeurIPS 2024
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA design
0.812024
Partial RNA design · Bioinform. 2024
Performance modeling and evaluation
benchmarking
0.812024
HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models · NeurIPS 2024
Hardware accelerators and domain-specific architectures › neural architecture search
hardware-aware neural architecture search
0.812024
HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.612022
Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule Design · NeurIPS 2022
Machine learning › Reinforcement learning
automated reinforcement learning
0.512021
Sample-Efficient Automated Deep Reinforcement Learning · ICLR 2021
Mathematical optimization › constrained optimization
augmented lagrangian method
0.212024
Improving Deep Learning Optimization through Constrained Parameter Regularization · NeurIPS 2024
Mathematical optimization
constrained optimization
0.212024
Improving Deep Learning Optimization through Constrained Parameter Regularization · NeurIPS 2024
Bioinformatics and computational biology › molecular informatics
molecular design
0.212022
Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule Design · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

prior-data fitted networks · 1.7bayesian inference · 1.7weight norm constraint · 0.9survey · 0.9semi-structured interviews · 0.9scalar multiplication · 0.9language modeling · 0.9hard view pretraining · 0.9eigenvalue analysis · 0.9contrastive learning · 0.9approximate normalization · 0.9adversarial view selection · 0.9weight sharing · 0.8surrogate prediction · 0.8reinforcement learning · 0.8multi-objective optimization · 0.8l2-norm constraint · 0.8augmented lagrangian method · 0.8
YearPublicationVenuePosition
2025 Beyond Random Augmentations: Pretraining with Hard Views
abstract
Self-Supervised Learning (SSL) methods typically rely on random image augmentations, or views, to make models invariant to different transformations. We hypothesize that the efficacy of pretraining pipelines based on conventional random view sampling can be enhanced by explicitly selecting views that benefit the learning progress. A simple yet effective approach is to select hard views that yield a higher loss. In this paper, we propose Hard View Pretraining (HVP), a learning-free strategy that extends random view generation by exposing models to more challenging samples during SSL pretraining. HVP encompasses the following iterative steps: 1) randomly sample multiple views and forward each view through the pretrained model, 2) create pairs of two views and compute their loss, 3) adversarially select the pair yielding the highest loss according to the current model state, and 4) perform a backward pass with the selected pair. In contrast to existing hard view literature, we are the first to demonstrate hard view pretraining's effectiveness at scale, particularly training on the full ImageNet-1k dataset, and evaluating across multiple SSL methods, Convolutional Networks, and Vision Transformers. As a result, HVP sets a new state-of-the-art on DINO ViT-B/16, reaching 78.8% linear evaluation accuracy (a 0.6% improvement) and consistent gains of 1% for both 100 and 300 epoch pretraining, with similar improvements across transfer tasks in DINO, SimSiam, iBOT, and SimCLR.
Fabio Ferreira, Ivo Rapant, Jörg K. H. Franke, Frank Hutter
ICLR3
2025 Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
abstract
Linear Recurrent Neural Networks (LRNNs) such as Mamba, RWKV, GLA, mLSTM, and DeltaNet have emerged as efficient alternatives to Transformers for long sequences. However, both Transformers and LRNNs struggle to perform state-tracking, which may impair performance in tasks such as code evaluation. In one forward pass, current architectures are unable to solve even parity, the simplest state-tracking task, which non-linear RNNs can handle effectively. Recently, Sarrof et al. (2024) demonstrated that the failure of LRNNs like Mamba to solve parity stems from restricting the value range of their diagonal state-transition matrices to $[0, 1]$ and that incorporating negative values can resolve this issue. We extend this result to non-diagonal LRNNs such as DeltaNet. We prove that finite precision LRNNs with state-transition matrices having only positive eigenvalues cannot solve parity, while non-triangular matrices are needed to count modulo $3$. Notably, we also prove that LRNNs can learn any regular language when their state-transition matrices are products of identity minus vector outer product matrices, each with eigenvalues in the range $[-1, 1]$. Our experiments confirm that extending the eigenvalue range of Mamba and DeltaNet to include negative values not only enables them to solve parity but consistently improves their performance on state-tracking tasks. We also show that state-tracking enabled LRNNs can be pretrained stably and efficiently at scale (1.3B parameters), achieving competitive performance on language modeling and showing promise on code and math tasks.
Riccardo Grazzi, Julien Siems, Arber Zela, Jörg K. H. Franke, Frank Hutter, Massimiliano Pontil
ICLR4
2025 KinPFN: Bayesian Approximation of RNA Folding Kinetics using Prior-Data Fitted Networks
abstract
RNA is a dynamic biomolecule crucial for cellular regulation, with its function largely determined by its folding into complex structures, while misfolding can lead to multifaceted biological sequelae. During the folding process, RNA traverses through a series of intermediate structural states, with each transition occurring at variable rates that collectively influence the time required to reach the functional form. Understanding these folding kinetics is vital for predicting RNA behavior and optimizing applications in synthetic biology and drug discovery. While in silico kinetic RNA folding simulators are often computationally intensive and time-consuming, accurate approximations of the folding times can already be very informative to assess the efficiency of the folding process. In this work, we present KinPFN, a novel approach that leverages prior-data fitted networks to directly model the posterior predictive distribution of RNA folding times. By training on synthetic data representing arbitrary prior folding times, KinPFN efficiently approximates the cumulative distribution function of RNA folding times in a single forward pass, given only a few initial folding time examples. Our method offers a modular extension to existing RNA kinetics algorithms, promising significant computational speed-ups orders of magnitude faster, while achieving comparable results. We showcase the effectiveness of KinPFN through extensive evaluations and real-world case studies, demonstrating its potential for RNA folding kinetics analysis, its practical relevance, and generalization to other biological data.
Dominik Scheuer, Frederic Runge, Jörg K. H. Franke, Michael T. Wolfinger, Christoph Flamm, Frank Hutter
ICLR3
2025 Learning in Compact Spaces with Approximately Normalized Transformer
abstract
The successful training of deep neural networks requires addressing challenges such as overfitting, numerical instabilities leading to divergence, and increasing variance in the residual stream. A common solution is to apply regularization and normalization techniques that usually require tuning additional hyperparameters. An alternative is to force all parameters and representations to lie on a hypersphere. This removes the need for regularization and increases convergence speed, but comes with additional costs. In this work, we propose a more holistic, approximate normalization via simple scalar multiplications motivated by the tight concentration of the norms of high-dimensional random vectors. Additionally, instead of applying strict normalization for the parameters, we constrain their norms. These modifications remove the need for weight decay and learning rate warm-up as well, but do not increase the total number of normalization layers. Our experiments with transformer architectures show up to 40% faster convergence compared to GPT models with QK normalization, with only 3% additional runtime cost. When deriving scaling laws, we found that our method enables training with larger batch sizes while preserving the favorable scaling characteristics of classic GPT architectures.
Jörg K. H. Franke, Urs Spiegelhalter, Marianna Nezhurina, Jenia Jitsev, Frank Hutter, Michael Hefenbrock
NeurIPS1
2025 Practitioner Motives to Use Different Hyperparameter Optimization Methods
abstract
Programmatic hyperparameter optimization (HPO) methods, such as Bayesian optimization and evolutionary algorithms, are known for their sample efficiency in identifying optimal configurations for machine learning (ML) models. However, practitioners often use less efficient methods, such as grid search, potentially resulting in under-optimized models. This discrepancy suggests that HPO method selection may be influenced by practitioner-specific motives, which remain insufficiently understood hindering user-centered advancement of HPO tools. To uncover these motives, we conducted 20 semi-structured interviews and an online survey with 49 ML practitioners. We revealed six primary goals (e.g., increasing ML model understanding) and 14 contextual factors (e.g., available computational resources) that influence practitioners’ choices of HPO methods. This study provides a conceptual foundation for understanding real-world HPO practices and informs the development of more user-centered and context-adaptive HPO tools in automated ML (AutoML).
Niclas Kannengießer, Niklas Hasebrook, Felix Morsbach, Marc-André Zöller, Jörg K. H. Franke, Marius Lindauer, Frank Hutter, Ali Sunyaev
ACM Trans. Comput. Hum. Interact.5
2024 Improving Deep Learning Optimization through Constrained Parameter Regularization
abstract
Regularization is a critical component in deep learning. The most commonly used approach, weight decay, applies a constant penalty coefficient uniformly across all parameters. This may be overly restrictive for some parameters, while insufficient for others. To address this, we present Constrained Parameter Regularization (CPR) as an alternative to traditional weight decay. Unlike the uniform application of a single penalty, CPR enforces an upper bound on a statistical measure, such as the L$_2$-norm, of individual parameter matrices. Consequently, learning becomes a constraint optimization problem, which we tackle using an adaptation of the augmented Lagrangian method. CPR introduces only a minor runtime overhead and only requires setting an upper bound. We propose simple yet efficient mechanisms for initializing this bound, making CPR rely on no hyperparameter or one, akin to weight decay. Our empirical studies on computer vision and language modeling tasks demonstrate CPR's effectiveness. The results show that CPR can outperform traditional weight decay and increase performance in pre-training and fine-tuning.
Jörg K. H. Franke, Michael Hefenbrock, Gregor Köhler, Frank Hutter
NeurIPS1
2024 HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
abstract
The increasing size of language models necessitates a thorough analysis across multiple dimensions to assess trade-offs among crucial hardware metrics such as latency, energy consumption, GPU memory usage, and performance. Identifying optimal model configurations under specific hardware constraints is becoming essential but remains challenging due to the computational load of exhaustive training and evaluation on multiple devices. To address this, we introduce HW-GPT-Bench, a hardware-aware benchmark that utilizes surrogate predictions to approximate various hardware metrics across 13 devices of architectures in the GPT-2 family, with architectures containing up to 1.55B parameters. Our surrogates, via calibrated predictions and reliable uncertainty estimates, faithfully model the heteroscedastic noise inherent in the energy and latency measurements. To estimate perplexity, we employ weight-sharing techniques from Neural Architecture Search (NAS), inheriting pretrained weights from the largest GPT-2 model. Finally, we demonstrate the utility of HW-GPT-Bench by simulating optimization trajectories of various multi-objective optimization algorithms in just a few seconds.
Rhea Sanjay Sukthanker, Arber Zela, Benedikt Staffler, Aaron Klein, Lennart Purucker, Jörg K. H. Franke, Frank Hutter
NeurIPS6
2024 RecycleNet: Latent Feature Recycling Leads to Iterative Decision Refinement
abstract
Despite the remarkable success of deep learning systems over the last decade, a key difference still remains between neural network and human decision-making: As humans, we can not only form a decision on the spot, but also ponder, revisiting an initial guess from different angles, distilling relevant information, arriving at a better decision. Here, we propose RecycleNet, a latent feature recycling method, instilling the pondering capability for neural networks to refine initial decisions over a number of recycling steps, where outputs are fed back into earlier network layers in an iterative fashion. This approach makes minimal assumptions about the neural network architecture and thus can be implemented in a wide variety of contexts. Using medical image segmentation as the evaluation environment, we show that latent feature recycling enables the network to iteratively refine initial predictions even beyond the iterations seen during training, converging towards an improved decision. We evaluate this across a variety of segmentation benchmarks and show consistent improvements even compared with top-performing segmentation methods. This allows trading increased computation time for improved performance, which can be beneficial, especially for safety-critical applications.
Gregor Köhler, Tassilo Wald, Constantin Ulrich, David Zimmerer, Paul F. Jaeger, Jörg K. H. Franke, Simon Kohl, Fabian Isensee, Klaus H. Maier-Hein
WACV6
2024 Partial RNA design
abstract
MOTIVATION: RNA design is a key technique to achieve new functionality in fields like synthetic biology or biotechnology. Computational tools could help to find such RNA sequences but they are often limited in their formulation of the search space. RESULTS: In this work, we propose partial RNA design, a novel RNA design paradigm that addresses the limitations of current RNA design formulations. Partial RNA design describes the problem of designing RNAs from arbitrary RNA sequences and structure motifs with multiple design goals. By separating the design space from the objectives, our formulation enables the design of RNAs with variable lengths and desired properties, while still allowing precise control over sequence and structure constraints at individual positions. Based on this formulation, we introduce a new algorithm, libLEARNA, capable of efficiently solving different constraint RNA design tasks. A comprehensive analysis of various problems, including a realistic riboswitch design task, reveals the outstanding performance of libLEARNA and its robustness. AVAILABILITY AND IMPLEMENTATION: libLEARNA is open-source and publicly available at: https://github.com/automl/learna_tools.
Frederic Runge, Jörg K. H. Franke, Daniel Fertmann, Rolf Backofen, Frank Hutter
Bioinform.2
2022 Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule Design
abstract
Our world is ambiguous and this is reflected in the data we use to train our algorithms. This is particularly true when we try to model natural processes where collected data is affected by noisy measurements and differences in measurement techniques. Sometimes, the process itself is ambiguous, such as in the case of RNA folding, where the same nucleotide sequence can fold into different structures. This suggests that a predictive model should have similar probabilistic characteristics to match the data it models. Therefore, we propose a hierarchical latent distribution to enhance one of the most successful deep learning models, the Transformer, to accommodate ambiguities and data distributions. We show the benefits of our approach (1) on a synthetic task that captures the ability to learn a hidden data distribution, (2) with state-of-the-art results in RNA folding that reveal advantages on highly ambiguous data, and (3) demonstrating its generative capabilities on property-based molecule design by implicitly learning the underlying distributions and outperforming existing work.
Jörg K. H. Franke, Frederic Runge, Frank Hutter
NeurIPS1
2021 Sample-Efficient Automated Deep Reinforcement Learning
Jörg K. H. Franke, Gregor Köhler, André Biedenkapp, Frank Hutter
ICLR1