Jan Disselhoff

dblp:334/7753 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0003-0379-5673ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 31% Trustworthy machine learning · 26% Knowledge representation and reasoning · 11%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning
abstract reasoning
0.912025
Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective · ICML 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › tree search
depth-first search
0.912025
Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective · ICML 2025
Natural language and speech › Language models and text generation
large language model
0.912025
Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective · ICML 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective · ICML 2025
Machine learning › Trustworthy machine learning › interpretability › attribution methods
feature attribution
0.712023
Initialization Noise in Image Gradients and Saliency Maps · CVPR 2023
Machine learning › Trustworthy machine learning
interpretability
0.712023
Initialization Noise in Image Gradients and Saliency Maps · CVPR 2023
Natural language and speech › Language models and text generation › text generation › surface realization
linearization
0.712023
Nonlinear Advantage: Trained Networks Might Not Be As Complex as You Think · ICML 2023
Machine learning › Efficient and distributed learning › model compression › sparsity
network sparsification
0.712023
Nonlinear Advantage: Trained Networks Might Not Be As Complex as You Think · ICML 2023
Machine learning › Deep learning architectures and training
neural network expressivity
0.712023
Nonlinear Advantage: Trained Networks Might Not Be As Complex as You Think · ICML 2023
Machine learning › Trustworthy machine learning › interpretability › visual explanation
saliency map
0.712023
Initialization Noise in Image Gradients and Saliency Maps · CVPR 2023
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.212023
Nonlinear Advantage: Trained Networks Might Not Be As Complex as You Think · ICML 2023

Methods — techniques the papers use, named apart from their topics

product of experts · 0.9depth-first search · 0.9data augmentation · 0.9smoothgrad · 0.7integrated gradients · 0.7average path length measure · 0.7SHAP · 0.7LIME · 0.7Grad-CAM · 0.7DeepLIFT · 0.7
YearPublicationVenuePosition
2025 Product of Experts with LLMs: Boosting Performance on ARC Is a Matter of Perspective
abstract
The Abstraction and Reasoning Corpus (ARC-AGI) poses a significant challenge for large language models (LLMs), exposing limitations in their abstract reasoning abilities. In this work, we leverage task-specific data augmentations throughout the training, generation, and scoring phases, and employ a depth-first search algorithm to generate diverse, high-probability candidate solutions. Furthermore, we utilize the LLM not only as a generator but also as a scorer, using its output probabilities to select the most promising solutions. Our method achieves a score of 71.6% (286.5/400 solved tasks) on the public ARC-AGI evaluation set, demonstrating state-of-the-art performance among publicly available approaches. While concurrent closed-source work has reported higher scores, our method distinguishes itself through its transparency, reproducibility, and remarkably low inference cost, averaging only around 2ct per task on readily available hardware.
Daniel Franzen, Jan Disselhoff, David Hartmann
ICML2
2023 Initialization Noise in Image Gradients and Saliency Maps
abstract
In this paper, we examine gradients of logits of image classification CNNs by input pixel values. We observe that these fluctuate considerably with training randomness, such as the random initialization of the networks. We extend our study to gradients of intermediate layers, obtained via GradCAM, as well as popular network saliency estimators such as DeepLIFT, SHAP, LIME, Integrated Gradients, and SmoothGrad. While empirical noise levels vary, qualitatively different attributions to image features are still possible with all of these, which comes with implications for interpreting such attributions, in particular when seeking data-driven explanations of the phenomenon generating the data. Finally, we demonstrate that the observed artefacts can be removed by marginalization over the initialization distribution by simple stochastic integration.
Ann-Christin Woerl, Jan Disselhoff, Michael Wand 0001
CVPR2
2023 Nonlinear Advantage: Trained Networks Might Not Be As Complex as You Think
abstract
We perform an empirical study of the behaviour of deep networks when fully linearizing some of its feature channels through a sparsity prior on the overall number of nonlinear units in the network. In experiments on image classification and machine translation tasks, we investigate how much we can simplify the network function towards linearity before performance collapses. First, we observe a significant performance gap when reducing nonlinearity in the network function early on as opposed to late in training, in-line with recent observations on the time-evolution of the data-dependent NTK. Second, we find that after training, we are able to linearize a significant number of nonlinear units while maintaining a high performance, indicating that much of a network's expressivity remains unused but helps gradient descent in early stages of training. To characterize the depth of the resulting partially linearized network, we introduce a measure called average path length, representing the average number of active nonlinearities encountered along a path in the network graph. Under sparsity pressure, we find that the remaining nonlinear units organize into distinct structures, forming core-networks of near constant effective depth and width, which in turn depend on task difficulty.
Christian H. X. Ali Mehmeti-Göpel, Jan Disselhoff
ICML2