Jonathan Kahana

dblp:317/0994 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-0775-197XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Representation and self-supervised learning · 23% Time series and sequential data · 19% Efficient and distributed learning · 12%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%
Software engineering, system software, and programming languages
1 paper
Software maintenance and evolution · 100%

Topics — the 11 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning › automated reasoning
model search
0.912025
Learning on Model Weights using Tree Experts · CVPR 2025
Machine learning › Representation and self-supervised learning
probing
0.912025
Deep Linear Probe Generators for Weight Space Learning · ICLR 2025
Machine learning › Transfer learning and domain adaptation › meta-learning
weight space learning
0.912025
Deep Linear Probe Generators for Weight Space Learning · ICLR 2025
Security and privacy of machine learning
model stealing
0.812024
Recovering the Pre-Fine-Tuning Weights of Generative Models · ICML 2024
Machine learning › Time series and sequential data
anomaly detection
0.712023
Red PANDA: Disambiguating Image Anomaly Detection by Removing Nuisance Factors · ICLR 2023
Machine learning › Trustworthy machine learning
interpretability
0.712023
Red PANDA: Disambiguating Image Anomaly Detection by Removing Nuisance Factors · ICLR 2023
Machine learning › Time series and sequential data › anomaly detection
visual anomaly detection
0.712023
Red PANDA: Disambiguating Image Anomaly Detection by Removing Nuisance Factors · ICLR 2023
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.612022
A Contrastive Objective for Learning Disentangled Representations · ECCV (26) 2022
Software maintenance and evolution › software ecosystems
model reuse
0.312025
We Should Chart an Atlas of All the World's Models · NeurIPS 2025
Software maintenance and evolution
software ecosystems
0.312025
We Should Chart an Atlas of All the World's Models · NeurIPS 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.212022
A Contrastive Objective for Learning Disentangled Representations · ECCV (26) 2022

Methods — techniques the papers use, named apart from their topics

spectral decomposition · 1.5LoRA fine-tuning · 1.5weight-language embedding · 0.9tree expert · 0.9linear probing · 0.9deep linear probe generators · 0.9disentanglement · 0.7anomaly detection · 0.7contrastive objective · 0.6
YearPublicationVenuePosition
2025 Learning on Model Weights using Tree Experts
abstract
The number of publicly available models is rapidly increasing, yet most remain undocumented. Users looking for suitable models for their tasks must first determine what each model does. Training machine learning models to infer missing documentation directly from model weights is challenging, as these weights often contain significant variation unrelated to model functionality (denoted nuisance). Here, we identify a key property of real-world models: most public models belong to a small set of Model Trees, where all models within a tree are fine-tuned from a common ancestor (e.g., a foundation model). Importantly, we find that within each tree there is less nuisance variation between models. Concretely, while learning across Model Trees requires complex architectures, even a linear classifier trained on a single model layer often works within trees. While effective, these linear classifiers are computationally expensive, especially when dealing with larger models that have many parameters. To address this, we introduce Probing Experts (ProbeX), a theoretically motivated and lightweight method. Notably, ProbeX is the first probing method specifically designed to learn from the weights of a single hidden model layer. We demonstrate the effectiveness of ProbeX by predicting the categories in a model’s training dataset based only on its weights. Excitingly, ProbeX can map the weights of Stable Diffusion into a weight-language embedding space, enabling model search via text, i.e., zero-shot model classification.
Eliahu Horwitz, Bar Cavia, Jonathan Kahana, Yedid Hoshen
CVPR3
2025 Deep Linear Probe Generators for Weight Space Learning
abstract
Weight space learning aims to extract information about a neural network, such as its training dataset or generalization error. Recent approaches learn directly from model weights, but this presents many challenges as weights are high-dimensional and include permutation symmetries between neurons. An alternative approach, Probing, represents a model by passing a set of learned inputs (probes) through the model, and training a predictor on top of the corresponding outputs. Although probing is typically not used as a stand alone approach, our preliminary experiment found that a vanilla probing baseline worked surprisingly well. However, we discover that current probe learning strategies are ineffective. We therefore propose Deep Linear Probe Generators (ProbeGen), a simple and effective modification to probing approaches. ProbeGen adds a shared generator module with a deep linear architecture, providing an inductive bias towards structured probes thus reducing overfitting. While simple, ProbeGen performs significantly better than the state-of-the-art and is very efficient, requiring between 30 to 1000 times fewer FLOPs than other top approaches.
Jonathan Kahana, Eliahu Horwitz, Imri Shuval, Yedid Hoshen
ICLR1
2025 We Should Chart an Atlas of All the World's Models
abstract
Public model repositories now contain millions of models, yet most remain undocumented and effectively lost: their capabilities, provenance, and constraints cannot be reliably determined. As a result, the field wastes training time and compute, propagates hidden biases, faces intellectual-property risks, and misses opportunities for model reuse and transfer. In this position paper, we advocate charting the world's model population in a unified structure we call the Model Atlas: a graph that captures models, their attributes, and the weight transformations connecting them. The Model Atlas enables applications in model forensics, meta-ML research, and model discovery, challenging tasks given today's unstructured model repositories. However, because most models lack documentation, large atlas regions remain uncharted. Addressing this gap motivates new machine learning methods that treat models themselves as data and infer properties such as functionality, performance, and lineage directly from their weights. We argue that a scalable path forward is to bypass the unique parameter symmetries that plague model weights. Charting all the world's models will require a community effort, and we hope its broad utility will rally researchers toward this goal.
Eliahu Horwitz, Nitzan Kurer, Jonathan Kahana, Liel Amar, Yedid Hoshen
NeurIPS3
2024 Recovering the Pre-Fine-Tuning Weights of Generative Models
abstract
The dominant paradigm in generative modeling consists of two steps: i) pre-training on a large-scale but unsafe dataset, ii) aligning the pre-trained model with human values via fine-tuning. This practice is considered safe, as no current method can recover the unsafe, pre-fine-tuning model weights. In this paper, we demonstrate that this assumption is often false. Concretely, we present Spectral DeTuning, a method that can recover the weights of the pre-fine-tuning model using a few low-rank (LoRA) fine-tuned models. In contrast to previous attacks that attempt to recover pre-fine-tuning capabilities, our method aims to recover the exact pre-fine-tuning weights. Our approach exploits this new vulnerability against large-scale models such as a personalized Stable Diffusion and an aligned Mistral. The code is available at https://vision.huji.ac.il/spectral_detuning/.
Eliahu Horwitz, Jonathan Kahana, Yedid Hoshen
ICML2
2023 Red PANDA: Disambiguating Image Anomaly Detection by Removing Nuisance Factors
Niv Cohen, Jonathan Kahana, Yedid Hoshen
ICLR2
2022 A Contrastive Objective for Learning Disentangled Representations
Jonathan Kahana, Yedid Hoshen
ECCV (26)1