VLDB 2026 Research / reviewers in the wild / expert
Jannes Elstner
dblp:312/6696
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Trustworthy machine learning · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability
representation engineering |
0.9 | 1 | 2025 | The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence · ICML 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
gradient-based representation engineering · 0.9activation intervention · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Geometry of Refusal in Large Language Models: Concept Cones and Representational IndependenceabstractThe safety alignment of large language models (LLMs) can be circumvented through adversarially crafted inputs, yet the mechanisms by which these attacks bypass safety barriers remain poorly understood. Prior work suggests that a *single* refusal direction in the model's activation space determines whether an LLM refuses a request. In this study, we propose a novel gradient-based approach to representation engineering and use it to identify refusal directions. Contrary to prior work, we uncover multiple independent directions and even multi-dimensional *concept cones* that mediate refusal. Moreover, we show that orthogonality alone does not imply independence under intervention, motivating the notion of *representational independence* that accounts for both linear and non-linear effects. Using this framework, we identify mechanistically independent refusal directions. We show that refusal mechanisms in LLMs are governed by complex spatial structures and identify functionally independent directions, confirming that multiple distinct mechanisms drive refusal behavior. Our gradient-based approach uncovers these mechanisms and can further serve as a foundation for future work on understanding LLMs. Tom Wollschläger, Jannes Elstner, Simon Geisler, Vincent Cohen-Addad, Stephan Günnemann, Johannes Gasteiger |
ICML | 2 |