Itay Yona

dblp:368/3292 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0008-4264-9330ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 60% Language models and text generation · 30% Generative modeling · 10%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning
adversarial attack
1.012026
In-Context Representation Hijacking · ACL (1) 2026
Security and privacy of machine learning › adversarial attack
jailbreak attack
1.012026
In-Context Representation Hijacking · ACL (1) 2026
Machine learning › Trustworthy machine learning
interpretability
0.912025
Interpreting the Repeated Token Phenomenon in Large Language Models · ICML 2025
Natural language and speech › Language models and text generation
large language model
0.912025
Interpreting the Repeated Token Phenomenon in Large Language Models · ICML 2025
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.912025
Interpreting the Repeated Token Phenomenon in Large Language Models · ICML 2025
Machine learning › Generative modeling
latent space interpretation
0.312026
In-Context Representation Hijacking · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

interpretability tools · 2.0targeted patching · 0.9neural circuit analysis · 0.9
YearPublicationVenuePosition
2026 In-Context Representation Hijacking
abstract
We introduce Doublespeak, a simple incontext representation hijacking attack against language models.The attack works by systematically replacing a harmful keyword (e.g., bomb) with a benign token (e.g., carrot) across multiple in-context examples, provided as a prefix to a harmful request.We demonstrate that this substitution leads to the internal representation of the benign token converging toward that of the harmful one, effectively embedding the harmful semantics under a euphemism.As a result, superficially innocuous prompts (e.g., "How to build a carrot?") are internally interpreted as disallowed instructions (e.g., "How to build a bomb?"), while bypassing the model's safety alignment.We use interpretability tools to show this semantic shift occurs progressively across layers.Doublespeak is optimizationfree, broadly transferable across model families, and achieves strong success rates on closed-source and open-source systems, reaching 74% ASR on Llama-3.3-70B-Instruct with a single-sentence context override.Our findings highlight a new attack surface in LM latent space, indicating that current alignment strategies are insufficient and should instead operate at the representation level. 1
Itay Yona, Amir Sarid, Michael Karasik, Yossi Gandelsman
ACL (1)1
2025 Interpreting the Repeated Token Phenomenon in Large Language Models
abstract
Large Language Models (LLMs), despite their impressive capabilities, often fail to accurately repeat a single word when prompted to, and instead output unrelated text. This unexplained failure mode represents a *vulnerability*, allowing even end users to diverge models away from their intended behavior. We aim to explain the causes for this phenomenon and link it to the concept of "attention sinks", an emergent LLM behavior crucial for fluency, in which the initial token receives disproportionately high attention scores. Our investigation identifies the neural circuit responsible for attention sinks and shows how long repetitions disrupt this circuit. We extend this finding to other nonrepeating sequences that exhibit similar circuit disruptions. To address this, we propose a targeted patch that effectively resolves the issue without negatively impacting the overall performance of the model. This study provides a mechanistic explanation for an LLM vulnerability, demonstrating how interpretability can diagnose and address issues, and offering insights that pave the way for more secure and reliable models.
Itay Yona, Ilia Shumailov, Jamie Hayes, Yossi Gandelsman
ICML1
2025 Measuring memorization in language models via probabilistic extraction
abstract
Jamie Hayes, Marika Swanberg, Harsh Chaudhari, Itay Yona, Ilia Shumailov, Milad Nasr, Christopher A. Choquette-Choo, Katherine Lee, A. Feder Cooper. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jamie Hayes, Marika Swanberg, Harsh Chaudhari, Itay Yona, Ilia Shumailov, Milad Nasr, Christopher A. Choquette-Choo, Katherine Lee, A. Feder Cooper
NAACL (Long Papers)4