Narutatsu Ri

dblp:347/9877 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 76% Language models and text generation · 24%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
1.022025
Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations · ICML 2024
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions · ICML 2025
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions · ICML 2025
Security and privacy of machine learning › adversarial attack
jailbreak attack
0.912025
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions · ICML 2025
Machine learning › Trustworthy machine learning › interpretability
explanation evaluation
0.812024
Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations · ICML 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations · ICML 2024
Machine learning › Trustworthy machine learning › interpretability
natural language explanation
0.812024
Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations · ICML 2024

Methods — techniques the papers use, named apart from their topics

multilingual prompting · 1.7multi-step interaction · 1.7counterfactual generation with LLMs · 0.8RLHF · 0.8
YearPublicationVenuePosition
2025 Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution
abstract
Recent state-of-the-art authorship attribution methods learn authorship representations of text in a latent, uninterpretable space, which hinders their usability in real-world applications. We propose a novel approach for interpreting learned embeddings by identifying representative points in the latent space and leveraging large language models to generate informative natural language descriptions of the writing style associated with each point. We evaluate the alignment between our interpretable and latent spaces and demonstrate superior prediction agreement over baseline methods. Additionally, we conduct a human evaluation to assess the quality of these style descriptions and validate their utility in explaining the latent space. Finally, we show that human performance on the challenging authorship attribution task improves by +20% on average when aided with explanations from our method.
Milad Alshomary, Narutatsu Ri, Marianna Apidianaki, Ajay Patel, Smaranda Muresan, Kathy McKeown
COLING2
2025 Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
abstract
Despite extensive safety alignment efforts, large language models (LLMs) remain vulnerable to jailbreak attacks that elicit harmful behavior. While existing studies predominantly focus on attack methods that require technical expertise, two critical questions remain underexplored: (1) Are jailbroken responses truly useful in enabling average users to carry out harmful actions? (2) Do safety vulnerabilities exist in more common, simple human-LLM interactions? In this paper, we demonstrate that LLM responses most effectively facilitate harmful actions when they are both *actionable* and *informative*---two attributes easily elicited in multi-step, multilingual interactions. Using this insight, we propose HarmScore, a jailbreak metric that measures how effectively an LLM response enables harmful actions, and Speak Easy, a simple multi-step, multilingual attack framework. Notably, by incorporating Speak Easy into direct request and jailbreak baselines, we see an average absolute increase of $0.319$ in Attack Success Rate and $0.426$ in HarmScore in both open-source and proprietary LLMs across four safety benchmarks. Our work reveals a critical yet often overlooked vulnerability: Malicious users can easily exploit common interaction patterns for harmful intentions.
Yik Siu Chan, Narutatsu Ri, Yuxin Xiao, Marzyeh Ghassemi
ICML2
2024 Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations
abstract
Large language models (LLMs) are trained to imitate humans to explain human decisions. However, do LLMs explain themselves? Can they help humans build mental models of how LLMs process different inputs? To answer these questions, we propose to evaluate $\textbf{counterfactual simulatability}$ of natural language explanations: whether an explanation can enable humans to precisely infer the model’s outputs on diverse counterfactuals of the explained input. For example, if a model answers ”$\textit{yes}$” to the input question ”$\textit{Can eagles fly?}$” with the explanation ”$\textit{all birds can fly}$”, then humans would infer from the explanation that it would also answer ”$\textit{yes}$” to the counterfactual input ”$\textit{Can penguins fly?}$”. If the explanation is precise, then the model’s answer should match humans’ expectations. We implemented two metrics based on counterfactual simulatability: precision and generality. We generated diverse counterfactuals automatically using LLMs. We then used these metrics to evaluate state-of-the-art LLMs (e.g., GPT-4) on two tasks: multi-hop factual reasoning and reward modeling. We found that LLM’s explanations have low precision and that precision does not correlate with plausibility. Therefore, naively optimizing human approvals (e.g., RLHF) may be insufficient.
Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 0013, He He 0001, Jacob Steinhardt, Kathy McKeown
ICML3