VLDB 2026 Research / reviewers in the wild / expert
Narutatsu Ri
dblp:347/9877
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 76% Language models and text generation · 24% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model |
1.0 | 2 | 2025 | Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations · ICML 2024 Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions · ICML 2025 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.9 | 1 | 2025 | Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions · ICML 2025 |
Security and privacy of machine learning › adversarial attack
jailbreak attack |
0.9 | 1 | 2025 | Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability
explanation evaluation |
0.8 | 1 | 2024 | Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations · ICML 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations · ICML 2024 |
Machine learning › Trustworthy machine learning › interpretability
natural language explanation |
0.8 | 1 | 2024 | Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
multilingual prompting · 1.7multi-step interaction · 1.7counterfactual generation with LLMs · 0.8RLHF · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Latent Space Interpretation for Stylistic Analysis and Explainable Authorship AttributionabstractRecent state-of-the-art authorship attribution methods learn authorship representations of text in a latent, uninterpretable space, which hinders their usability in real-world applications. We propose a novel approach for interpreting learned embeddings by identifying representative points in the latent space and leveraging large language models to generate informative natural language descriptions of the writing style associated with each point. We evaluate the alignment between our interpretable and latent spaces and demonstrate superior prediction agreement over baseline methods. Additionally, we conduct a human evaluation to assess the quality of these style descriptions and validate their utility in explaining the latent space. Finally, we show that human performance on the challenging authorship attribution task improves by +20% on average when aided with explanations from our method. Milad Alshomary, Narutatsu Ri, Marianna Apidianaki, Ajay Patel, Smaranda Muresan, Kathy McKeown |
COLING | 2 |
| 2025 | Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple InteractionsabstractDespite extensive safety alignment efforts, large language models (LLMs) remain vulnerable to jailbreak attacks that elicit harmful behavior. While existing studies predominantly focus on attack methods that require technical expertise, two critical questions remain underexplored: (1) Are jailbroken responses truly useful in enabling average users to carry out harmful actions? (2) Do safety vulnerabilities exist in more common, simple human-LLM interactions? In this paper, we demonstrate that LLM responses most effectively facilitate harmful actions when they are both *actionable* and *informative*---two attributes easily elicited in multi-step, multilingual interactions. Using this insight, we propose HarmScore, a jailbreak metric that measures how effectively an LLM response enables harmful actions, and Speak Easy, a simple multi-step, multilingual attack framework. Notably, by incorporating Speak Easy into direct request and jailbreak baselines, we see an average absolute increase of $0.319$ in Attack Success Rate and $0.426$ in HarmScore in both open-source and proprietary LLMs across four safety benchmarks. Our work reveals a critical yet often overlooked vulnerability: Malicious users can easily exploit common interaction patterns for harmful intentions. Yik Siu Chan, Narutatsu Ri, Yuxin Xiao, Marzyeh Ghassemi |
ICML | 2 |
| 2024 | Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsabstractLarge language models (LLMs) are trained to imitate humans to explain human decisions. However, do LLMs explain themselves? Can they help humans build mental models of how LLMs process different inputs? To answer these questions, we propose to evaluate $\textbf{counterfactual simulatability}$ of natural language explanations: whether an explanation can enable humans to precisely infer the model’s outputs on diverse counterfactuals of the explained input. For example, if a model answers ”$\textit{yes}$” to the input question ”$\textit{Can eagles fly?}$” with the explanation ”$\textit{all birds can fly}$”, then humans would infer from the explanation that it would also answer ”$\textit{yes}$” to the counterfactual input ”$\textit{Can penguins fly?}$”. If the explanation is precise, then the model’s answer should match humans’ expectations. We implemented two metrics based on counterfactual simulatability: precision and generality. We generated diverse counterfactuals automatically using LLMs. We then used these metrics to evaluate state-of-the-art LLMs (e.g., GPT-4) on two tasks: multi-hop factual reasoning and reward modeling. We found that LLM’s explanations have low precision and that precision does not correlate with plausibility. Therefore, naively optimizing human approvals (e.g., RLHF) may be insufficient. Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 0013, He He 0001, Jacob Steinhardt, Kathy McKeown |
ICML | 3 |