Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sheridan Feucht

dblp:306/1619 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 92% Representation and self-supervised learning · 8%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › machine unlearning
concept unlearning
0.912025
Erasing Conceptual Knowledge from Language Models · NeurIPS 2025
Machine learning › Trustworthy machine learning
machine unlearning
0.912025
Erasing Conceptual Knowledge from Language Models · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.812024
Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs · EMNLP 2024
Machine learning › Representation and self-supervised learning › word representation
token embedding
0.212024
Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

low-rank updates · 1.7introspective classification · 1.7distribution matching · 1.7layer-wise representation analysis · 0.8
YearPublicationVenuePosition
2025 Erasing Conceptual Knowledge from Language Models
abstract
In this work, we introduce Erasure of Language Memory (ELM), a principled approach to concept-level unlearning that operates by matching distributions defined by the model's own introspective classification capabilities. Our key insight is that effective unlearning should leverage the model's ability to evaluate its own knowledge, using the language model itself as a classifier to identify and reduce the likelihood of generating content related to undesired concepts. ELM applies this framework to create targeted low-rank updates that reduce generation probabilities for concept-specific content while preserving the model's broader capabilities. We demonstrate ELM's efficacy on biosecurity, cybersecurity, and literary domain erasure tasks. Comparative evaluation reveals that ELM-modified models achieve near-random performance on assessments targeting erased concepts, while simultaneously preserving generation coherence, maintaining benchmark performance on unrelated tasks, and exhibiting strong robustness to adversarial attacks.
Rohit Gandikota, Sheridan Feucht, Samuel Marks, David Bau
NeurIPS2
2024 Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs
abstract
LLMs process text as sequences of tokens that roughly correspond to words, where less common words are represented by multiple tokens.However, individual tokens are often semantically unrelated to the meanings of the words/concepts they comprise.For example, Llama-2-7b's tokenizer splits the word "northeastern" into the tokens [_n, ort, he, astern], none of which correspond to semantically meaningful units like "north" or "east."Similarly, the overall meanings of named entities like "Neil Young" and multi-word expressions like "break a leg" cannot be directly inferred from their constituent tokens.Mechanistically, how do LLMs convert such arbitrary groups of tokens into useful higher-level representations?In this work, we find that last token representations of named entities and multi-token words exhibit a pronounced "erasure" effect, where information about previous and current tokens is rapidly forgotten in early layers.Using this observation, we propose a method to "read out" the implicit vocabulary of an autoregressive LLM by examining differences in token representations across layers, and present results of this method for Llama-2-7b and Llama-3-8b.To our knowledge, this is the first attempt to probe the implicit vocabulary of an LLM. 1
Sheridan Feucht, David Atkinson, Byron C. Wallace, David Bau
EMNLP1
2021 The Anatomy of Discourse: Linguistic Predictors of Narrative and Argument Quality
Sheridan Feucht, Babak Hemmatian, Rachel Avram, Alexander Wey, Kate Spitalnic, Muskaan Garg, Carsten Eickhoff, Ellie Pavlick, Björn Sandstede, Steven A. Sloman
CogSci1
2021 Can computers tell a story? Discourse Structure in Computer-generated Text and Humans
Alexander Wey, Babak Hemmatian, Rachel Avram, Sheridan Feucht, Kate Spitalnic, Muskaan Garg, Carsten Eickhoff, Ellie Pavlick, Björn Sandstede, Steven A. Sloman
CogSci4