Oam Patel

dblp:348/9558 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 70% Trustworthy machine learning · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering
0.712023
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model · NeurIPS 2023
Natural language and speech › Language models and text generation
inference-time intervention
0.712023
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model · NeurIPS 2023
Machine learning › Trustworthy machine learning
interpretability
0.712023
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model · NeurIPS 2023
Natural language and speech › Language models and text generation › instruction following
instruction-following language models
0.212023
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

attention head analysis · 0.7activation intervention · 0.7
YearPublicationVenuePosition
2023 Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
abstract
We introduce Inference-Time Intervention (ITI), a technique designed to enhance the "truthfulness" of large language models (LLMs). ITI operates by shifting model activations during inference, following a learned set of directions across a limited number of attention heads. This intervention significantly improves the performance of LLaMA models on the TruthfulQA benchmark. On an instruction-finetuned LLaMA called Alpaca, ITI improves its truthfulness from $32.5\%$ to $65.1\%$. We identify a tradeoff between truthfulness and helpfulness and demonstrate how to balance it by tuning the intervention strength. ITI is minimally invasive and computationally inexpensive. Moreover, the technique is data efficient: while approaches like RLHF require extensive annotations, ITI locates truthful directions using only few hundred examples. Our findings suggest that LLMs may have an internal representation of the likelihood of something being true, even as they produce falsehoods on the surface.
Kenneth Li 0002, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister, Martin Wattenberg
NeurIPS2