EDBT 2026 Demo / reviewers in the wild / expert
Oam Patel
dblp:348/9558
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 70% Trustworthy machine learning · 30% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering |
0.7 | 1 | 2023 | Inference-Time Intervention: Eliciting Truthful Answers from a Language Model · NeurIPS 2023 |
Natural language and speech › Language models and text generation
inference-time intervention |
0.7 | 1 | 2023 | Inference-Time Intervention: Eliciting Truthful Answers from a Language Model · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | Inference-Time Intervention: Eliciting Truthful Answers from a Language Model · NeurIPS 2023 |
Natural language and speech › Language models and text generation › instruction following
instruction-following language models |
0.2 | 1 | 2023 | Inference-Time Intervention: Eliciting Truthful Answers from a Language Model · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
attention head analysis · 0.7activation intervention · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelabstractWe introduce Inference-Time Intervention (ITI), a technique designed to enhance the "truthfulness" of large language models (LLMs). ITI operates by shifting model activations during inference, following a learned set of directions across a limited number of attention heads. This intervention significantly improves the performance of LLaMA models on the TruthfulQA benchmark. On an instruction-finetuned LLaMA called Alpaca, ITI improves its truthfulness from $32.5\%$ to $65.1\%$. We identify a tradeoff between truthfulness and helpfulness and demonstrate how to balance it by tuning the intervention strength. ITI is minimally invasive and computationally inexpensive. Moreover, the technique is data efficient: while approaches like RLHF require extensive annotations, ITI locates truthful directions using only few hundred examples. Our findings suggest that LLMs may have an internal representation of the likelihood of something being true, even as they produce falsehoods on the surface. Kenneth Li 0002, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister, Martin Wattenberg |
NeurIPS | 2 |