VLDB 2026 Research / reviewers in the wild / expert
Meg Tong
dblp:356/2704
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 68% Trustworthy machine learning · 22% Knowledge representation and reasoning · 11% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › interpretability › representation engineering
activation editing |
0.8 | 1 | 2024 | Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024 |
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering |
0.8 | 1 | 2024 | Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Towards Understanding Sycophancy in Language Models · ICLR 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024 |
Natural language and speech › Language models and text generation › model steering
language model steering |
0.8 | 1 | 2024 | Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024 |
Natural language and speech › Language models and text generation
large language model safety |
0.8 | 1 | 2024 | Many-shot Jailbreaking · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model reasoning › reasoning robustness
reversal curse |
0.8 | 1 | 2024 | The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" · ICLR 2024 |
Natural language and speech › Language models and text generation › alignment
sycophancy |
0.8 | 1 | 2024 | Towards Understanding Sycophancy in Language Models · ICLR 2024 |
Security and privacy of machine learning
adversarial attack |
0.8 | 1 | 2024 | Many-shot Jailbreaking · NeurIPS 2024 |
Security and privacy of machine learning › adversarial attack
jailbreak attack |
0.8 | 1 | 2024 | Many-shot Jailbreaking · NeurIPS 2024 |
Natural language and speech › Language models and text generation › alignment
LLM behavior control |
0.2 | 1 | 2024 | Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024 |
Security and privacy of machine learning
prompt injection |
0.2 | 1 | 2024 | Many-shot Jailbreaking · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
in-context learning · 2.3long-context prompting · 1.5reinforcement learning from human feedback · 0.8preference modeling · 0.8fine-tuning · 0.8contrastive activation addition · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Steering Llama 2 via Contrastive Activation AdditionabstractNina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, Alexander Turner. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, Alexander Matt Turner |
ACL (1) | 4 |
| 2024 | The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"abstractWe expose a surprising failure of generalization in auto-regressive large language models (LLMs). If a model is trained on a sentence of the form ''_A_ is _B_'', it will not automatically generalize to the reverse direction ''_B_ is _A_''. This is the **Reversal Curse**. For instance, if a model is trained on ''Valentina Tereshkova was the first woman to travel to space'', it will not automatically be able to answer the question, ''Who was the first woman to travel to space?''. Moreover, the likelihood of the correct answer (''Valentina Tershkova'') will not be higher than for a random name. Thus, models do not generalize a prevalent pattern in their training set: if ''_A_ is _B_'' occurs, ''_B_ is _A_'' is more likely to occur. It is worth noting, however, that if ''_A_ is _B_'' appears _in-context_, models can deduce the reverse relationship.
We provide evidence for the Reversal Curse by finetuning GPT-3 and Llama-1 on fictitious statements such as ''Uriah Hawthorne is the composer of _Abyssal Melodies_'' and showing that they fail to correctly answer ''Who composed _Abyssal Melodies?_''. The Reversal Curse is robust across model sizes and model families and is not alleviated by data augmentation.
We also evaluate ChatGPT (GPT-3.5 and GPT-4) on questions about real-world celebrities, such as ''Who is Tom Cruise's mother? [A: Mary Lee Pfeiffer]'' and the reverse ''Who is Mary Lee Pfeiffer's son?''. GPT-4 correctly answers questions like the former 79\% of the time, compared to 33\% for the latter.
Code available at: https://github.com/lukasberglund/reversal_curse. Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, Owain Evans |
ICLR | 2 |
| 2024 | Towards Understanding Sycophancy in Language ModelsabstractReinforcement learning from human feedback (RLHF) is a popular technique for training high-quality AI assistants. However, RLHF may also encourage model responses that match user beliefs over truthful responses, a behavior known as sycophancy. We investigate the prevalence of sycophancy in RLHF-trained models and whether human preference judgments are responsible. We first demonstrate that five state-of-the-art AI assistants consistently exhibit sycophancy behavior across four varied free-form text-generation tasks. To understand if human preferences drive this broadly observed behavior of RLHF models, we analyze existing human preference data. We find that when a response matches a user's views, it is more likely to be preferred. Moreover, both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time. Optimizing model outputs against PMs also sometimes sacrifices truthfulness in favor of sycophancy. Overall, our results indicate that sycophancy is a general behavior of RLHF models, likely driven in part by human preference judgments favoring sycophantic responses. Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Miranda Zhang, Ethan Perez |
ICLR | 2 |
| 2024 | Many-shot JailbreakingabstractWe investigate a family of simple long-context attacks on large language models: prompting with hundreds of demonstrations of undesirable behavior. This attack is newly feasible with the larger context windows recently deployed by language model providers like Google DeepMind, OpenAI and Anthropic. We find that in diverse, realistic circumstances, the effectiveness of this attack follows a power law, up to hundreds of shots. We demonstrate the success of this attack on the most widely used state-of-the-art closed-weight models, and across various tasks. Our results suggest very long contexts present a rich new attack surface for LLMs. Cem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Meg Tong, Jesse Mu, Daniel Ford, Francesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Carson Denison, Evan Hubinger, Yuntao Bai, Trenton Bricken, Timothy Maxwell, Nicholas Schiefer, James Sully, Alex Tamkin, Tamera Lanham, Karina Nguyen, Tomek Korbak, Jared Kaplan, Deep Ganguli, Samuel R. Bowman, Ethan Perez, Roger B. Grosse, David Duvenaud |
NeurIPS | 8 |