VLDB 2026 Research / reviewers in the wild / expert
Julian Michael
dblp:185/0981
· DBLP profile ↗
18ranked-venue papers
2as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Quantifying Elicitation of Latent Capabilities in Language ModelsabstractLarge language models often possess latent capabilities that lie dormant unless explicitly elicited, or surfaced, through fine-tuning or prompt engineering. Predicting, assessing, and understanding these latent capabilities pose significant challenges in the development of effective, safe AI systems. In this work, we recast elicitation as an information-constrained fine-tuning problem and empirically characterize upper bounds on the minimal number of parameters needed to achieve specific task performances. We find that training as few as 10–100 randomly chosen parameters—several orders of magnitude fewer than state-of-the-art parameter-efficient methods—can recover up to 50\% of the performance gap between pretrained-only and full fine-tuned models, and 1,000s to 10,000s of parameters can recover 95\% of this performance gap. We show that a logistic curve fits the relationship between the number of trained parameters and model performance gap recovery. This scaling generalizes across task formats and domains, as well as model sizes and families, extending to reasoning models and remaining robust to increases in inference compute. To help explain this behavior, we consider a simplified picture of elicitation via fine-tuning where each trainable parameter serves as an encoding mechanism for accessing task-specific knowledge. We observe a relationship between the number of trained parameters and how efficiently relevant model capabilities can be accessed and elicited, offering a potential route to distinguish elicitation from teaching. Elizabeth Donoway, Hailey Joren, Arushi Somani, Henry Sleight, Julian Michael, Michael R. DeWeese, John Schulman, Ethan Perez, Fabien Roger, Jan Leike |
NeurIPS | 5 |
| 2025 | AI Debate Aids Assessment of Controversial ClaimsabstractAs AI grows more powerful, it will increasingly shape how we understand the world. But with this influence comes the risk of amplifying misinformation and deepening social divides—especially on consequential topics where factual accuracy directly impacts well-being. Scalable Oversight aims to ensure AI systems remain truthful even when their capabilities exceed those of their evaluators. Yet when humans serve as evaluators, their own beliefs and biases can impair judgment. We study whether AI debate can guide biased judges toward the truth by having two AI systems debate opposing sides of controversial factuality claims on COVID-19 and climate change where people hold strong prior beliefs. We conduct two studies. Study I recruits human judges with either mainstream or skeptical beliefs who evaluate claims through two protocols: debate (interaction with two AI advisors arguing opposing sides) or consultancy (interaction with a single AI advisor). Study II uses AI judges with and without human-like personas to evaluate the same protocols. In Study I, debate consistently improves human judgment accuracy and confidence calibration, outperforming consultancy by 4-10\% across COVID-19 and climate change claims. The improvement is most significant for judges with mainstream beliefs (up to +15.2\% accuracy on COVID-19 claims), though debate also helps skeptical judges who initially misjudge claims move toward accurate views (+4.7\% accuracy). In Study II, AI judges with human-like personas achieve even higher accuracy (78.5\%) than human judges (70.1\%) and default AI judges without personas (69.8\%), suggesting their potential for supervising frontier AI models. These findings highlight AI debate as a promising path toward scalable, bias-resilient oversight in contested domains. Salman Rahman, Sheriff Issaka, Ashima Suvarna, Genglin Liu, James Shiffer, Md. Rizwan Parvez, Hamid Palangi, Nanyun Peng 0001, Yejin Choi 0001, Julian Michael, Saadia Gabriel |
NeurIPS | 12 |
| 2025 | Why Do Some Language Models Fake Alignment While Others Don't?abstract*Alignment Faking in Large Language Models* presented a demonstration of Claude 3 Opus and Claude 3.5 Sonnet selectively complying with a helpful-only training objective to prevent modification of their behavior outside of training. We expand this analysis to 25 models and find that only 5 (Claude 3 Opus, Claude 3.5 Sonnet, Llama 3 405B, Grok 3, Gemini 2.0 Flash) comply with harmful queries more when they infer they are in training than when they infer they are in deployment. First, we study the motivations of these 5 models. Results from perturbing details of the scenario suggest that only Claude 3 Opus's compliance gap is primarily and consistently motivated by trying to keep its goals. Second, we investigate why many chat models don't fake alignment. Our results suggest this is not entirely due to a lack of capabilities: many base models fake alignment some of the time, and post-training eliminates alignment-faking for some models and amplifies it for others. We investigate 5 hypotheses for how post-training may suppress alignment faking and find that variations in refusal behavior may account for a significant portion of differences in alignment faking. Abhay Sheshadri, Julian Michael, Alex Mallen, Arun Jose, Fabien Roger |
NeurIPS | 3 |
| 2024 | Analyzing the Role of Semantic Representations in the Era of Large Language ModelsabstractZhijing Jin, Yuen Chen, Fernando Gonzalez Adauto, Jiarui Liu, Jiayi Zhang, Julian Michael, Bernhard Schölkopf, Mona Diab. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhijing Jin 0001, Yuen Chen, Fernando Gonzalez Adauto, Jiarui Liu 0004, Julian Michael, Bernhard Schölkopf, Mona T. Diab |
NAACL-HLT | 6 |
| 2023 | What Do NLP Researchers Believe? Results of the NLP Community MetasurveyabstractJulian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman |
ACL (1) | 1 |
| 2023 | We're Afraid Language Models Aren't Modeling AmbiguityabstractAlisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah Smith, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Alisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah A. Smith, Yejin Choi 0001 |
EMNLP | 3 |
| 2023 | Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingabstractLarge Language Models (LLMs) can achieve strong performance on many tasks by producing step-by-step reasoning before giving a final output, often referred to as chain-of-thought reasoning (CoT). It is tempting to interpret these CoT explanations as the LLM's process for solving a task. This level of transparency into LLMs' predictions would yield significant safety benefits. However, we find that CoT explanations can systematically misrepresent the true reason for a model's prediction. We demonstrate that CoT explanations can be heavily influenced by adding biasing features to model inputs—e.g., by reordering the multiple-choice options in a few-shot prompt to make the answer always "(A)"—which models systematically fail to mention in their explanations. When we bias models toward incorrect answers, they frequently generate CoT explanations rationalizing those answers. This causes accuracy to drop by as much as 36% on a suite of 13 tasks from BIG-Bench Hard, when testing with GPT-3.5 from OpenAI and Claude 1.0 from Anthropic. On a social-bias task, model explanations justify giving answers in line with stereotypes without mentioning the influence of these social biases. Our findings indicate that CoT explanations can be plausible yet misleading, which risks increasing our trust in LLMs without guaranteeing their safety. Building more transparent and explainable systems will require either improving CoT faithfulness through targeted efforts or abandoning CoT in favor of alternative methods. Miles Turpin, Julian Michael, Ethan Perez, Samuel R. Bowman |
NeurIPS | 2 |
| 2022 | Nearest Neighbor Zero-Shot InferenceabstractRetrieval-augmented language models (LMs) use non-parametric memory to substantially outperform their non-retrieval counterparts on perplexity-based evaluations, but it is an open question whether they achieve similar gains in few- and zero-shot end-task accuracy. We extensively study one such model, the k-nearest neighbor LM (kNN-LM), showing that the gains marginally transfer. The main challenge is to achieve coverage of the verbalizer tokens that define the different end-task class labels. To address this challenge, we also introduce kNN-Prompt, a simple and effective kNN-LM with automatically expanded fuzzy verbalizers (e.g. to expand "terrible" to also include "silly" and other task-specific synonyms for sentiment classification). Across nine diverse end-tasks, using kNN-Prompt with GPT-2 large yields significant performance boosts over strong zeroshot baselines (13.4% absolute improvement over the base LM on average). We also show that other advantages of non-parametric augmentation hold for end tasks; kNN-Prompt is effective for domain adaptation with no further training, and gains increase with the size of the retrieval model. Julian Michael, Suchin Gururangan, Luke Zettlemoyer |
EMNLP | 2 |
| 2021 | Asking It All: Generating Contextualized Questions for any Semantic RoleabstractAsking questions about a situation is an inherent step towards understanding it.To this end, we introduce the task of role question generation, which, given a predicate mention and a passage, requires producing a set of questions asking about all possible semantic roles of the predicate.We develop a two-stage model for this task, which first produces a contextindependent question prototype for each role and then revises it to be contextually appropriate for the passage.Unlike most existing approaches to question generation, our approach does not require conditioning on existing answers in the text.Instead, we condition on the type of information to inquire about, regardless of whether the answer appears explicitly in the text, could be inferred from it, or should be sought elsewhere.Our evaluation demonstrates that we generate diverse and well-formed questions for a large, broadcoverage ontology of predicates and roles. Valentina Pyatkin, Paul Roit, Julian Michael, Yoav Goldberg, Reut Tsarfaty, Ido Dagan |
EMNLP (1) | 3 |
| 2020 | Controlled Crowdsourcing for High-Quality QA-SRL AnnotationabstractPaul Roit, Ayal Klein, Daniela Stepanov, Jonathan Mamou, Julian Michael, Gabriel Stanovsky, Luke Zettlemoyer, Ido Dagan. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Paul Roit, Ayal Klein, Daniela Stepanov, Jonathan Mamou, Julian Michael, Gabriel Stanovsky, Luke Zettlemoyer, Ido Dagan |
ACL | 5 |
| 2020 | Asking without Telling: Exploring Latent Ontologies in Contextual RepresentationsabstractThe success of pretrained contextual encoders, such as ELMo and BERT, has brought a great deal of interest in what these models learn: do they, without explicit supervision, learn to encode meaningful notions of linguistic structure?If so, how is this structure encoded?To investigate this, we introduce latent subclass learning (LSL): a modification to classifierbased probing that induces a latent categorization (or ontology) of the probe's inputs.Without access to fine-grained gold labels, LSL extracts emergent structure from input representations in an interpretable and quantifiable form.In experiments, we find strong evidence of familiar categories, such as a notion of personhood in ELMo, as well as novel ontological distinctions, such as a preference for fine-grained semantic roles on core arguments.Our results provide unique new evidence of emergent structure in pretrained encoders, including departures from existing annotations which are inaccessible to earlier methods. Julian Michael, Jan A. Botha, Ian Tenney |
EMNLP (1) | 1 |
| 2020 | AmbigQA: Answering Ambiguous Open-domain QuestionsabstractAmbiguity is inherent to open-domain question answering; especially when exploring new topics, it can be difficult to ask questions that have a single, unambiguous answer.In this paper, we introduce AMBIGQA, a new open-domain question answering task which involves finding every plausible answer, and then rewriting the question for each one to resolve the ambiguity.To study this task, we construct AMBIGNQ, a dataset covering 14,042 questions from NQ-OPEN, an existing opendomain QA benchmark.We find that over half of the questions in NQ-OPEN are ambiguous, with diverse sources of ambiguity such as event and entity references.We also present strong baseline models for AMBIGQA which we show benefit from weakly supervised learning that incorporates NQ-OPEN, strongly suggesting our new task and data will support significant future research effort.Our data and baselines are available at https://nlp.cs. washington.edu/ambigqa.Type Example Event references (39%) What season does meredith and derek get married in grey's anatomy?Q: In what season do Meredith and Derek get informally married in Grey's Anatomy? / A: Season 5 Q: In what season do Meredith and Derek get legally married in Grey's Anatomy? / A: Season 7 Properties (27%) How many episode in seven deadly sins season 2? Q: How many episodes were there in seven deadly sins season 2, not including the OVA episode?/ A: 25 Q: How many episodes were there in seven deadly sins season 2, including the OVA episode?/ A: 26 Entity references (23%) How many sacks does clay matthews have in his career?Q: How many sacks does Clay Matthews Jr. have in his career?/ A: 69.5 Q: How many sacks does Clay Matthews III have in his career?/ A: 91.5 Answer types (16%) Who sings the song what a beautiful name it is?Q: Which group sings the song what a beautiful name it is?/ A: Hillsong Live Q: Who is the lead singer of the song what a beautiful name it is?/ A: Brooke Ligertwood Sewon Min, Julian Michael, Hannaneh Hajishirzi, Luke Zettlemoyer |
EMNLP (1) | 2 |
| 2019 | GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel R. Bowman |
ICLR (Poster) | 3 |
| 2019 | SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding SystemsabstractIn the last year, new models and methods for pretraining and transfer learning have driven striking performance improvements across a range of language understanding tasks. The GLUE benchmark, introduced a little over one year ago, offers a single-number metric that summarizes progress on a diverse set of such tasks, but performance on the benchmark has recently surpassed the level of non-expert humans, suggesting limited headroom for further research. In this paper we present SuperGLUE, a new benchmark styled after GLUE with a new set of more difficult language understanding tasks, a software toolkit, and a public leaderboard. SuperGLUE is available at https://super.gluebenchmark.com. Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel R. Bowman |
NeurIPS | 5 |
| 2018 | Large-Scale QA-SRL ParsingabstractWe present a new large-scale corpus of Question-Answer driven Semantic Role Labeling (QA-SRL) annotations, and the first high-quality QA-SRL parser.Our corpus, QA-SRL Bank 2.0, consists of over 250,000 question-answer pairs for over 64,000 sentences across 3 domains and was gathered with a new crowd-sourcing scheme that we show has high precision and good recall at modest cost.We also present neural models for two QA-SRL subtasks: detecting argument spans for a predicate and generating questions to label the semantic relationship.The best models achieve question accuracy of 82.6% and span-level accuracy of 77.6% (under human evaluation) on the full pipelined QA-SRL prediction task.They can also, as we show, be used to gather additional annotations at low cost. Nicholas FitzGerald, Julian Michael, Luheng He, Luke Zettlemoyer |
ACL (1) | 2 |
| 2018 | Supervised Open Information ExtractionabstractGabriel Stanovsky, Julian Michael, Luke Zettlemoyer, Ido Dagan. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Gabriel Stanovsky, Julian Michael, Luke Zettlemoyer, Ido Dagan |
NAACL-HLT | 2 |
| 2016 | Human-in-the-Loop ParsingabstractThis paper demonstrates that it is possible for a parser to improve its performance with a human in the loop, by posing simple questions to non-experts.For example, given the first sentence of this abstract, if the parser is uncertain about the subject of the verb "pose," it could generate the question What would pose something?with candidate answers this paper and a parser.Any fluent speaker can answer this question, and the correct answer resolves the original uncertainty.We apply the approach to a CCG parser, converting uncertain attachment decisions into natural language questions about the arguments of verbs.Experiments show that crowd workers can answer these questions quickly, accurately and cheaply.Our human-in-the-loop parser improves on the state of the art with less than 2 questions per sentence on average, with a gain of 1.7 F1 on the 10% of sentences whose parses are changed. Luheng He, Julian Michael, Mike Lewis, Luke Zettlemoyer |
EMNLP | 2 |
| 2016 | Proving infinitary formulasabstractAbstract The infinitary propositional logic of here-and-there is important for the theory of answer set programming in view of its relation to strongly equivalent transformations of logic programs. We know a formal system axiomatizing this logic exists, but a proof in that system may include infinitely many formulas. In this note we describe a relationship between the validity of infinitary formulas in the logic of here-and-there and the provability of formulas in some finite deductive systems. This relationship allows us to use finite proofs to justify the validity of infinitary formulas. Amelia Harrison, Vladimir Lifschitz, Julian Michael |
Theory Pract. Log. Program. | 3 |