Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ouail Kitouni

dblp:285/7983 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 36% Deep learning architectures and training · 24% Language models and text generation · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.422024
From Neurons to Neutrons: A Case Study in Interpretability · ICML 2024
Expressive Monotonic Neural Networks · ICLR 2023
Natural language and speech › Information extraction and text analysis › information retrieval
knowledge retrieval
0.812024
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model
0.812024
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.812024
From Neurons to Neutrons: A Case Study in Interpretability · ICML 2024
Machine learning › Generative modeling › autoregressive model
next-token prediction
0.812024
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More · NeurIPS 2024
Machine learning › Deep learning architectures and training
training objective
0.812024
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › explainable AI › interpretable neural network
monotonic neural networks
0.712023
Expressive Monotonic Neural Networks · ICLR 2023
Machine learning › Deep learning architectures and training › training dynamics
grokking
0.612022
Towards Understanding Grokking: An Effective Theory of Representation Learning · NeurIPS 2022
Machine learning › Representation and self-supervised learning
structured representation
0.612022
Towards Understanding Grokking: An Effective Theory of Representation Learning · NeurIPS 2022
Machine learning › Deep learning architectures and training
training dynamics
0.612022
Towards Understanding Grokking: An Effective Theory of Representation Learning · NeurIPS 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.212024
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

mechanistic interpretability · 1.5controlled experiment · 0.8bidirectional-attention training · 0.8monotonic neural network · 0.7phase diagrams · 0.6effective theory · 0.6
YearPublicationVenuePosition
2024 From Neurons to Neutrons: A Case Study in Interpretability
abstract
Mechanistic Interpretability (MI) proposes a path toward fully understanding how neural networks make their predictions. Prior work demonstrates that even when trained to perform simple arithmetic, models can implement a variety of algorithms (sometimes concurrently) depending on initialization and hyperparameters. Does this mean neuron-level interpretability techniques have limited applicability? Here, we argue that high-dimensional neural networks can learn useful low-dimensional representations of the data they were trained on, going beyond simply making good predictions: Such representations can be understood with the MI lens and provide insights that are surprisingly faithful to human-derived domain knowledge. This indicates that such approaches to interpretability can be useful for deriving a new understanding of a problem from models trained to solve it. As a case study, we extract nuclear physics concepts by studying models trained to reproduce nuclear data.
Ouail Kitouni, Niklas Nolte, Víctor Samuel Pérez-Díaz, Sokratis Trifinopoulos, Mike Williams
ICML1
2024 The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
abstract
Today's best language models still struggle with "hallucinations", factually incorrect generations, which impede their ability to reliably retrieve information seen during training. The *reversal curse*, where models cannot recall information when probed in a different order than was encountered during training, exemplifies limitations in information retrieval. To better understand these limitations, we reframe the reversal curse as a *factorization curse* --- a failure of models to learn the same joint distribution under different factorizations. We more closely simulate finetuning workflows which train pretrained models on specialized knowledge by introducing *WikiReversal*, a realistic testbed based on Wikipedia knowledge graphs. Through a series of controlled experiments with increasing levels of realism, including non-reciprocal relations, we find that reliable information retrieval is an inherent failure of the next-token prediction objective used in popular large language models. Moreover, we demonstrate reliable information retrieval cannot be solved with scale, reversed tokens, or even naive bidirectional-attention training. Consequently, various approaches to finetuning on specialized data would necessarily provide mixed results on downstream tasks, unless the model has already seen the right sequence of tokens. Across five tasks of varying levels of complexity, our results uncover a promising path forward: factorization-agnostic objectives can significantly mitigate the reversal curse and hint at improved knowledge storage and planning capabilities.
Ouail Kitouni, Niklas Nolte, Adina Williams, Michael G. Rabbat, Diane Bouchacourt, Mark Ibrahim
NeurIPS1
2023 Expressive Monotonic Neural Networks
Niklas Nolte, Ouail Kitouni, Mike Williams
ICLR2
2022 Towards Understanding Grokking: An Effective Theory of Representation Learning
abstract
We aim to understand grokking, a phenomenon where models generalize long after overfitting their training set. We present both a microscopic analysis anchored by an effective theory and a macroscopic analysis of phase diagrams describing learning performance across hyperparameters. We find that generalization originates from structured representations, whose training dynamics and dependence on training set size can be predicted by our effective theory (in a toy setting). We observe empirically the presence of four learning phases: comprehension, grokking, memorization, and confusion. We find representation learning to occur only in a "Goldilocks zone" (including comprehension and grokking) between memorization and confusion. Compared to the comprehension phase, the grokking phase stays closer to the memorization phase, leading to delayed generalization. The Goldilocks phase is reminiscent of "intelligence from starvation" in Darwinian evolution, where resource limitations drive discovery of more efficient solutions. This study not only provides intuitive explanations of the origin of grokking, but also highlights the usefulness of physics-inspired tools, e.g., effective theories and phase diagrams, for understanding deep learning.
Ziming Liu 0001, Ouail Kitouni, Niklas Nolte, Eric J. Michaud, Max Tegmark, Mike Williams
NeurIPS2