Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yair Lakretz

dblp:166/5196 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-8774-6427ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 46% Transfer learning and domain adaptation · 17% Knowledge representation and reasoning · 17%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 61% Medical and health informatics · 39%

Topics — the 8 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
meta-learning
0.912025
Meta-Learning Neural Mechanisms rather than Bayesian Priors · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model
large language model representation
0.812024
A Polar coordinate system represents syntax in large language models · NeurIPS 2024
Natural language and speech › Language models and text generation › text representation
syntactic representation
0.812024
A Polar coordinate system represents syntax in large language models · NeurIPS 2024
Natural language and speech › Language models and text generation
neural language model
0.612022
Neural Language Models are not Born Equal to Fit Brain Data, but Training Helps · ICML 2022
Natural language and speech › Language models and text generation
language acquisition
0.312025
Meta-Learning Neural Mechanisms rather than Bayesian Priors · ACL (1) 2025
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.212015
Probabilistic Graphical Models of Dyslexia · KDD 2015
Bioinformatics and computational biology › computational neuroscience › neural coding
brain encoding
0.212022
Neural Language Models are not Born Equal to Fit Brain Data, but Training Helps · ICML 2022
Bioinformatics and computational biology
computational neuroscience
0.212022
Neural Language Models are not Born Equal to Fit Brain Data, but Training Helps · ICML 2022

Methods — techniques the papers use, named apart from their topics

representational similarity analysis · 1.1fMRI encoding models · 1.1neural network analysis · 0.9meta-learning · 0.9structural probe · 0.8probing classifier · 0.8polar probe · 0.8naive bayes · 0.4latent dirichlet allocation · 0.4
YearPublicationVenuePosition
2025 Meta-Learning Neural Mechanisms rather than Bayesian Priors
abstract
Children acquire language despite being exposed to several orders of magnitude less data than large language models require.Metalearning has been proposed as a way to integrate human-like learning biases into neuralnetwork architectures, combining both the structured generalizations of symbolic models with the scalability of neural-network models.But what does meta-learning exactly imbue the model with?We investigate the meta-learning of formal languages and find that, contrary to previous claims, meta-trained models are not learning simplicity-based priors when metatrained on datasets organised around simplicity.Rather, we find evidence that meta-training imprints neural mechanisms (such as counters) into the model, which function like cognitive primitives for the network on downstream tasks.Most surprisingly, we find that meta-training on a single formal language can provide as much improvement to a model as meta-training on 5000 different formal languages, provided that the formal language incentivizes the learning of useful neural mechanisms.Taken together, our findings provide practical implications for efficient meta-learning paradigms and new theoretical insights into linking symbolic theories and neural mechanisms.
Michael Eric Goodale, Salvador Mascarenhas, Yair Lakretz
ACL (1)3
2024 Automatic Recognition of Gesture Identity and Onset of Cued-Speech
abstract
Cued speech is a communication system based on hand gestures used in certain communities of deaf people around the world. Hand gestures in cued speech convey complementary information to that available from lip-reading alone, which helps to communicate accurate phonological information. Cued speech provides several unique opportunities for scientific research, such as studying phonological processing via the visual modality. However, cued speech has been studied only scarcely since its invention, and there are only a few empirical datasets and standardized methods available to the scientific community. Here, we suggest several contributions to advance research in the field: (1) A new dataset on cued speech annotated for various linguistic features, (2) a new approach to automatically identify gesture identity and gesture onset from raw videos, using cuedspeech-specific features, which achieves relatively high performance (AUCidentity> 0.95, Erroronset= 93ms), and (3) additional insights into the relationship between sound and gesture production in cued speech, showing that syllable acoustic onset precedes gesture onset by around 150ms on average and that this time difference is more sensitive to consonant rather than to vowel identity. We make the new dataset and all associated tools publicly available.
Annahita Sarré, Hagar Salpeter, Deliane Bechar, Laurent Cohen, Yair Lakretz
ICASSP5
2024 A Polar coordinate system represents syntax in large language models
abstract
Originally formalized with symbolic representations, syntactic trees may also be effectively represented in the activations of large language models (LLMs). Indeed, a ''Structural Probe'' can find a subspace of neural activations, where syntactically-related words are relatively close to one-another. However, this syntactic code remains incomplete: the distance between the Structural Probe word embeddings can represent the \emph{existence} but not the type and direction of syntactic relations. Here, we hypothesize that syntactic relations are, in fact, coded by the relative direction between nearby embeddings. To test this hypothesis, we introduce a ''Polar Probe'' trained to read syntactic relations from both the distance and the direction between word embeddings. Our approach reveals three main findings. First, our Polar Probe successfully recovers the type and direction of syntactic relations, and substantially outperforms the Structural Probe by nearly two folds. Second, we confirm that this polar coordinate system exists in a low-dimensional subspace of the intermediate layers of many LLMs and becomes increasingly precise in the latest frontier models. Third, we demonstrate with a new benchmark that similar syntactic relations are coded similarly across the nested levels of syntactic trees. Overall, this work shows that LLMs spontaneously learn a geometry of neural activations that explicitly represents the main symbolic structures of linguistic theory.
Pablo Diego-Simón, Stéphane d'Ascoli, Emmanuel Chemla, Yair Lakretz, Jean-Rémi King
NeurIPS4
2022 Can Transformers Process Recursive Nested Constructions, Like Humans?
abstract
Recursive processing is considered a hallmark of human linguistic abilities. A recent study evaluated recursive processing in recurrent neural language models (RNN-LMs) and showed that such models perform below chance level on embedded dependencies within nested constructions – a prototypical example of recursion in natural language. Here, we study if state-of-the-art Transformer LMs do any better. We test eight different Transformer LMs on two different types of nested constructions, which differ in whether the embedded (inner) dependency is short or long range. We find that Transformers achieve near-perfect performance on short-range embedded dependencies, significantly better than previous results reported for RNN-LMs and humans. However, on long-range embedded dependencies, Transformers’ performance sharply drops below chance level. Remarkably, the addition of only three words to the embedded dependency caused Transformers to fall from near-perfect to below-chance performance. Taken together, our results reveal how brittle syntactic processing is in Transformers, compared to humans.
Yair Lakretz, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene
COLING1
2022 Neural Language Models are not Born Equal to Fit Brain Data, but Training Helps
abstract
Neural Language Models (NLMs) have made tremendous advances during the last years, achieving impressive performance on various linguistic tasks. Capitalizing on this, studies in neuroscience have started to use NLMs to study neural activity in the human brain during language processing. However, many questions remain unanswered regarding which factors determine the ability of a neural language model to capture brain activity (aka its ’brain score’). Here, we make first steps in this direction and examine the impact of test loss, training corpus and model architecture (comparing GloVe, LSTM, GPT-2 and BERT), on the prediction of functional Magnetic Resonance Imaging time-courses of participants listening to an audiobook. We find that (1) untrained versions of each model already explain significant amount of signal in the brain by capturing similarity in brain responses across identical words, with the untrained LSTM outperforming the transformer-based models, being less impacted by the effect of context; (2) that training NLP models improves brain scores in the same brain regions irrespective of the model’s architecture; (3) that Perplexity (test loss) is not a good predictor of brain score; (4) that training data have a strong influence on the outcome and, notably, that off-the-shelf models may lack statistical power to detect brain activations. Overall, we outline the impact of model-training choices, and suggest good practices for future studies aiming at explaining the human language system using neural language models.
Alexandre Pasquiou, Yair Lakretz, John T. Hale, Bertrand Thirion, Christophe Pallier
ICML2
2015 Probabilistic Graphical Models of Dyslexia
abstract
Reading is a complex cognitive process, errors in which may assume diverse forms. In this study, introducing a novel approach, we use two families of probabilistic graphical models to analyze patterns of reading errors made by dyslexic people: an LDA-based model and two Naëve Bayes models which differ by their assumptions about the generation process of reading errors. The models are trained on a large corpus of reading errors. Results show that a Naëve Bayes model achieves highest accuracy compared to labels given by clinicians (AUC = 0.801 ± 0.05), thus providing the first automated and objective diagnosis tool for dyslexia which is solely based on reading errors data. Results also show that the LDA-based model best captures patterns of reading errors and could therefore contribute to the understanding of dyslexia and to future improvement of the diagnostic procedure. Finally, we draw on our results to shed light on a theoretical debate about the definition and heterogeneity of dyslexia. Our results support a model assuming multiple dyslexia subtypes, that of a heterogeneous view of dyslexia.
Yair Lakretz, Gal Chechik, Naama Friedmann, Michal Rosen-Zvi
KDD1