EDBT 2026 Demo / reviewers in the wild / expert
Emmanuel Chemla
dblp:85/8850
· DBLP profile ↗
7ranked-venue papers
1as first author
4since 2021 · last 2024
0000-0002-8423-5880ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Speech recognition and synthesis · 28% Language models and text generation · 28% Multi-agent systems · 16% | |
| Theoretical computer science
1 paper |
Automata and formal languages · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory
generalization |
0.8 | 1 | 2024 | Bridging the Empirical-Theoretical Gap in Neural Network Formal Language Learning Using Minimum Description Length · ACL (1) 2024 |
Natural language and speech › Language models and text generation › large language model
large language model representation |
0.8 | 1 | 2024 | A Polar coordinate system represents syntax in large language models · NeurIPS 2024 |
Natural language and speech › Speech recognition and synthesis
speech language model |
0.8 | 1 | 2024 | Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach · EMNLP 2024 |
Natural language and speech › Speech recognition and synthesis
speech representation learning |
0.8 | 1 | 2024 | Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach · EMNLP 2024 |
Natural language and speech › Language models and text generation › text representation
syntactic representation |
0.8 | 1 | 2024 | A Polar coordinate system represents syntax in large language models · NeurIPS 2024 |
Automata and formal languages › language learning
formal language learning |
0.8 | 1 | 2024 | Bridging the Empirical-Theoretical Gap in Neural Network Formal Language Learning Using Minimum Description Length · ACL (1) 2024 |
Knowledge, reasoning and agents › Multi-agent systems › emergent communication
language emergence |
0.4 | 1 | 2020 | On the Spontaneous Emergence of Discrete and Compositional Signals · ACL 2020 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
signaling games |
0.4 | 1 | 2020 | On the Spontaneous Emergence of Discrete and Compositional Signals · ACL 2020 |
Methods — techniques the papers use, named apart from their topics
regularization · 1.5minimum description length · 1.5structural probe · 0.8probing classifier · 0.8polar probe · 0.8phoneme classification · 0.8fine-tuning · 0.8neural agent · 0.4backpropagation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Bridging the Empirical-Theoretical Gap in Neural Network Formal Language Learning Using Minimum Description LengthabstractNeural networks offer good approximation to many tasks but consistently fail to reach perfect generalization, even when theoretical work shows that such perfect solutions can be expressed by certain architectures.Using the task of formal language learning, we focus on one simple formal language and show that the theoretically correct solution is in fact not an optimum of commonly used objectiveseven with regularization techniques that according to common wisdom should lead to simple weights and good generalization (L1, L2) or other meta-heuristics (early-stopping, dropout).On the other hand, replacing standard targets with the Minimum Description Length objective (MDL) results in the correct solution being an optimum. Nur Geffen Lan, Emmanuel Chemla, Roni Katzir |
ACL (1) | 2 |
| 2024 | Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning ApproachabstractRecent progress in Spoken Language Modeling has shown that learning language directly from speech is feasible.Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and nonverbal vocalizations.Modeling directly from speech opens up the path to more natural and expressive systems.On the other hand, speechonly systems require up to three orders of magnitude more data to catch up to their text-based counterparts in terms of their semantic abilities.We show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data. Maxime Poli, Emmanuel Chemla, Emmanuel Dupoux |
EMNLP | 2 |
| 2024 | A Polar coordinate system represents syntax in large language modelsabstractOriginally formalized with symbolic representations, syntactic trees may also be effectively represented in the activations of large language models (LLMs). Indeed, a ''Structural Probe'' can find a subspace of neural activations, where syntactically-related words are relatively close to one-another. However, this syntactic code remains incomplete: the distance between the Structural Probe word embeddings can represent the \emph{existence} but not the type and direction of syntactic relations. Here, we hypothesize that syntactic relations are, in fact, coded by the relative direction between nearby embeddings. To test this hypothesis, we introduce a ''Polar Probe'' trained to read syntactic relations from both the distance and the direction between word embeddings. Our approach reveals three main findings. First, our Polar Probe successfully recovers the type and direction of syntactic relations, and substantially outperforms the Structural Probe by nearly two folds. Second, we confirm that this polar coordinate system exists in a low-dimensional subspace of the intermediate layers of many LLMs and becomes increasingly precise in the latest frontier models. Third, we demonstrate with a new benchmark that similar syntactic relations are coded similarly across the nested levels of syntactic trees. Overall, this work shows that LLMs spontaneously learn a geometry of neural activations that explicitly represents the main symbolic structures of linguistic theory. Pablo Diego-Simón, Stéphane d'Ascoli, Emmanuel Chemla, Yair Lakretz, Jean-Rémi King |
NeurIPS | 3 |
| 2022 | Minimum Description Length Recurrent Neural NetworksabstractAbstract We train neural networks to optimize a Minimum Description Length score, that is, to balance between the complexity of the network and its accuracy at a task. We show that networks optimizing this objective function master tasks involving memory challenges and go beyond context-free languages. These learners master languages such as anbn, anbncn, anb2n, anbmcn +m, and they perform addition. Moreover, they often do so with 100% accuracy. The networks are small, and their inner workings are transparent. We thus provide formal proofs that their perfect accuracy holds not only on a given test set, but for any input sequence. To our knowledge, no other connectionist model has been shown to capture the underlying grammars for these languages in full generality. Nur Geffen Lan, Michal Geyer, Emmanuel Chemla, Roni Katzir |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | On the Spontaneous Emergence of Discrete and Compositional SignalsabstractWe propose a general framework to study language emergence through signaling games with neural agents.Using a continuous latent space, we are able to (i) train using backpropagation, (ii) show that discrete messages nonetheless naturally emerge.We explore whether categorical perception effects follow and show that the messages are not compositional. Nur Geffen Lan, Emmanuel Chemla, Shane Steinert-Threlkeld |
ACL | 2 |
| 2017 | Characterizing logical consequence in many-valued logicabstractSeveral definitions of logical consequence have been proposed in many-valued logic, which coincide in the two-valued case, but come apart as soon as three truth values come into play. Those definitions include so-called pure consequence, order-theoretic consequence and mixed consequence. In this article, we examine whether those definitions together carve out a natural class of consequence relations. We respond positively by identifying a small set of properties that we see instantiated in those various consequence relations, namely truth-relationality, value-monotonicity, validity-coherence and a constraint of bivalence-compliance, provably replaceable by a structural requisite of nontriviality. Our main result is that the class of consequence relations satisfying those properties coincides exactly with the class of mixed consequence relations and their intersections, including pure consequence relations and order-theoretic consequence. We provide an enumeration of the set of those relations in finite many-valued logics of two extreme kinds: those in which truth values are well-ordered and those in which values between 0 and 1 are incomparable. Emmanuel Chemla, Paul Egré, Benjamin Spector |
J. Log. Comput. | 1 |
| 2013 | Pragmatic priming and the search for alternatives
Lewis Bott, Emmanuel Chemla |
CogSci | 2 |