Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Benjamin Muller

dblp:304/3570 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 46% Transfer learning and domain adaptation · 16% Question answering and dialogue systems · 13%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › language modeling
byte-level language model
0.912025
Byte Latent Transformer: Patches Scale Better Than Tokens · ACL (1) 2025
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.912025
Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models · ICLR 2025
Natural language and speech › Language models and text generation › language modeling
language model architecture
0.912025
Byte Latent Transformer: Patches Scale Better Than Tokens · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model
0.912025
Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models · ICLR 2025
Natural language and speech › Language models and text generation
mathematical reasoning
0.912025
Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models · ICLR 2025
Machine learning › Efficient and distributed learning
model merging
0.912025
Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models · ICLR 2025
Natural language and speech › Language models and text generation
tokenization
0.912025
Byte Latent Transformer: Patches Scale Better Than Tokens · ACL (1) 2025
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.912025
Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models · ICLR 2025
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.812024
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants · ACL (1) 2024
Machine learning › Trustworthy machine learning › interpretability › explanation evaluation
attribution evaluation
0.712023
Evaluating and Modeling Attribution for Cross-Lingual Question Answering · EMNLP 2023
Natural language and speech › Question answering and dialogue systems › multilingual question answering
cross-lingual question answering
0.712023
Evaluating and Modeling Attribution for Cross-Lingual Question Answering · EMNLP 2023
Machine learning › Trustworthy machine learning
interpretability
0.712023
Evaluating and Modeling Attribution for Cross-Lingual Question Answering · EMNLP 2023
Natural language and speech › Language models and text generation
pre-trained language model
0.412020
CamemBERT: a Tasty French Language Model · ACL 2020
Natural language and speech › Information extraction and text analysis › data annotation › corpus annotation
treebank construction
0.412020
Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell · ACL 2020
Information retrieval
evaluation
0.212024
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants · ACL (1) 2024
Information retrieval › evaluation › benchmark
multilingual benchmark
0.212024
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants · ACL (1) 2024
Natural language and speech › Language models and text generation
masked language modeling
0.112020
CamemBERT: a Tasty French Language Model · ACL 2020
Computational social science and digital humanities
user-generated content
0.112020
Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell · ACL 2020

Methods — techniques the papers use, named apart from their topics

dataset construction · 1.5model merging · 0.9layer swapping · 0.9latent transformer · 0.9fine-tuning · 0.9masked language modeling · 0.4
YearPublicationVenuePosition
2025 Byte Latent Transformer: Patches Scale Better Than Tokens
abstract
Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason E Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srini Iyer. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez 0001, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srinivasan Iyer 0001
ACL (1)5
2025 Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models
abstract
Model merging, such as model souping, is the practice of combining different models with the same architecture together without further training. In this work, we present a model merging methodology that addresses the difficulty of fine-tuning Large Language Models (LLMs) for target tasks in non-English languages, where task-specific data is often unavailable. We focus on mathematical reasoning and without in-language math data, facilitate cross-lingual transfer by composing language and math capabilities. Starting from the same pretrained model, we fine-tune separate "experts" on math instruction data in English and on generic instruction data in the target language. We then replace the top and bottom transformer layers of the math expert directly with layers from the language expert, which consequently enhances math performance in the target language. The resulting merged models outperform the individual experts and other merging methods on the math benchmark, MGSM, by 10% across four major languages where math instruction data is scarce. In addition, this layer swapping is simple, inexpensive, and intuitive, as it is based on an interpretative analysis of the most important parameter changes during the fine-tuning of each expert. The ability to successfully re-compose LLMs for cross-lingual transfer in this manner opens up future possibilities to combine model expertise, create modular solutions, and transfer reasoning capabilities across languages all post hoc.
Lucas Bandarkar, Benjamin Muller, Pritish Yuvraj, Nayan Singhal, Hongjiang Lv
ICLR2
2024 The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
abstract
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, Madian Khabsa. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal 0001, Abhinandan Krishnan, Luke Zettlemoyer, Madian Khabsa
ACL (1)3
2023 Evaluating and Modeling Attribution for Cross-Lingual Question Answering
abstract
Benjamin Muller, John Wieting, Jonathan Clark, Tom Kwiatkowski, Sebastian Ruder, Livio Soares, Roee Aharoni, Jonathan Herzig, Xinyi Wang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Benjamin Muller, John Wieting, Jonathan H. Clark, Tom Kwiatkowski, Sebastian Ruder, Livio B. Soares, Roee Aharoni, Jonathan Herzig, Xinyi Wang 0001
EMNLP1
2021 First Align, then Predict: Understanding the Cross-Lingual Ability of Multilingual BERT
abstract
Multilingual pretrained language models have demonstrated remarkable zero-shot crosslingual transfer capabilities.Such transfer emerges by fine-tuning on a task of interest in one language and evaluating on a distinct language, not seen during the fine-tuning.Despite promising results, we still lack a proper understanding of the source of this transfer.Using a novel layer ablation technique and analyses of the model's internal representations, we show that multilingual BERT, a popular multilingual language model, can be viewed as the stacking of two sub-networks: a multilingual encoder followed by a taskspecific language-agnostic predictor.While the encoder is crucial for cross-lingual transfer and remains mostly unchanged during finetuning, the task predictor has little importance on the transfer and can be reinitialized during fine-tuning.We present extensive experiments with three distinct tasks, seventeen typologically diverse languages and multiple domains to support our hypothesis.RANDOM-INIT of layers SRC-TRG REF ∆1-2 ∆3-4 ∆5-6 ∆7-8 ∆9-10 ∆11-12 Parsing EN -EN 88.98 -0.96 -0.66 -0.93 -0.55 0.04 -0.09RU -RU 85.15 -0.82 -1.38 -1.51 -0.86 -0.29 0.18 AR -AR 59.54 -0.78 -2.14 -1.20 -0.67 -0.27 0.08 EN -X 53.23 -15.77 -6.51 -3.39 -1.47 0.29 1.00 RU -X 55.41 -7.69 -3.71 -3.13 -1.70 0.92 0.94 AR -X 27.97 -4.91 -3.17 -1.48 -1.68 -0.36 -0.14 POS EN -EN 96.51 -0.30 -0.25 -0.40 -0.00 0.05 0.02 RU -RU 96.90 -0.52 -0.55 -0.40 -0.07 0.02 -0.03 AR -AR 79.28 -0.35 -0.49-0.36 -0.19 -0.05 -0.00 EN -X 79.37 -8.94 -2.49-1.66 -0.88 0.20 -0.14 RU -X 79.25 -10.08 -2.83 -1.65 -2.74 0.01 -0.45 AR -X 64.81 -6.73 -3.50 -1.63 -1.56 -0.73 -1.29 NER EN -EN 83.30 -2.66 -2.14 -1.43 -0.63 -0.23 -0.12 RU -RU 88.20 -2.08 -2.13 -1.52 -0.64 -0.33 -0.13 AR -AR 87.97 -2.37 -2.11 -0.96 -0.39 -0.15 0.21 EN -X 64.17 -8.28 -5.09 -3.07 -0.79 -0.47 -0.13 RU -X 62.13 -15.85 -9.36 -5.50 -2.44 -1.16 -0.06 AR -X 65.59 -16.10 -8.42 -3.
Benjamin Muller, Yanai Elazar, Benoît Sagot, Djamé Seddah
EACL1
2021 When Being Unseen from mBERT is just the Beginning: Handling New Languages With Multilingual Language Models
abstract
Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, Djamé Seddah. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, Djamé Seddah
NAACL-HLT1
2020 CamemBERT: a Tasty French Language Model
abstract
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, Benoît Sagot. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Louis Martin, Benjamin Muller, Pedro Ortiz Suarez, Yoann Dupont, Laurent Romary, Éric Villemonte de la Clergerie, Djamé Seddah, Benoît Sagot
ACL2
2020 Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell
abstract
Djamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral, Benjamin Muller, Pedro Javier Ortiz Suárez, Benoît Sagot, Abhishek Srivastava. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Djamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral, Benjamin Muller, Pedro Ortiz Suarez, Benoît Sagot
ACL5
2020 Establishing a New State-of-the-Art for French Named Entity Recognition
abstract
The French TreeBank developed at the University Paris 7 is the main source of morphosyntactic and syntactic annotations for French. However, it does not include explicit information related to named entities, which are among the most useful information for several natural language processing tasks and applications. Moreover, no large-scale French corpus with named entity annotations contain referential information, which complement the type and the span of each mention with an indication of the entity it refers to. We have manually annotated the French TreeBank with such information, after an automatic pre-annotation step. We sketch the underlying annotation guidelines and we provide a few figures about the resulting annotations.
Pedro Ortiz Suarez, Yoann Dupont, Benjamin Muller, Laurent Romary, Benoît Sagot
LREC3