EDBT 2026 Demo / reviewers in the wild / expert
Kelly Marchisio
dblp:247/6476
· DBLP profile ↗
9ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 36% Transfer learning and domain adaptation · 17% Representation and self-supervised learning · 15% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
bilingual lexicon induction |
1.1 | 2 | 2022 | IsoVec: Controlling the Relative Isomorphism of Word Embedding Spaces · EMNLP 2022 Bilingual Lexicon Induction for Low-Resource Languages using Graph Matching via Optimal Transport · EMNLP 2022 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.9 | 2 | 2024 | Improving Language Plasticity via Pretraining with Active Forgetting · NeurIPS 2023 RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs · EMNLP 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.9 | 1 | 2025 | Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation · ACL (1) 2025 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs · EMNLP 2024 |
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation |
0.8 | 1 | 2024 | Understanding and Mitigating Language Confusion in LLMs · EMNLP 2024 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.8 | 1 | 2024 | RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs · EMNLP 2024 |
Machine learning › Transfer learning and domain adaptation
language adaptation |
0.7 | 1 | 2023 | Improving Language Plasticity via Pretraining with Active Forgetting · NeurIPS 2023 |
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
0.7 | 1 | 2023 | Improving Language Plasticity via Pretraining with Active Forgetting · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › word representation › word embedding
cross-lingual word embedding |
0.6 | 1 | 2022 | IsoVec: Controlling the Relative Isomorphism of Word Embedding Spaces · EMNLP 2022 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.6 | 1 | 2022 | IsoVec: Controlling the Relative Isomorphism of Word Embedding Spaces · EMNLP 2022 |
Natural language and speech › Language models and text generation
multilingual language models |
0.3 | 1 | 2025 | Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation · ACL (1) 2025 |
Machine learning › Representation and self-supervised learning › representation learning
embedding learning |
0.2 | 1 | 2023 | Improving Language Plasticity via Pretraining with Active Forgetting · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
multilingual benchmarking · 0.9reinforcement learning from human feedback · 0.8preference tuning · 0.8preference optimization · 0.8multilingual supervised fine-tuning · 0.8few-shot prompting · 0.8meta-learning · 0.7embedding reset · 0.7active forgetting · 0.7graph matching · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual EvaluationabstractShivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, Andre Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, André F. T. Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker |
ACL (1) | 8 |
| 2024 | RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMsabstractPreference optimization techniques have become a standard final stage for training state-ofart large language models (LLMs).However, despite widespread adoption, the vast majority of work to-date has focused on a small set of high-resource languages like English and Chinese.This captures a small fraction of the languages in the world, but also makes it unclear which aspects of current state-of-the-art research transfer to a multilingual setting.In this work, we perform an exhaustive study to achieve a new state of the art in aligning multilingual LLMs.We introduce a novel, scalable method for generating high-quality multilingual feedback data to balance data coverage.We establish the benefits of cross-lingual transfer and increased dataset size in preference training.Our preference-trained model achieves a 54.4% win-rate against Aya 23 8B, the current state-of-the-art multilingual LLM in its parameter class, and a 69.5% win-rate or higher against widely used models like Gemma, Mistral and Llama 3. As a result of our efforts, we expand the frontier of alignment techniques to 23 languages, covering approximately half of the world's population. John Dang, Arash Ahmadian, Kelly Marchisio, Julia Kreutzer, Ahmet Üstün, Sara Hooker |
EMNLP | 3 |
| 2024 | Understanding and Mitigating Language Confusion in LLMsabstractWe investigate a surprising limitation of LLMs: their inability to consistently generate text in a user's desired language.We create the Language Confusion Benchmark (LCB) to evaluate such failures, covering 15 typologically diverse languages with existing and newly-created English and multilingual prompts.We evaluate a range of LLMs on monolingual and crosslingual generation reflecting practical use cases, finding that Llama Instruct and Mistral models exhibit high degrees of language confusion and even the strongest models fail to consistently respond in the correct language.We observe that base and English-centric instruct models are more prone to language confusion, which is aggravated by complex prompts and high sampling temperatures.We find that language confusion can be partially mitigated via fewshot prompting, multilingual SFT and preference tuning.We release our language confusion benchmark, which serves as a first layer of efficient, scalable multilingual evaluation. 1 Kelly Marchisio, Wei-Yin Ko, Alexandre Berard, Théo Dehaze, Sebastian Ruder |
EMNLP | 1 |
| 2023 | Improving Language Plasticity via Pretraining with Active ForgettingabstractPretrained language models (PLMs) are today the primary model for natural language processing. Despite their impressive downstream performance, it can be difficult to apply PLMs to new languages, a barrier to making their capabilities universally accessible. While prior work has shown it possible to address this issue by learning a new embedding layer for the new language, doing so is both data and compute inefficient. We propose to use an active forgetting mechanism during pretraining, as a simple way of creating PLMs that can quickly adapt to new languages. Concretely, by resetting the embedding layer every K updates during pretraining, we encourage the PLM to improve its ability of learning new embeddings within limited number of updates, similar to a meta-learning effect. Experiments with RoBERTa show that models pretrained with our forgetting mechanism not only demonstrate faster convergence during language adaptation, but also outperform standard ones in a low-data regime, particularly for languages that are distant from English. Code will be available at https://github.com/facebookresearch/language-model-plasticity. Kelly Marchisio, Roberta Raileanu, David Ifeoluwa Adelani, Pontus Stenetorp, Sebastian Riedel 0001, Mikel Artetxe |
NeurIPS | 2 |
| 2022 | Bilingual Lexicon Induction for Low-Resource Languages using Graph Matching via Optimal TransportabstractBilingual lexicons form a critical component of various natural language processing applications, including unsupervised and semisupervised machine translation and crosslingual information retrieval.We improve bilingual lexicon induction performance across 40 language pairs with a graph-matching method based on optimal transport.The method is especially strong with low amounts of supervision. Kelly Marchisio, Ali Saad-Eldin, Kevin Duh, Carey E. Priebe, Philipp Koehn |
EMNLP | 1 |
| 2022 | IsoVec: Controlling the Relative Isomorphism of Word Embedding SpacesabstractThe ability to extract high-quality translation dictionaries from monolingual word embedding spaces depends critically on the geometric similarity of the spaces-their degree of "isomorphism."We address the root-cause of faulty cross-lingual mapping: that word embedding training resulted in the underlying spaces being non-isomorphic.We incorporate global measures of isomorphism directly into the Skip-gram loss function, successfully increasing the relative isomorphism of trained word embedding spaces and improving their ability to be mapped to a shared crosslingual space.The result is improved bilingual lexicon induction in general data conditions, under domain mismatch, and with training algorithm dissimilarities.We release IsoVec at https://github.com/ kellymarchisio/isovec. Kelly Marchisio, Neha Verma 0001, Kevin Duh, Philipp Koehn |
EMNLP | 1 |
| 2022 | On Systematic Style Differences between Unsupervised and Supervised MT and an Application for High-Resource Machine TranslationabstractModern unsupervised machine translation (MT) systems reach reasonable translation quality under clean and controlled data conditions.As the performance gap between supervised and unsupervised MT narrows, it is interesting to ask whether the different training methods result in systematically different output beyond what is visible via quality metrics like adequacy or BLEU.We compare translations from supervised and unsupervised MT systems of similar quality, finding that unsupervised output is more fluent and more structurally different in comparison to human translation than is supervised MT.We then demonstrate a way to combine the benefits of both methods into a single system which results in improved adequacy and fluency as rated by human evaluators.Our results open the door to interesting discussions about how supervised and unsupervised MT might be different yet mutually-beneficial. Kelly Marchisio, Markus Freitag, David Grangier |
NAACL-HLT | 1 |
| 2021 | An Alignment-Based Approach to Semi-Supervised Bilingual Lexicon Induction with Small Parallel CorporaabstractAimed at generating a seed lexicon for use in downstream natural language tasks and unsupervised methods for bilingual lexicon induction have received much attention in the academic literature recently. While interesting and fully unsupervised settings are unrealistic; small amounts of bilingual data are usually available due to the existence of massively multilingual parallel corpora and or linguists can create small amounts of parallel data. In this work and we demonstrate an effective bootstrapping approach for semi-supervised bilingual lexicon induction that capitalizes upon the complementary strengths of two disparate methods for inducing bilingual lexicons. Whereas statistical methods are highly effective at inducing correct translation pairs for words frequently occurring in a parallel corpus and monolingual embedding spaces have the advantage of having been trained on large amounts of data and and therefore may induce accurate translations for words absent from the small corpus. By combining these relative strengths and our method achieves state-of-the-art results on 3 of 4 language pairs in the challenging VecMap test set using minimal amounts of parallel data and without the need for a translation dictionary. We release our implementation at www.blind-review.code. Kelly Marchisio, Philipp Koehn, Conghao Xiong |
MTSummit (1) | 1 |
| 2019 | Controlling the Reading Level of Machine Translation Output
Kelly Marchisio, Jialiang Guo, Cheng-I Lai, Philipp Koehn |
MTSummit (1) | 1 |