Yihong Liu 0001

dblp:86/3284-1 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0003-1073-0958ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 8 first-author · 14 since 2021
YearPublicationVenuePosition
2026 Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation
abstract
Qianli Wang, Van Bach Nguyen, Yihong Liu, Fedor Splitt, Nils Feldhus, Christin Seifert, Hinrich Schuetze, Sebastian Möller, Vera Schmitt. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Van Bach Nguyen, Yihong Liu 0001, Fedor Splitt, Nils Feldhus, Christin Seifert, Hinrich Schütze, Sebastian Möller 0001, Vera Schmitt
ACL (1)3
2026 Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
abstract
Large language models (LLMs) are increasingly deployed in multilingual, real-world applications with user inputs -naturally introducing typographical errors (typos).Yet most benchmarks assume clean input, leaving the robustness of LLMs to typos across languages largely underexplored.To address this gap, we introduce MULTYPO, a multilingual typo generation algorithm that simulates human-like errors based on language-specific keyboard layouts and typing behavior.We evaluate 18 opensource LLMs across three model families and five downstream tasks spanning language inference, multi-choice question answering, mathematical reasoning, and machine translation tasks.Our results show that typos consistently degrade performance, particularly in generative tasks and those requiring reasoning -while the natural language inference task is comparatively more robust.Instruction tuning improves clean-input performance but may increase brittleness under noise.We also observe language-dependent robustness: high-resource languages are generally more robust than lowresource ones, and translation from English is more robust than translation into English.Our findings underscore the need for noise-aware training and multilingual robustness evaluation.We release a Python package for MULTYPO and make the source code publicly available at https://github.com/cisnlp/multypo. Error Example Sentence NoneColorless green ideas smell furiously.Replacement Colorless green ideaa smell furiously.Insertion Colorless greenm ideas smell furiously.Deletion Coorless green ideas smell furiously.Transposition Colorless green ideas smell furioulsy.
Raoyuan Zhao, Yihong Liu 0001, Lena Altinger, Hinrich Schütze, Michael A. Hedderich
ACL (1)2
2025 LangSAMP: Language-Script Aware Multilingual Pretraining
abstract
Recent multilingual pretrained language models (mPLMs) often avoid using language embeddings -learnable vectors assigned to individual languages.However, this places a significant burden on token representations to encode all language-specific information, which may hinder language neutrality.To address this limitation, we propose Language-Script Aware Multilingual Pretraining (LANGSAMP), a method that incorporates both language and script embeddings to enhance representation learning.Specifically, we integrate these embeddings into the output of the Transformer blocks before passing the final representations to the language modeling head for prediction.We apply LANGSAMP to the continual pretraining of XLM-R (Conneau et al., 2020) on a highly multilingual corpus covering more than 500 languages.The resulting model consistently outperforms the baseline in zero-shot crosslingual transfer across diverse downstream tasks.Extensive analysis reveals that language and script embeddings capture language-and script-specific nuances, which benefits more language-neutral representations, proven by improved pairwise cosine similarity.In our case study, we also show that language and script embeddings can be used to select better source languages for crosslingual transfer.We make our code and models publicly available at https://github. com/cisnlp/LangSAMP.
Yihong Liu 0001, Haotian Ye, Chunlan Ma, Mingyang Wang 0003, Hinrich Schütze
ACL (1)1
2025 Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models
abstract
Mingyang Wang, Heike Adel, Lukas Lange, Yihong Liu, Ercong Nie, Jannik Strötgen, Hinrich Schuetze. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Mingyang Wang 0003, Heike Adel, Lukas Lange, Yihong Liu 0001, Ercong Nie, Jannik Strötgen, Hinrich Schütze
ACL (1)4
2025 Understanding In-Context Machine Translation for Low-Resource Languages: A Case Study on Manchu
abstract
In-context machine translation (MT) with large language models (LLMs) is a promising approach for low-resource MT, as it can readily take advantage of linguistic resources such as grammar books and dictionaries.Such resources are usually selectively integrated into the prompt so that LLMs can directly perform translation without any specific training, via their in-context learning capability (ICL).However, the relative importance of each type of resource, e.g., dictionary, grammar book, and retrieved parallel examples, is not entirely clear.To address this gap, this study systematically investigates how each resource and its quality affect the translation performance, with the Manchu language as our case study. To remove any prior knowledge of Manchu encoded in the LLM parameters and single out the effect of ICL, we also experiment with an enciphered version of Manchu texts.Our results indicate that high-quality dictionaries and good parallel examples are very helpful, while grammars hardly help.In a follow-up study, we showcase a promising application of in-context MT: parallel data augmentation as a way to bootstrap a conventional MT model. When monolingual data abound, generating synthetic parallel data through in-context MT offers a pathway to mitigate data scarcity and build effective and efficient low-resource neural MT systems.
Renhao Pei, Yihong Liu 0001, Peiqin Lin, François Yvon, Hinrich Schütze
ACL (1)2
2025 TransMI: A Framework to Create Strong Baselines from Multilingual Pretrained Language Models for Transliterated Data
abstract
Transliterating related languages that use different scripts into a common script is effective for improving crosslingual transfer in downstream tasks. However, this methodology often makes pretraining a model from scratch unavoidable, as transliteration brings about new subwords not covered in existing multilingual pretrained language models (mPLMs). This is undesirable because it requires a large computation budget. A more promising way is to make full use of available mPLMs. To this end, this paper proposes a simple but effective framework: Transliterate-Merge-Initialize (TransMI). TransMI can create strong baselines for data that is transliterated into a common script by exploiting an existing mPLM and its tokenizer without any training. TransMI has three stages: (a) transliterate the vocabulary of an mPLM into a common script; (b) merge the new vocabulary with the original vocabulary; and (c) initialize the embeddings of the new subwords. We apply TransMI to three strong recent mPLMs. Our experiments demonstrate that TransMI not only preserves the mPLM’s ability to handle non-transliterated data, but also enables it to effectively process transliterated data, thereby facilitating crosslingual transfer across scripts. The results show consistent improvements of 3% to 34% for different mPLMs and tasks. We make our code and models publicly available at https://github.com/cisnlp/TransMI.
Yihong Liu 0001, Chunlan Ma, Haotian Ye, Hinrich Schütze
COLING1
2025 How Transliterations Improve Crosslingual Alignment
abstract
Recent studies have shown that post-aligning multilingual pretrained language models (mPLMs) using alignment objectives on both original and transliterated data can improve crosslingual alignment. This improvement further leads to better crosslingual transfer performance. However, it remains unclear how and why a better crosslingual alignment is achieved, as this technique only involves transliterations, and does not use any parallel data. This paper attempts to explicitly evaluate the crosslingual alignment and identify the key elements in transliteration-based approaches that contribute to better performance. For this, we train multiple models under varying setups for two pairs of related languages: (1) Polish and Ukrainian and (2) Hindi and Urdu. To assess alignment, we define four types of similarities based on sentence representations. Our experimental results show that adding transliterations alone improves the overall similarities, even for random sentence pairs. With the help of auxiliary transliteration-based alignment objectives, especially the contrastive objective, the model learns to distinguish matched from random pairs, leading to better crosslingual alignment. However, we also show that better alignment does not always yield better downstream performance, suggesting that further research is needed to clarify the connection between alignment and performance. The code implementation is based on https://github.com/cisnlp/Transliteration-PPA.
Yihong Liu 0001, Mingyang Wang 0003, Amir Hossein Kargaran, Ayyoob Imani, Orgest Xhelili, Haotian Ye, Chunlan Ma, François Yvon, Hinrich Schütze
COLING1
2025 On Relation-Specific Neurons in Large Language Models
abstract
Yihong Liu, Runsheng Chen, Lea Hirlimann, Ahmad Dawar Hakimi, Mingyang Wang, Amir Hossein Kargaran, Sascha Rothe, François Yvon, Hinrich Schuetze. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yihong Liu 0001, Runsheng Chen, Lea Hirlimann, Ahmad Dawar Hakimi, Mingyang Wang 0003, Amir Hossein Kargaran, Sascha Rothe, François Yvon, Hinrich Schütze
EMNLP1
2025 M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment Analysis
abstract
ChengYan Wu, Bolei Ma, Yihong Liu, Zheyu Zhang, Ningyuan Deng, Yanshu Li, Baolan Chen, Yi Zhang, Yun Xue, Barbara Plank. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
ChengYan Wu, Bolei Ma, Yihong Liu 0001, Zheyu Zhang 0007, Ningyuan Deng, Yanshu Li, Baolan Chen, Barbara Plank
EMNLP3
2025 Refusal Direction is Universal Across Safety-Aligned Languages
abstract
Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass refusals. While this is primarily demonstrated in an English-centric context, appropriate refusal behavior is important for any language, but poorly understood. In this paper, we investigate the refusal behavior in LLMs across 14 languages using \textit{PolyRefuse}, a multilingual safety dataset created by translating malicious and benign English prompts into these languages. We uncover the surprising cross-lingual universality of the refusal direction: a vector extracted from English can bypass refusals in other languages with near-perfect effectiveness, without any additional fine-tuning. Even more remarkably, refusal directions derived from any safety-aligned language transfer seamlessly to others. We attribute this transferability to the parallelism of refusal vectors across languages in the embedding space and identify the underlying mechanism behind cross-lingual jailbreaks. These findings provide actionable insights for building more robust multilingual safety defenses and pave the way for a deeper mechanistic understanding of cross-lingual vulnerabilities in LLMs.
Xinpeng Wang 0003, Mingyang Wang 0003, Yihong Liu 0001, Hinrich Schütze, Barbara Plank
NeurIPS3
2024 TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language Models
abstract
The world's more than 7000 languages are written in at least 293 scripts. 1 Due to various reasons, many closely related languages use different scripts, which poses a difficulty for multilingual pretrained language models (mPLMs) in learning crosslingual knowledge through lexical overlap.As a consequence, mPLMs are faced with a script barrier: representations from different scripts are located in different subspaces, which can result in crosslingual transfer involving languages of different scripts performing suboptimally.To address this problem, we propose TRANSLICO, a framework that optimizes the Transliteration Contrastive Modeling (TCM) objective to fine-tune an mPLM by contrasting sentences in its training data and their transliterations in a unified script (in our case Latin 2 ), which enhances uniformity in the representation space for different scripts.Using Glot500-m (ImaniGooghari et al., 2023), an mPLM pretrained on over 500 languages, as our source model, we fine-tune it on a small portion (5%) of its training data, and refer to the resulting model as FURINA.We show that FURINA not only better aligns representations from distinct scripts but also outperforms the original Glot500-m on various zero-shot crosslingual transfer tasks.Additionally, we achieve consistent improvement in a case study on the Indic group where the languages exhibit areal features but use different scripts.We make our code and models publicly available.3
Yihong Liu 0001, Chunlan Ma, Haotian Ye, Hinrich Schütze
ACL (1)1
2023 A Crosslingual Investigation of Conceptualization in 1335 Languages
abstract
Yihong Liu, Haotian Ye, Leonie Weissweiler, Philipp Wicke, Renhao Pei, Robert Zangenfeind, Hinrich Schütze. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yihong Liu 0001, Haotian Ye, Leonie Weissweiler, Philipp Wicke, Renhao Pei, Robert Zangenfeind, Hinrich Schütze
ACL (1)1
2022 Flow-Adapter Architecture for Unsupervised Machine Translation
abstract
In this work, we propose a flow-adapter architecture for unsupervised NMT.It leverages normalizing flows to explicitly model the distributions of sentence-level latent representations, which are subsequently used in conjunction with the attention mechanism for the translation task.The primary novelties of our model are: (a) capturing language-specific sentence representations separately for each language using normalizing flows and (b) using a simple transformation of these latent representations for translating from one language to another.This architecture allows for unsupervised training of each language independently.While there is prior work on latent variables for supervised MT, to the best of our knowledge, this is the first work that uses latent variables and normalizing flows for unsupervised MT.We obtain competitive results on several unsupervised MT benchmarks.
Yihong Liu 0001, Haris Jabbar, Hinrich Schütze
ACL (1)1
2021 A label-oriented loss function for learning sentence representations
Yihong Liu 0001, Dongxu Lu, Xianchun Zou
Comput. Speech Lang.1