Miaoran Zhang

dblp:302/4697 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Machine translation · 94% Information extraction and text analysis · 6%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
document-level machine translation
0.912025
AFRIDOC-MT: Document-level MT Corpus for African Languages · EMNLP 2025
Natural language and speech › Machine translation
low-resource machine translation
0.912025
AFRIDOC-MT: Document-level MT Corpus for African Languages · EMNLP 2025
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation
0.812024
Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice? · EMNLP 2024
Privacy and data protection › anonymization › identity obfuscation
authorship obfuscation
0.512021
Preventing Author Profiling through Zero-Shot Multilingual Back-Translation · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis › multilingual NLP
multilingual language resources
0.312025
AFRIDOC-MT: Document-level MT Corpus for African Languages · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

style transfer · 1.0back-translation · 1.0corpus construction · 0.9synthetic data · 0.8fine-tuning · 0.8
YearPublicationVenuePosition
2025 AFRIDOC-MT: Document-level MT Corpus for African Languages
abstract
Jesujoba Oluwadara Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, Dawei Zhu, David Ifeoluwa Adelani, Clement Oyeleke Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jesujoba O. Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, David Ifeoluwa Adelani, Clement Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow
EMNLP3
2024 Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?
abstract
Traditionally, success in multilingual machine translation can be attributed to three key factors in training data: large volume, diverse translation directions, and high quality.In the current practice of fine-tuning large language models (LLMs) for translation, we revisit the importance of these factors.We find that LLMs display strong translation capability after being fine-tuned on as few as 32 parallel sentences and that fine-tuning on a single translation direction enables translation in multiple directions.However, the choice of direction is critical: fine-tuning LLMs with only English on the target side can lead to task misinterpretation, which hinders translation into non-English languages.Problems also arise when noisy synthetic data is placed on the target side, especially when the target language is wellrepresented in LLM pre-training.Yet interestingly, synthesized data in an under-represented language has a less pronounced effect.Our findings suggest that when adapting LLMs to translation, the requirement on data quantity can be eased but careful considerations are still crucial to prevent an LLM from exploiting unintended data biases.
Pinzhen Chen, Miaoran Zhang, Barry Haddow, Xiaoyu Shen 0001, Dietrich Klakow
EMNLP3
2022 MCSE: Multimodal Contrastive Learning of Sentence Embeddings
abstract
Miaoran Zhang, Marius Mosbach, David Adelani, Michael Hedderich, Dietrich Klakow. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Miaoran Zhang, Marius Mosbach, David Ifeoluwa Adelani, Michael A. Hedderich, Dietrich Klakow
NAACL-HLT1
2021 Preventing Author Profiling through Zero-Shot Multilingual Back-Translation
abstract
Documents as short as a single sentence may inadvertently reveal sensitive information about their authors, including e.g.their gender or ethnicity.Style transfer is an effective way of transforming texts in order to remove any information that enables author profiling.However, for a number of current state-of-theart approaches the improved privacy is accompanied by an undesirable drop in the downstream utility of the transformed data.In this paper, we propose a simple, zero-shot way to effectively lower the risk of author profiling through multilingual back-translation using off-the-shelf translation models.We compare our models with five representative text style transfer models on three datasets across different domains.Results from both an automatic and a human evaluation show that our approach achieves the best overall performance while requiring no training data.We are able to lower the adversarial prediction of gender and race by up to 22% while retaining 95% of the original utility on downstream tasks.
David Ifeoluwa Adelani, Miaoran Zhang, Xiaoyu Shen 0001, Ali Davody, Thomas Kleinbauer, Dietrich Klakow
EMNLP (1)2