VLDB 2026 Research / reviewers in the wild / expert
Miaoran Zhang
dblp:302/4697
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Machine translation · 94% Information extraction and text analysis · 6% | |
| Network and information security
1 paper |
Privacy and data protection · 100% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
document-level machine translation |
0.9 | 1 | 2025 | AFRIDOC-MT: Document-level MT Corpus for African Languages · EMNLP 2025 |
Natural language and speech › Machine translation
low-resource machine translation |
0.9 | 1 | 2025 | AFRIDOC-MT: Document-level MT Corpus for African Languages · EMNLP 2025 |
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation |
0.8 | 1 | 2024 | Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice? · EMNLP 2024 |
Privacy and data protection › anonymization › identity obfuscation
authorship obfuscation |
0.5 | 1 | 2021 | Preventing Author Profiling through Zero-Shot Multilingual Back-Translation · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › multilingual NLP
multilingual language resources |
0.3 | 1 | 2025 | AFRIDOC-MT: Document-level MT Corpus for African Languages · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
style transfer · 1.0back-translation · 1.0corpus construction · 0.9synthetic data · 0.8fine-tuning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AFRIDOC-MT: Document-level MT Corpus for African LanguagesabstractJesujoba Oluwadara Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, Dawei Zhu, David Ifeoluwa Adelani, Clement Oyeleke Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jesujoba O. Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, David Ifeoluwa Adelani, Clement Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow |
EMNLP | 3 |
| 2024 | Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?abstractTraditionally, success in multilingual machine translation can be attributed to three key factors in training data: large volume, diverse translation directions, and high quality.In the current practice of fine-tuning large language models (LLMs) for translation, we revisit the importance of these factors.We find that LLMs display strong translation capability after being fine-tuned on as few as 32 parallel sentences and that fine-tuning on a single translation direction enables translation in multiple directions.However, the choice of direction is critical: fine-tuning LLMs with only English on the target side can lead to task misinterpretation, which hinders translation into non-English languages.Problems also arise when noisy synthetic data is placed on the target side, especially when the target language is wellrepresented in LLM pre-training.Yet interestingly, synthesized data in an under-represented language has a less pronounced effect.Our findings suggest that when adapting LLMs to translation, the requirement on data quantity can be eased but careful considerations are still crucial to prevent an LLM from exploiting unintended data biases. Pinzhen Chen, Miaoran Zhang, Barry Haddow, Xiaoyu Shen 0001, Dietrich Klakow |
EMNLP | 3 |
| 2022 | MCSE: Multimodal Contrastive Learning of Sentence EmbeddingsabstractMiaoran Zhang, Marius Mosbach, David Adelani, Michael Hedderich, Dietrich Klakow. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Miaoran Zhang, Marius Mosbach, David Ifeoluwa Adelani, Michael A. Hedderich, Dietrich Klakow |
NAACL-HLT | 1 |
| 2021 | Preventing Author Profiling through Zero-Shot Multilingual Back-TranslationabstractDocuments as short as a single sentence may inadvertently reveal sensitive information about their authors, including e.g.their gender or ethnicity.Style transfer is an effective way of transforming texts in order to remove any information that enables author profiling.However, for a number of current state-of-theart approaches the improved privacy is accompanied by an undesirable drop in the downstream utility of the transformed data.In this paper, we propose a simple, zero-shot way to effectively lower the risk of author profiling through multilingual back-translation using off-the-shelf translation models.We compare our models with five representative text style transfer models on three datasets across different domains.Results from both an automatic and a human evaluation show that our approach achieves the best overall performance while requiring no training data.We are able to lower the adversarial prediction of gender and race by up to 22% while retaining 95% of the original utility on downstream tasks. David Ifeoluwa Adelani, Miaoran Zhang, Xiaoyu Shen 0001, Ali Davody, Thomas Kleinbauer, Dietrich Klakow |
EMNLP (1) | 2 |