VLDB 2026 Research / reviewers in the wild / expert
Zhongtao Miao
dblp:300/4125
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0002-0785-4714ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement LearningabstractNeologism-aware machine translation 1 aims to translate source sentences containing neologisms into target languages.This field remains underexplored compared with general machine translation (MT).In this paper, we propose an agentic framework, NeoAMT, for neologism-aware machine translation equipped with a Wiktionary-based search toolkit.Specifically, we first construct a dedicated dataset for neologism-aware machine translation and build a search toolkit grounded in Wiktionary.The dataset covers 16 languages and 75 translation directions in total, derived from approximately 10 million records of an English Wiktionary dump.The retrieval corpus of the search toolkit is also constructed from around 3 million cleaned records of the same dump.We then leverage the dataset and toolkit to train a translation agent via reinforcement learning (RL) and to evaluate the accuracy of neologismaware machine translation.Furthermore, we propose an RL training framework featuring a novel reward design and an adaptive rollout generation strategy that exploits "translation difficulty" to further improve the translation quality of translation agents using our search toolkit 2 . Zhongtao Miao, Kaiyan Zhao, Masaaki Nagata, Yoshimasa Tsuruoka |
ACL (1) | 1 |
| 2024 | Word Alignment as Preference for Machine TranslationabstractThe problem of hallucination and omission, a long-standing problem in machine translation (MT), is more pronounced when a large language model (LLM) is used in MT because an LLM itself is susceptible to these phenomena.In this work, we mitigate the problem in an LLM-based MT model by guiding it to better word alignment.We first study the correlation between word alignment and the phenomena of hallucination and omission in MT.Then we propose to utilize word alignment as preference to optimize the LLM-based MT model.The preference data are constructed by selecting chosen and rejected translations from multiple MT tools.Subsequently, direct preference optimization is used to optimize the LLM-based model towards the preference signal.Given the absence of evaluators specifically designed for hallucination and omission in MT, we further propose selecting hard instances and utilizing GPT-4 to directly evaluate the performance of the models in mitigating these issues.We verify the rationality of these designed evaluation methods by experiments, followed by extensive results demonstrating the effectiveness of word alignment-based preference optimization to mitigate hallucination and omission.On the other hand, although it shows promise in mitigating hallucination and omission, the overall performance of MT in different language directions remains mixed, with slight increases in BLEU and decreases in COMET. Qiyu Wu 0001, Masaaki Nagata, Zhongtao Miao, Yoshimasa Tsuruoka |
EMNLP | 3 |
| 2023 | Bugs4Q: A benchmark of existing bugs to enable controlled testing and debugging studies for quantum programs
Pengzhan Zhao, Zhongtao Miao, Shuhan Lan, Jianjun Zhao 0001 |
J. Syst. Softw. | 2 |
| 2022 | A Comprehensive Study of Bug Fixes in Quantum ProgramsabstractAs quantum programming evolves, more and more quantum programming languages are being developed. As a result, debugging and testing quantum programs have become increasingly important. While bug fixing in classical programs has come a long way, there is a lack of research in quantum programs. To this end, this paper presents a comprehensive study on bug fixing in quantum programs. We collect and investigate 96 real-world bugs and their fixes from four popular quantum programming languages (Qiskit, Cirq, Q#, and ProjectQ). Our study shows that a high proportion of bugs in quantum programs are quantum-specific bugs (over 80%), which requires further research in the bug fixing domain. We also summarize and extend the bug patterns in quantum programs and subdivide the most critical part, math-related bugs, to make it more applicable to the study of quantum programs. Our findings summarize the characteristics of bugs in quantum programs and provide a basis for studying testing and debugging quantum programs. Junjie Luo 0005, Pengzhan Zhao, Zhongtao Miao, Shuhan Lan, Jianjun Zhao 0001 |
SANER | 3 |
| 2021 | Bugs4Q: A Benchmark of Real Bugs for Quantum ProgramsabstractRealistic benchmarks of reproducible bugs and fixes are vital to good experimental evaluation of debugging and testing approaches. However, there is no suitable benchmark suite that can systematically evaluate the debugging and testing methods of quantum programs until now. This paper proposes Bugs4Q, a benchmark of thirty-six real, manually validated Qiskit bugs from four popular Qiskit elements (Terra, Aer, Ignis, and Aqua), supplemented with the test cases for reproducing buggy behaviors. Bugs4Q also provides interfaces for accessing the buggy and fixed versions of the Qiskit programs and executing the corresponding test cases, facilitating the reproducible empirical studies and comparisons of Qiskit program debugging and testing tools. Bugs4Q is publicly available at https://github.com/Z-928/Bugs4Q Pengzhan Zhao, Jianjun Zhao 0001, Zhongtao Miao, Shuhan Lan |
ASE | 3 |