Zhongtao Miao

dblp:300/4125 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0002-0785-4714ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning
abstract
Neologism-aware machine translation 1 aims to translate source sentences containing neologisms into target languages.This field remains underexplored compared with general machine translation (MT).In this paper, we propose an agentic framework, NeoAMT, for neologism-aware machine translation equipped with a Wiktionary-based search toolkit.Specifically, we first construct a dedicated dataset for neologism-aware machine translation and build a search toolkit grounded in Wiktionary.The dataset covers 16 languages and 75 translation directions in total, derived from approximately 10 million records of an English Wiktionary dump.The retrieval corpus of the search toolkit is also constructed from around 3 million cleaned records of the same dump.We then leverage the dataset and toolkit to train a translation agent via reinforcement learning (RL) and to evaluate the accuracy of neologismaware machine translation.Furthermore, we propose an RL training framework featuring a novel reward design and an adaptive rollout generation strategy that exploits "translation difficulty" to further improve the translation quality of translation agents using our search toolkit 2 .
Zhongtao Miao, Kaiyan Zhao, Masaaki Nagata, Yoshimasa Tsuruoka
ACL (1)1
2024 Word Alignment as Preference for Machine Translation
abstract
The problem of hallucination and omission, a long-standing problem in machine translation (MT), is more pronounced when a large language model (LLM) is used in MT because an LLM itself is susceptible to these phenomena.In this work, we mitigate the problem in an LLM-based MT model by guiding it to better word alignment.We first study the correlation between word alignment and the phenomena of hallucination and omission in MT.Then we propose to utilize word alignment as preference to optimize the LLM-based MT model.The preference data are constructed by selecting chosen and rejected translations from multiple MT tools.Subsequently, direct preference optimization is used to optimize the LLM-based model towards the preference signal.Given the absence of evaluators specifically designed for hallucination and omission in MT, we further propose selecting hard instances and utilizing GPT-4 to directly evaluate the performance of the models in mitigating these issues.We verify the rationality of these designed evaluation methods by experiments, followed by extensive results demonstrating the effectiveness of word alignment-based preference optimization to mitigate hallucination and omission.On the other hand, although it shows promise in mitigating hallucination and omission, the overall performance of MT in different language directions remains mixed, with slight increases in BLEU and decreases in COMET.
Qiyu Wu 0001, Masaaki Nagata, Zhongtao Miao, Yoshimasa Tsuruoka
EMNLP3
2023 Bugs4Q: A benchmark of existing bugs to enable controlled testing and debugging studies for quantum programs
Pengzhan Zhao, Zhongtao Miao, Shuhan Lan, Jianjun Zhao 0001
J. Syst. Softw.2
2022 A Comprehensive Study of Bug Fixes in Quantum Programs
abstract
As quantum programming evolves, more and more quantum programming languages are being developed. As a result, debugging and testing quantum programs have become increasingly important. While bug fixing in classical programs has come a long way, there is a lack of research in quantum programs. To this end, this paper presents a comprehensive study on bug fixing in quantum programs. We collect and investigate 96 real-world bugs and their fixes from four popular quantum programming languages (Qiskit, Cirq, Q#, and ProjectQ). Our study shows that a high proportion of bugs in quantum programs are quantum-specific bugs (over 80%), which requires further research in the bug fixing domain. We also summarize and extend the bug patterns in quantum programs and subdivide the most critical part, math-related bugs, to make it more applicable to the study of quantum programs. Our findings summarize the characteristics of bugs in quantum programs and provide a basis for studying testing and debugging quantum programs.
Junjie Luo 0005, Pengzhan Zhao, Zhongtao Miao, Shuhan Lan, Jianjun Zhao 0001
SANER3
2021 Bugs4Q: A Benchmark of Real Bugs for Quantum Programs
abstract
Realistic benchmarks of reproducible bugs and fixes are vital to good experimental evaluation of debugging and testing approaches. However, there is no suitable benchmark suite that can systematically evaluate the debugging and testing methods of quantum programs until now. This paper proposes Bugs4Q, a benchmark of thirty-six real, manually validated Qiskit bugs from four popular Qiskit elements (Terra, Aer, Ignis, and Aqua), supplemented with the test cases for reproducing buggy behaviors. Bugs4Q also provides interfaces for accessing the buggy and fixed versions of the Qiskit programs and executing the corresponding test cases, facilitating the reproducible empirical studies and comparisons of Qiskit program debugging and testing tools. Bugs4Q is publicly available at https://github.com/Z-928/Bugs4Q
Pengzhan Zhao, Jianjun Zhao 0001, Zhongtao Miao, Shuhan Lan
ASE3