VLDB 2026 Research / reviewers in the wild / expert
Ming Zhong 0016
dblp:92/2292-16
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0002-7814-7523ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Multi-Modal Retrieval-Augmented Framework for Compiler Backend Generation with LLMs
Ming Zhong 0016, Hongna Geng, Lulin Wang, Lei Qiu 0007, Huimin Cui, Xiaobing Feng 0002 |
SANER | 1 |
| 2026 | BePilot: An AI Programming Assistant for Compiler Backend DevelopmentabstractCompiler backends are tasked with generating executable machine code for various processors. As the diversity of processors continues to grow, it is imperative for programmers to tailor specific compiler backends to accommodate each one. However, compiler backend development remains a labor-intensive and time-consuming process, with limited automation tools available. Although large language models (LLMs) have demonstrated strong abilities in code completion and code generation tasks, the lack of appropriate datasets for compiler backend development limits the application of LLMs in this field. In this article, we introduce ComBack++, a multilingual dataset covering C/C++, machine description, and TableGen, with 184 backends from GCC and LLVM, four backend-specific tasks. Based on ComBack++, we present BePilot, a compiler backend-specific LLM available in two sizes: BePilot-1.5B and BePilot-7B. We also introduce CB-Retriever , a retriever that constructs few-shot prompts via in-context learning to improve vanilla LLM performance in resource-constrained settings. Experimental results show that BePilot-1.5B and BePilot-7B achieve significantly higher accuracy across four tasks in ComBack++ compared to 12 baseline LLMs (125M–34B parameters). In addition, CB-Retriever consistently boosts the accuracy of six mainstream LLMs. Both BePilot-1.5B and BePilot-7B, as well as vanilla LLMs augmented with CB-Retriever , outperform the traditional manual compiler backend development approach (Fork-Flow) in efficiency across all four tasks in ComBack++. Furthermore, human evaluation by four experienced compiler backend developers confirms that BePilot not only improves development efficiency over Fork-Flow but also surpasses commercial AI programming assistants such as GPT-4o-mini and Gemini2-Flash in terms of code quality. These findings confirm that BePilot and CB-Retriever can substantially enhance compiler backend development efficiency. Ming Zhong 0016, Lulin Wang, Hongna Geng, Lei Qiu 0007, Huimin Cui, Xiaobing Feng 0002 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2025 | VEGA: Automatically Generating Compiler Backends using a Pre-trained Transformer ModelabstractWe introduce VEGA, an AI-driven system aimed at easing the development of compiler backends for new targets. Our approach involves categorizing functions from existing backends into function groups, each comprising various target-specific implementations of a standard compiler interface function, abstracted as a single function template. Therefore, generating a new backend involves customizing these function templates to specific target requirements. To capitalize on AI's capabilities in code generation, VEGA maps statements in a target-specific version of a function template into feature vectors, distinguishing between target-independent and target-specific properties. Leveraging a pre-trained model, VEGA can efficiently auto-generate a version of each function template tailored to a specific target, thereby enabling the construction of a complete compiler backend for a new target based solely on its target description files. We evaluated VEGA on three distinct targets: a CPU processor (RISC-V), a customized processor with instruction extensions (RI5CY), and an IoT processor (xCORE). VEGA demonstrated high efficiency, generating compiler backends under an hour, which can substantially enhance developer productivity. Across the three targets, VEGA achieved accuracy rates of 71.5%, 73.2%, and 62.2% for all generated functions, significantly outperforming the traditional fork-flow method, which yielded less than 8% accuracy. Moreover, VEGA provides explicit confidence scores for generated functions and statements, allowing developers to easily identify areas requiring minimal manual intervention. This research has the potential to improve the effectiveness of traditional compiler backend development. Ming Zhong 0016, Lulin Wang, Lei Qiu 0007, Ying Liu 0055, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
CGO | 1 |
| 2025 | ReLOpt: A Retriever-Augmented Framework for Optimizing Code with Long-Range Dependencies
Lei Qiu 0007, Fang Lyu, Ming Zhong 0016, Lulin Wang, Xiaobing Feng 0002 |
ICONIP (1) | 3 |
| 2025 | IR-OptSet: An Optimization-Sensitive Dataset for Advancing LLM-Based IR OptimizerabstractCompiler optimization is essential for improving program performance, yet modern compilers still depend on manually crafted transformation rules over intermediate representations (IRs). As compilers grow in complexity, maintaining these rule-based optimizations becomes increasingly labor-intensive and difficult to scale. Recent advances in large language models (LLMs) offer a promising alternative, but their effectiveness in compiler optimization remains limited—primarily due to the lack of IR-oriented datasets that expose models to diverse transformation samples in real-world scenarios (optimization-sensitive samples), hindering LLMs from learning rich and generalizable optimization strategies.In this paper, we introduce IR-OptSet, the first public optimization-sensitive dataset for advancing LLM-based IR optimizers. It comprises 170K LLVM IR samples from open-source repositories across 8 representative optimization domains. IR-OptSet defines two core tasks: Code Analysis and Optimized Code Generation, and provides tools for correctness verification, performance evaluation, and dataset expansion. In our experiments, fine-tuning three representative LLMs on IR-OptSet leads to significant accuracy improvements across both tasks. Moreover, the LLM fine-tuned with IR-OptSet outperforms traditional compiler with the -O3 option in 64 test cases in terms of performance. Further analysis reveals that IR-OptSet provides greater transformation diversity and representativeness than three widely used IR-oriented datasets, highlighting its potential to drive model-based IR optimization. IR-OptSet is publicly available at https://huggingface.co/datasets/YangziResearch/IR-OptSet. Lei Qiu 0007, Fang Lyu, Ming Zhong 0016, ZhiLei Chai, Haojie Zhou, Huimin Cui, Xiaobing Feng 0002 |
NeurIPS | 4 |
| 2025 | Boosting Large Language Models for System Software Retargeting: A Preliminary StudyabstractSystem software bridges hardware platforms and high-level applications. As new hardware platforms emerge, developers must customize code to support various system software, a process known as “retargeting”. This process is time-consuming and poorly automated. While large language models (LLMs) are proficient in general code generation tasks, their effectiveness in retargeting is limited by code complexity and abstract function descriptions. This paper presents TeSyn, a novel framework to enhance the code generation capabilities for system software retargeting. TeSyn comprises three steps: target-specific value extraction, common code clustering, and template synthesis. To evaluate TeSyn's effectiveness, we intro-duce SysRetar, the first dataset for system software retargeting, covering four types of system software and 195 hardware platforms. In our experiments, we select five LLMs and fine-tune CodeLLaMA-7B-Instruct on SysRetar to create SysRetar-LLM. Results show that TeSyn significantly enhances retargeting performance across five LLMs. Furthermore, code generated by SysRetar- LLM requires substantially less modification than the manual retargeting approach (Fork-Flow), suggesting potential improvements in efficiency. Given these promising results, we outline future research directions for advancing retargeting through LLMs. The dataset and code are publicly available at https://huggingface.co/doczll05/SysRetar-LLM. Ming Zhong 0016, Lulin Wang, Lei Qiu 0007, Hongna Geng, Huimin Cui, Xiaobing Feng 0002 |
SANER | 1 |
| 2024 | ComBack: A Versatile Dataset for Enhancing Compiler Backend Development EfficiencyabstractCompiler backends are tasked with generating executable machine code for processors. With the proliferation of diverse processors, it is imperative for programmers to tailor specific compiler backends to accommodate each one. Meanwhile, compiler backend development is a laborious and time-consuming task, lacking effective automation methods. Although language models have demonstrated strong abilities in code related tasks, the lack of appropriate datasets for compiler backend development limits the application of language models in this field.In this paper, we introduce ComBack, the first public dataset designed for improving compiler backend development capabilities of language models. ComBack includes 178 backends for mainstream compilers and three tasks including statement-level completion, next-statement suggestion and code generation, representing common development scenarios. We conducted experiments by fine-tuning six pre-trained language models with ComBack, demonstrating its effectiveness in enhancing model accuracy across the three tasks. We further evaluated the top-performing model(CodeT5+) across the three tasks for new targets, comparing its accuracy with conventional methods (Fork-Flow), ChatGPT-3.5-Turbo, and Code-LLaMA-34B-Instruct. Remarkably, fine-tuned CodeT5+ with only 220M parameters on ComBack outperformed Fork-Flow methods significantly and surpassed ChatGPT and Code-LLaMA. This suggests potential efficiency improvements in compiler development. ComBack is avaliable at https://huggingface.co/datasets/docz1105/ComBack. Ming Zhong 0016, Fang Lyu, Lulin Wang, Hongna Geng, Lei Qiu 0007, Huimin Cui, Xiaobing Feng 0002 |
NeurIPS | 1 |
| 2023 | OPTango: Multi-central Representation Learning against Innumerable Compiler Optimization for Binary DiffingabstractBinary diffing, which quantitatively measures the difference between given binaries, has been broadly used in critical security areas. Previous studies have been tackling the challenge of default compiler optimization, as it can affect binary representation but overlooked the exploration of non-default optimization settings, which can also significantly affect the accuracy of diffing. Recent research indicates a growing trend of compiling applications with non-default optimization settings to magnify binary code discrepancies, enabling them to evade detection by binary diffing tools. This paper takes the first step to systematically studying the resistance of compiler optimization (including default and non-default optimization settings) on binary diffing tasks. To this end, we construct a diverse and unique dataset, OPTBinary, with 3.6 million functions compiled from 514 optimization settings. Then, we propose OPTango, an innovative transformer-based multi-central representation learning approach, exploring the solution to build a compiler optimization-agnostic binary diffing tool. We conduct extensive experiments and benchmark OPTango with state-of-the-art binary diffing approaches. Evaluation results show that OPTango is more robust and significantly outperforms existing methods against both default and non-default compiler optimization. Hongna Geng, Ming Zhong 0016, Peihua Zhang, Xiaobing Feng 0002 |
ISSRE | 2 |
| 2023 | Automatic Target Description File Generation
Hongna Geng, Fang Lyu, Ming Zhong 0016, Huimin Cui, Jingling Xue, Xiaobing Feng 0002 |
J. Comput. Sci. Technol. | 3 |