VLDB 2026 Research / reviewers in the wild / expert
Yun Miao
dblp:166/3160
· DBLP profile ↗
8ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Source Code Summarization in the Era of Large Language ModelsabstractTo support software developers in understanding and maintaining programs, various automatic (source) code summarization techniques have been proposed to generate a concise natural language summary (i.e., comment) for a given code snippet. Recently, the emergence of large language models (LLMs) has led to a great boost in the performance of coderelated tasks. In this paper, we undertake a systematic and comprehensive study on code summarization in the era of LLMs, which covers multiple aspects involved in the workflow of LLMbased code summarization. Specifically, we begin by examining prevalent automated evaluation methods for assessing the quality of summaries generated by LLMs and find that the results of the GPT-4 evaluation method are most closely aligned with human evaluation. Then, we explore the effectiveness of five prompting techniques (zero-shot, few-shot, chain-of-thought, critique, and expert) in adapting LLMs to code summarization tasks. Contrary to expectations, advanced prompting techniques may not outperform simple zero-shot prompting. Next, we investigate the impact of LLMs' model settings (including top_p and temperature parameters) on the quality of generated summaries. We find the impact of the two parameters on summary quality varies by the base LLM and programming language, but their impacts are similar. Moreover, we canvass LLMs' abilities to summarize code snippets in distinct types of programming languages. The results reveal that LLMs perform suboptimally when summarizing code written in logic programming languages compared to other language types (e.g., procedural and object-oriented programming languages). Finally, we unexpectedly find that CodeLlamaInstruct with 7B parameters can outperform advanced GPT-4 in generating summaries describing code design rationale and asserting code properties. We hope that our findings can provide a comprehensive understanding of code summarization in the era of LLMs. Weisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang 0002, Chunrong Fang, Yi Liu 0069, Gelei Deng, Yang Liu 0003, Zhenyu Chen 0001 |
ICSE | 2 |
| 2024 | On the Effectiveness of Large Language Models in Statement-level Code SummarizationabstractCode comments are crucial for program comprehension, and the automated generation of comments greatly enhances the efficiency of code commenting. Statement-level code summarization represents the finest granularity of code summarization, typically encompassing explanations of the purpose and functionality of code statements. Large Language Models (LLMs) are deep learning models trained on massive text data. They possess not only the capability to generate natural language text but also to deeply understand the meaning of text, applicable to various natural language processing tasks such as text summarization, question answering, and translation. Currently, LLMs have demonstrated the ability to generate summaries for code. In this paper, we systematically investigate the capability of LLMs to generate statement-level code summaries. we construct a dataset for statement-level code summarization and evaluate the ability of large language models on statement-level code summarization. For further research, we investigate the impact of different prompting techniques. We also study the effect of the temperature parameter on the quality of generated summaries. Additionally, we compare large language models with the state-of-the-art pretrained model CodeT5 and find out that large language models have great potential to replace pre-trained models and become the new state-of-the-art models on statement-level code summarization. To ensure the reliability of automatic evaluation, we also conduct human evaluation. Our findings emphasize the transformative potential of LLMs on statement-level code summarization and the challenges yet to be overcome. Yun Miao, Junwu Zhu, Xiaolei Sun |
QRS | 2 |
| 2023 | Perturbations of Tensor-Schur decomposition and its applications to multilinear control systems and facial recognitions
Juefei Chen, Wanli Ma 0004, Yun Miao, Yimin Wei 0001 |
Neurocomputing | 3 |
| 2023 | TitleGen-FL: Quality prediction-based filter for automated issue title generation
Xiang Chen 0005, Xuejiao Chen, Zhanqi Cui, Yun Miao, Jianmin Wang 0015 |
J. Syst. Softw. | 5 |
| 2023 | T-ADAF: Adaptive Data Augmentation Framework for Image Classification Network Based on Tensor T-product Operator
Feiyang Han, Yun Miao, Zhaoyi Sun, Yimin Wei 0001 |
Neural Process. Lett. | 2 |
| 2018 | Seeking the user interface
Steven P. Reiss, Yun Miao, Qi Xin 0001 |
Autom. Softw. Eng. | 2 |
| 2017 | A silicon based fdNIRS system with integrated tDCS on chip for non-invasive closed-loop neuro stimulationabstractA silicon based frequency domain near infrared spectroscopy (fdNIRS) with integrated transcranial direct-current stimulation (tDCS) on chip is designed to enable non-invasive closed-loop brain stimulation for neuro diseases treatment and cognitive performance enhancement. The dual channel fdNIRS system amplifies, filters, and downconverts the received photocurrent in a heterodyning scheme with a dynamic range of 10nA to 450nA and a detectable phase shift down to 0.2°. Each channel provides a transimpedance gain of 120dBΩ at 80MHz with a power consumption of 30mW while tolerating up to 8pF input capacitance. The on chip programmable stimulator can support a stimulating current from 0.6mA to 2.2mA with 1% accuracy, which covers the required current range of tDCS. A time-multiplexer is also embedded to support multi-distance fdNIRS measurements. The chip is fabricated in a standard 130nm CMOS process and occupies an area of 2.25mm2. Yun Miao, Valencia Joyner Koomson |
ISCAS | 1 |
| 2015 | A 350μW Sign-Bit architecture for multi-parameter estimation during OFDM acquisition in 65nm CMOSabstractCorrect estimation of symbol timing, Carrier Frequency Offset (CFO), and Signal-to-Noise Ratio (SNR) is crucial in Orthogonal Frequency Division Multiplexing (OFDM) communication. Typically, high estimation accuracy is desired, but often comes with increased complexity. Which has a direct repercussion in energy consumption. In this article, an architecture based on Sign-Bit estimation with low complexity, and hence low power dissipation, is presented. The architecture, is capable of estimating the afore-mentioned parameters in virtually any OFDM standard. The proof of concept has been fabricated in 65nm CMOS technology with low-power high-VT cells. Measurements performed with supply voltage of 1.2V. resulted in a power dissipation of 350 μW, 6 times smaller to that of an equivalent 8-bit architecture, and the lowest power density reported in literature. Isael Diaz, Siyu Tan, Yun Miao, Leif R. Wilhelmsson, Ove Edfors, Viktor Öwall |
ISCAS | 3 |