VLDB 2026 Research / reviewers in the wild / expert
Chunguang Zhao
dblp:405/7506
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 43% Machine translation · 14% Trustworthy machine learning · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
instruction tuning |
2.0 | 2 | 2026 | M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning Datasets · SIGIR 2026 MIDB: Multilingual Instruction Data Booster for Enhancing Cultural Equality in Multilingual Instruction Synthesis · AAAI 2026 |
Machine learning › Efficient and distributed learning
data selection |
1.0 | 1 | 2026 | M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning Datasets · SIGIR 2026 |
Machine learning › Deep learning architectures and training
diversity-aware sampling |
1.0 | 1 | 2026 | M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning Datasets · SIGIR 2026 |
Machine learning › Trustworthy machine learning
fairness |
1.0 | 1 | 2026 | MIDB: Multilingual Instruction Data Booster for Enhancing Cultural Equality in Multilingual Instruction Synthesis · AAAI 2026 |
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation |
1.0 | 1 | 2026 | The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models · ACL (1) 2026 |
Computational social science and digital humanities › cultural analysis
cross-cultural analysis |
0.3 | 1 | 2026 | The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
benchmark construction · 2.0quality scoring model · 1.0maximal marginal relevance · 1.0instruction synthesis · 1.0fine-tuning · 1.0data revision · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIDB: Multilingual Instruction Data Booster for Enhancing Cultural Equality in Multilingual Instruction SynthesisabstractDespite doubts on data quality, instruction synthesis has been widely applied into instruction tuning (IT) of LLMs as an economic and rapid alternative. Recent endeavors focus on improving data quality for synthesized instruction pairs in English and have facilitated IT of English-centric LLMs. However, data quality issues in multilingual synthesized instruction pairs are even more severe, since the common synthesizing practice is to translate English synthesized data into other languages using machine translation (MT). Besides the known content errors in these English synthesized data, multilingual synthesized instruction data are further exposed to defects introduced by MT and face insufficient localization of the target languages, leading to cultural inequality in trained LLMs. In this paper, we propose MIDB, a Multilingual Instruction Data Booster to automatically address the quality issues in multilingual synthesized data. MIDB is trained on around 36.8k revision examples across 16 languages by human linguistic experts, thereby can boost the low-quality data by addressing content errors and MT defects, and improving localization in these synthesized data. Both automatic and human evaluation indicate that not only MIDB steadily improved instruction data quality in 16 languages, but also the instruction-following and cultural-understanding abilities of multilingual LLMs fine-tuned on MIDB-boosted data were significantly enhanced, suggesting an improved linguistic and cultural equality. Yilun Liu 0001, Chunguang Zhao, Xinhua Yang, Hongyong Zeng, Shimin Tao, Weibin Meng, Minggui He, Hongxia Ma, Daimeng Wei, Boxing Chen |
AAAI | 2 |
| 2026 | The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language ModelsabstractYilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui HE, Chenxin Liu, Zhang Li, Mahongxia, Jiaxin Guo, Chen Liu, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yilun Liu 0001, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui He, Chenxin Liu, Hongxia Ma, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao |
ACL (1) | 2 |
| 2026 | M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning DatasetsabstractMultilingual instruction fine-tuning (IFT) empowers large language models to generalize across diverse linguistic and cultural contexts; however, high-quality, systematically curated multilingual IFT datasets remain scarce. To address this gap, we propose M-DaQ (Multilingual Diversity and Quality), a diversity-aware sampling framework that jointly optimizes instruction-response quality and cross-lingual semantic diversity. M-DaQ leverages a fine-tuned Quality Scoring Model alongside a maximal marginal relevance-inspired selection strategy to construct balanced, high-fidelity training data. Furthermore, we present the first systematic investigation of the Superficial Alignment Hypothesis in multilingual settings. Extensive evaluations across 18 languages demonstrate that models trained on M-DaQ-curated data achieve average win rates exceeding 60% against strong baselines on Alpaca-Eval and MT-Bench. Complementary human evaluations corroborate these gains, highlighting significant improvements in cultural relevance, contextual appropriateness, and instruction-following capability. The code are publicly released to facilitate reproducibility and future research. Chunguang Zhao, Yilun Liu 0001, Pufan Zeng, Yuanchang Luo, Shimin Tao, Minggui He, Weibin Meng, Hongxia Ma, Boxing Chen, Daimeng Wei |
SIGIR | 1 |
| 2025 | SuperFC: Selective Data Utilization for a Sustainable and Effective Function-Calling AgentabstractThe function-calling agent is obtained by performing agent tuning to the large language model (LLM) on function-calling dataset. However, even state-of-the-art datasets (e.g., xlam-function-calling-60k datasets) still contain numerous misleading examples of low-quality data, wasting significant computational resources and result in an unnecessary carbon footprint. Furthermore, such inductive bad data negatively impacts the performance of the agent. In this paper, we propose a set of scoring criteria specifically tailored to evaluate function-calling data and use these criteria to develop a data filtering framework. By applying this framework to filter out low-quality data, we fine-tuned SuperFC, which demonstrates substantial improvements in both sustainability and performance. The SuperFC-7B training process reduced training time from 455 minutes to 85 minutes, resulting in a 80.02% reduction in carbon footprint. Simultaneously, fine-tuning on high-quality data subsets led to performance improvements of up to 3.68%. Additionally, we provide an in-depth analysis of the causes behind the low quality of synthetic function-calling data, offering valuable insights for future data synthesis in this domain. We have also released a high-quality function-calling dataset, available at: https://github.com/Zire-Young/SuperFC Xinhua Yang, Yilun Liu 0001, Shimin Tao, Chunguang Zhao, Weibin Meng, Minggui He, Chang Su 0001, Hongxia Ma, Jingzhou Du, Hao Yang 0006, Boxing Chen, Chuanwen Li |
IJCNN | 4 |
| 2025 | An integrated deep learning method and optimization algorithm framework for energy-saving steel billet heating
Wenchao Ji, Chunguang Zhao, Zhi Yi, Linyang Wei, Shuangcheng Sun |
Eng. Appl. Artif. Intell. | 3 |