Yijie Li 0005

dblp:54/8054-5 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0003-1735-6804ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
1.012026
Diversity in Unity, Theory in Practice: Hierarchical Multitask Benchmarks for Chinese Minority Languages · ACL (1) 2026
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation
1.012026
Diversity in Unity, Theory in Practice: Hierarchical Multitask Benchmarks for Chinese Minority Languages · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

benchmark construction · 2.0LLM-as-a-judge · 2.0
YearPublicationVenuePosition
2026 Diversity in Unity, Theory in Practice: Hierarchical Multitask Benchmarks for Chinese Minority Languages
abstract
Despite the rapid advancement of LLMs, their performance on linguistically and culturally diverse minority languages within a unified national context remains underexplored.We present CMiLBench, a collection of hierarchical multitask benchmarks designed to translate theoretical notions of diversity in unity (in Chinese: "美美与共") into practical evaluation for three representative Chinese minority languages: Tibetan, Mongolian, and Uyghur.CMiLBench comprises 24,663 instances across 5 difficulty levels and 17 tasks spanning foundational ability, cultural specificity, and safety alignment.We adopt existing dataset adaptation, minority knowledge construction, and high-resource benchmark translation to construct CMiLBench.We assess 14 state-of-the-art commercial and open-source LLMs with a hybrid framework that integrates automatic metrics and LLM-as-a-Judge scoring.The comparative experimental results reveal the gap between theoretical capability and practical utility.CMiLBench serves as a foundational and scalable evaluation resource to bridge the digital language divide and promote the informatization and intelligentization of lowresource Chinese minority languages.
Yijie Li 0005, Yuan Sun 0010, Quulgan Minggad, Abdulla Ablikim, Jia Qing Cai Wang
ACL (1)1
2026 A fine-grained evaluation framework for language models: Combining pointwise grading and pairwise comparison
Yijie Li 0005, Yuan Sun 0010
Inf. Process. Manag.1
2025 Tibetan Question Generation Based on Key Sentence and Knowledge Graph
abstract
Question generation aims to generate questions according to the given context and answer, and it has made significant progress in both Chinese and English languages. However, research on Tibetan question generation is still in the early stages, with key challenges including the omission of crucial keywords that render questions unanswerable. Existing large-scale models do not provide robust support for low-resource languages, such as GPT or BERT. To solve the problem, this article proposes to generate Tibetan questions based on key sentences and the knowledge graph. The question generator is based on the Transformer model to better understand context and multiple sources of input information. We identify key sentences to leverage closely related information, and construct a knowledge graph to incorporate more distantly related information. The results show that the BLEU-4 reaches 43.92 on TibetanQA, surpassing existing models in Tibetan question generation and significantly improving the answerability of the generated questions.
Yan Zhuang 0007, Yuan Sun 0010, Yijie Li 0005
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2024 AlpaCream: an Effective Method of Data Selection on Alpaca
abstract
Instruction Fine-Tuning (IFT) optimizes Large Language Models (LLMs) to enhance the comprehension and execution of user instructions through extensive instruction datasets. However, these datasets are often voluminous, repetitive, and contain a substantial proportion of low-quality data, necessitating effective selection. Traditional data selection methods are plagued by challenges such as insufficient diver-sity, unclear selection criteria, biases from external LLMs, and excessive resource consumption. This paper proposes AlpaCream, a novel instruction data selection methodology designed for industrial applications that aligns with expert insights and ensures maximum diversity of the data. Our approach involves initially categorizing the instruction data into dense clusters using a topic model and then employing a quality assessment model to isolate high-quality instruction data subsets based on their categorizations. The selected data is further improved by using prompt to enhance the data quality. Finally, the augmented data are used to fine-tune the base LLM to get a model with instruction-following solid capability. It can achieve an average performance increase of 31.3% to 39.8% on the Alpaca_52k dataset, which uses 5.76% of the total instruction data. In general, AlpaCream achieved superior performance over comparable instruction data selection methods.
Yijie Li 0005, Yuan Sun 0010
SMC1