VLDB 2026 Research / reviewers in the wild / expert
Yuan Sun 0010
dblp:75/5247-10
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-0565-9659ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 68% Transfer learning and domain adaptation · 16% Machine translation · 16% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model |
1.0 | 1 | 2026 | Diversity in Unity, Theory in Practice: Hierarchical Multitask Benchmarks for Chinese Minority Languages · ACL (1) 2026 |
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation |
1.0 | 1 | 2026 | Diversity in Unity, Theory in Practice: Hierarchical Multitask Benchmarks for Chinese Minority Languages · ACL (1) 2026 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.9 | 1 | 2025 | Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource Languages · ACL (1) 2025 |
Natural language and speech › Language models and text generation
multilingual language models |
0.9 | 1 | 2025 | Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource Languages · ACL (1) 2025 |
Natural language and speech › Language models and text generation
tokenization |
0.9 | 1 | 2025 | Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource Languages · ACL (1) 2025 |
Natural language and speech › Machine translation
transliteration |
0.9 | 1 | 2025 | Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource Languages · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
benchmark construction · 2.0LLM-as-a-judge · 2.0huffman coding · 0.9character transliteration · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diversity in Unity, Theory in Practice: Hierarchical Multitask Benchmarks for Chinese Minority LanguagesabstractDespite the rapid advancement of LLMs, their performance on linguistically and culturally diverse minority languages within a unified national context remains underexplored.We present CMiLBench, a collection of hierarchical multitask benchmarks designed to translate theoretical notions of diversity in unity (in Chinese: "美美与共") into practical evaluation for three representative Chinese minority languages: Tibetan, Mongolian, and Uyghur.CMiLBench comprises 24,663 instances across 5 difficulty levels and 17 tasks spanning foundational ability, cultural specificity, and safety alignment.We adopt existing dataset adaptation, minority knowledge construction, and high-resource benchmark translation to construct CMiLBench.We assess 14 state-of-the-art commercial and open-source LLMs with a hybrid framework that integrates automatic metrics and LLM-as-a-Judge scoring.The comparative experimental results reveal the gap between theoretical capability and practical utility.CMiLBench serves as a foundational and scalable evaluation resource to bridge the digital language divide and promote the informatization and intelligentization of lowresource Chinese minority languages. Yijie Li 0005, Yuan Sun 0010, Quulgan Minggad, Abdulla Ablikim, Jia Qing Cai Wang |
ACL (1) | 3 |
| 2026 | A fine-grained evaluation framework for language models: Combining pointwise grading and pairwise comparison
Yijie Li 0005, Yuan Sun 0010 |
Inf. Process. Manag. | 2 |
| 2025 | Enhancing Cross-Lingual Transfer through Reversible Transliteration: A Huffman-Based Approach for Low-Resource LanguagesabstractAs large language models (LLMs) are trained on increasingly diverse and extensive multilingual corpora, they demonstrate cross-lingual transfer capabilities. However, these capabilities often fail to effectively extend to low-resource languages, particularly those utilizing non-Latin scripts. While transliterating low-resource languages into Latin script presents a natural solution, there currently lacks a comprehensive framework for integrating transliteration into LLMs training and deployment. Taking a pragmatic approach, this paper innovatively combines character transliteration with Huffman coding to design a complete transliteration framework. Our proposed framework offers the following advantages: 1) Compression: Reduces storage requirements for low-resource language content, achieving up to 50% reduction in file size and 50-80% reduction in token count. 2) Accuracy: Guarantees 100% lossless conversion from transliterated text back to the source language. 3) Efficiency: Eliminates the need for vocabulary expansion for low-resource languages, improving training and inference efficiency. 4) Scalability: The framework can be extended to other low-resource languages. We validate the effectiveness of our framework across multiple downstream tasks, including text classification, machine reading comprehension, and machine translation. Experimental results demonstrate that our method significantly enhances the model's capability to process low-resource languages while maintaining performance on high-resource languages. Our data and code are publicly available at https://github.com/CMLI-NLP/HuffmanTranslit. Wenhao Zhuang, Yuan Sun 0010 |
ACL (1) | 2 |
| 2025 | Tibetan Question Generation Based on Key Sentence and Knowledge GraphabstractQuestion generation aims to generate questions according to the given context and answer, and it has made significant progress in both Chinese and English languages. However, research on Tibetan question generation is still in the early stages, with key challenges including the omission of crucial keywords that render questions unanswerable. Existing large-scale models do not provide robust support for low-resource languages, such as GPT or BERT. To solve the problem, this article proposes to generate Tibetan questions based on key sentences and the knowledge graph. The question generator is based on the Transformer model to better understand context and multiple sources of input information. We identify key sentences to leverage closely related information, and construct a knowledge graph to incorporate more distantly related information. The results show that the BLEU-4 reaches 43.92 on TibetanQA, surpassing existing models in Tibetan question generation and significantly improving the answerability of the generated questions. Yan Zhuang 0007, Yuan Sun 0010, Yijie Li 0005 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | AlpaCream: an Effective Method of Data Selection on AlpacaabstractInstruction Fine-Tuning (IFT) optimizes Large Language Models (LLMs) to enhance the comprehension and execution of user instructions through extensive instruction datasets. However, these datasets are often voluminous, repetitive, and contain a substantial proportion of low-quality data, necessitating effective selection. Traditional data selection methods are plagued by challenges such as insufficient diver-sity, unclear selection criteria, biases from external LLMs, and excessive resource consumption. This paper proposes AlpaCream, a novel instruction data selection methodology designed for industrial applications that aligns with expert insights and ensures maximum diversity of the data. Our approach involves initially categorizing the instruction data into dense clusters using a topic model and then employing a quality assessment model to isolate high-quality instruction data subsets based on their categorizations. The selected data is further improved by using prompt to enhance the data quality. Finally, the augmented data are used to fine-tune the base LLM to get a model with instruction-following solid capability. It can achieve an average performance increase of 31.3% to 39.8% on the Alpaca_52k dataset, which uses 5.76% of the total instruction data. In general, AlpaCream achieved superior performance over comparable instruction data selection methods. Yijie Li 0005, Yuan Sun 0010 |
SMC | 2 |
| 2023 | MiLMo: Minority Multilingual Pre-Trained Language ModelabstractPre-trained language models are trained on large-scale unsupervised data, and they can fine-tune the model only on small-scale labeled datasets, and achieve good results. Multilingual pre-trained language models can be trained on multiple languages, and the model can understand multiple languages at the same time. At present, the search on pre-trained models mainly focuses on rich resources, while there is relatively little research on low-resource languages such as minority languages, and the public multilingual pre-trained language model can not work well for minority languages. Therefore, this paper constructs a multilingual pre-trained model named MiLMo that performs better on minority language tasks, including Mongolian, Tibetan, Uyghur, Kazakh and Korean. To solve the problem of scarcity of datasets on minority languages and verify the effectiveness of the MiLMo model, this paper constructs a minority multilingual text classification dataset named MiTC, and trains a word2vec model for each language. By comparing the word2vec model and the pre-trained model in the text classification task, this paper provides an optimal scheme for the downstream task research of minority languages. The final experimental results show that the performance of the pre-trained model is better than the word2vec model, and it has achieved the best results in minority multilingual text classification. The multilingual pre-trained model MiLMo, multilingual word2vec model and multilingual text classification dataset MiTC are published on https://milmo.cmli-nlp.com/. Junjie Deng, Hanru Shi, Xinhe Yu, Wugedele Bao, Yuan Sun 0010 |
SMC | 5 |
| 2022 | Question Generation Based on Grammar Knowledge and Fine-grained ClassificationabstractQuestion generation is the task of automatically generating questions based on given context and answers, and there are problems that the types of questions and answers do not match. In minority languages such as Tibetan, since the grammar rules are complex and the training data is small, the related research on question generation is still in its infancy. To solve the above problems, this paper constructs a question type classifier and a question generator. We perform fine-grained division of question types and integrate grammatical knowledge into question type classifiers to improve the accuracy of question types. Then, the types predicted by the question type classifier are fed into the question generator. Our model improves the accuracy of interrogative words in generated questions, and the BLEU-4 on SQuAD reaches 17.52, the BLEU-4 on HotpotQA reaches 19.31, the BLEU-4 on TibetanQA reaches 25.58. Yuan Sun 0010, Zhengcuo Dan |
COLING | 1 |
| 2022 | TiBERT: Tibetan Pre-trained Language ModelabstractThe pre-trained language model is trained on large-scale unlabeled text and can achieve state-of-the-art results in many different downstream tasks. However, the current pre-trained language model is mainly concentrated in the Chinese and English fields. For low resource language such as Tibetan, there is lack of a monolingual pre-trained model. To promote the development of Tibetan natural language processing tasks, this paper collects the large-scale training data from Tibetan websites and constructs a vocabulary that can cover 99.95% of the words in the corpus by using Sentencepiece. Then, we train the Tibetan monolingual pre-trained language model named TiBERT on the data and vocabulary. Finally, we apply TiBERT to the downstream tasks of text classification and question generation, and compare it with classic models and multilingual pre-trained models, the experimental results show that TiBERT can achieve the best performance. Our model is published in http://tibert.cmli-nlp.con Junjie Deng, Yuan Sun 0010 |
SMC | 3 |
| 2021 | A Joint Model for Representation Learning of Tibetan Knowledge Graph Based on EncyclopediaabstractLearning the representation of a knowledge graph is critical to the field of natural language processing. There is a lot of research for English knowledge graph representation. However, for the low-resource languages, such as Tibetan, how to represent sparse knowledge graphs is a key problem. In this article, aiming at scarcity of Tibetan knowledge graphs, we extend the Tibetan knowledge graph by using the triples of the high-resource language knowledge graphs and Point of Information map information. To improve the representation learning of the Tibetan knowledge graph, we propose a joint model to merge structure and entity description information based on the Translating Embeddings and Convolution Neural Networks models. In addition, to solve the segmentation errors, we use character and word embedding to learn more complex information in Tibetan. Finally, the experimental results show that our model can make a better representation of the Tibetan knowledge graph than the baseline. Yuan Sun 0010, Andong Chen 0001, Tianci Xia |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2014 | Tibetan-Chinese cross language named entity extraction based on comparable corpus and naturally annotated resourcesabstractTibetan-Chinese named entity extraction can effectively improve the performance of Tibetan-Chinese cross language question answering system, information retrieval, machine translation and other researches. In the condition of no practical Tibetan named entity recognition system and Tibetan-Chinese translation model, this paper proposes a method to extract Tibetan-Chinese entities based on comparable corpus and naturally annotated resources from webs. The main work of this paper is in the following: (1) Tibetan-Chinese comparable corpus construction. (2) Combining sentence length, word matching and boundary term features, using multi-feature fusion algorithm to obtain parallel sentences from comparable corpus. (3) Tibetan-Chinese entity mapping based on the maximum word continuous intersection model of parallel sentence. Finally, the experimental results show that our approach can effectively find Tibetan-Chinese cross language named entity. Yuan Sun 0010 |
CIDM | 1 |