Kexin Yang 0002

dblp:54/774-2 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0003-1136-6801ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
abstract
Yuzhen Shi, Huanghai Liu, Yiran HU, Song Gaojie, Xu Xinran, Yubo Ma, Tianyi Tang, Li Zhang, Qingjing Chen, Feng Di, Wenbo Lv, Weiheng Wu, Kexin Yang, Sen Yang, Wei Wang, Rongyao Shi, Qiu Yuanyang, Yuemeng Qi, Zhang Jingwen, Sui Xiaoyu, Yifan Chen, Zhang Yi, An Yang, Bowen Yu, Dayiheng Liu, Junyang Lin, Weixing Shen, Bing Zhao, Charles L. A. Clarke, HU Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuzhen Shi, Huanghai Liu, Yiran Hu, Gaojie Song, Xinran Xu, Yubo Ma, Qingjing Chen, Di Feng, Wenbo Lv, Weiheng Wu, Kexin Yang 0002, Wei Wang 0225, Rongyao Shi, Yuanyang Qiu, Yuemeng Qi, Xiaoyu Sui, Yi Zhang 0101, An Yang, Bowen Yu 0002, Dayiheng Liu, Junyang Lin, Weixing Shen, Charles L. A. Clarke, Hu Wei
ACL (1)13
2025 NOVA-63: Native Omni-lingual Versatile Assessments of 63 Disciplines
abstract
The multilingual capabilities of large language models (LLMs) have attracted considerable attention over the past decade. Assessing the accuracy with which LLMs provide answers in multilingual contexts is essential for determining their level of multilingual proficiency. Nevertheless, existing multilingual benchmarks generally reveal severe drawbacks, such as overly translated content (translationese), the absence of difficulty control, constrained diversity, and disciplinary imbalance, making the benchmarking process unreliable and showing low convincingness. To alleviate those shortcomings, we introduce NOVA-63 (Native Omni-lingual Versatile Assessments of 63 Disciplines), a comprehensive, difficult multilingual benchmark featuring 93,536 questions sourced from native speakers across 14 languages and 63 academic disciplines. Leveraging a robust pipeline that integrates LLM-assisted formatting, expert quality verification, and multi-level difficulty screening, NOVA-63 is balanced on disciplines with consistent difficulty standards while maintaining authentic linguistic elements. Extensive experimentation with current LLMs has shown significant insights into cross-lingual consistency among language families, and exposed notable disparities in models’ capabilities across various disciplines. This work provides valuable benchmarking data for the future development of multilingual models. Furthermore, our findings underscore the importance of moving beyond overall scores and instead conducting fine-grained analyses of model performance.
Kexin Yang 0002, Yu Wan 0004, Muyang Ye, Baosong Yang, Junyang Lin, Dayiheng Liu
EMNLP2
2025 DataMan: Data Manager for Pre-training Large Language Models
abstract
The performance emergence of large language models (LLMs) driven by data scaling laws makes the selection of pre-training data increasingly important. However, existing methods rely on limited heuristics and human intuition, lacking comprehensive and clear guidelines. To address this, we are inspired by *``reverse thinking''* -- prompting LLMs to self-identify which criteria benefit its performance. As its pre-training capabilities are related to perplexity (PPL), we derive 14 quality criteria from the causes of text perplexity anomalies and introduce 15 common application domains to support domain mixing. In this paper, we train a **Data** **Man**ager (**DataMan**) to learn quality ratings and domain recognition from pointwise rating, and use it to annotate a 447B token pre-training corpus with 14 quality ratings and domain type. Our experiments validate our approach, using DataMan to select 30B tokens to train a 1.3B-parameter language model, demonstrating significant improvements in in-context learning (ICL), perplexity, and instruction-following ability over the state-of-the-art baseline. The best-performing model, based on the *Overall Score l=5* surpasses a model trained with 50% more data using uniform sampling. We continue pre-training with high-rated, domain-specific data annotated by DataMan to enhance domain-specific ICL performance and thus verify DataMan's domain mixing ability. Our findings emphasize the importance of quality ranking, the complementary nature of quality criteria, and their low correlation with perplexity, analyzing misalignment between PPL and ICL performance. We also thoroughly analyzed our pre-training dataset, examining its composition, the distribution of quality ratings, and the original document sources.
Ru Peng, Kexin Yang 0002, Yawen Zeng, Junyang Lin, Dayiheng Liu, Junbo Zhao 0002
ICLR2
2023 Tailor: A Soft-Prompt-Based Approach to Attribute-Based Controlled Text Generation
abstract
Kexin Yang, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Mingfeng Xue, Boxing Chen, Jun Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Kexin Yang 0002, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Mingfeng Xue, Boxing Chen
ACL (1)1
2023 Fantastic Expressions and Where to Find Them: Chinese Simile Generation with Multiple Constraints
abstract
Kexin Yang, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Xiangpeng Wei, Zhengyuan Liu, Jun Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Kexin Yang 0002, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Xiangpeng Wei, Zhengyuan Liu
ACL (1)1
2023 Fantastic Gradients and Where to Find Them: Improving Multi-attribute Text Style Transfer by Quadratic Program
Qian Qu, Jian Wang 0124, Kexin Yang 0002, Hang Zhang 0029, Jiancheng Lv 0001
NLPCC (3)3
2022 CoupGAN: Chinese couplet generation via encoder-decoder model and adversarial training under global control
Qian Qu, Jiancheng Lv 0001, Dayiheng Liu, Kexin Yang 0002
Soft Comput.4
2021 POS-Constrained Parallel Decoding for Non-autoregressive Generation
abstract
Kexin Yang, Wenqiang Lei, Dayiheng Liu, Weizhen Qi, Jiancheng Lv. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Kexin Yang 0002, Wenqiang Lei, Dayiheng Liu, Weizhen Qi, Jiancheng Lv 0001
ACL/IJCNLP (1)1
2021 AnchiBERT: A Pre-Trained Model for Ancient Chinese Language Understanding and Generation
abstract
Ancient Chinese is the essence of Chinese culture. There are several natural language processing tasks of ancient Chinese domain, such as ancient-modern Chinese translation, poem generation, and couplet generation. Previous studies usually use the supervised models which deeply rely on parallel data. However, it is difficult to obtain large-scale parallel data of ancient Chinese. In order to make full use of the more easily available monolingual ancient Chinese corpora, we release An-chiBERT, a pre-trained language model based on the architecture of BERT, which is trained on large-scale ancient Chinese corpora. We evaluate AnchiBERT on both language understanding and generation tasks, including poem classification, ancient-modern Chinese translation, poem generation, and couplet generation. The experimental results show that AnchiBERT outperforms BERT as well as the non-pretrained models and achieves state-of - the-art results in all cases.
Huishuang Tian, Kexin Yang 0002, Dayiheng Liu, Jiancheng Lv 0001
IJCNN2
2021 An automatic evaluation metric for Ancient-Modern Chinese translation
Kexin Yang 0002, Dayiheng Liu, Qian Qu, Yongsheng Sang, Jiancheng Lv 0001
Neural Comput. Appl.1
2020 Herb-Know: Knowledge Enhanced Prescription Generation for Traditional Chinese Medicine
abstract
Prescription generation of traditional Chinese medicine (TCM) is a meaningful and challenging problem. Previous researches mainly model the relationship between symptoms and herbal prescription directly. However, TCM practitioners often take herb effects into consideration when prescribing. Few works focus on fusing the external knowledge of herbs. In this paper, we explore how to generate a prescription with the knowledge of herb effects under the given symptoms. We propose Herb-Know, a sequence to sequence (seq2seq) model with pointer network, where the prescription is conditioned over two inputs (symptoms and pre-selected herb candidates). To the best of our knowledge, this is the first attempt to generate a prescription with a knowledge enhanced seq2seq model. The experimental results demonstrate that our method can make use of knowledge to generate informative and reasonable herbs, which outperforms other baseline models.
Chanjuan Li, Dayiheng Liu, Kexin Yang 0002, Jiancheng Lv 0001
BIBM3
2020 The COVID-19 outbreak in Sichuan, China: Epidemiology and impact of interventions
abstract
In January 2020, a COVID-19 outbreak was detected in Sichuan Province of China. Six weeks later, the outbreak was successfully contained. The aim of this work is to characterize the epidemiology of the Sichuan outbreak and estimate the impact of interventions in limiting SARS-CoV-2 transmission. We analyzed patient records for all laboratory-confirmed cases reported in the province for the period of January 21 to March 16, 2020. To estimate the basic and daily reproduction numbers, we used a Bayesian framework. In addition, we estimated the number of cases averted by the implemented control strategies. The outbreak resulted in 539 confirmed cases, lasted less than two months, and no further local transmission was detected after February 27. The median age of local cases was 8 years older than that of imported cases. We estimated R0 at 2.4 (95% CI: 1.6-3.7). The epidemic was self-sustained for about 3 weeks before going below the epidemic threshold 3 days after the declaration of a public health emergency by Sichuan authorities. Our findings indicate that, were the control measures be adopted four weeks later, the epidemic could have lasted 49 days longer (95% CI: 31-68 days), causing 9,216 more cases (95% CI: 1,317-25,545).
Quanhui Liu, Ana I. Bento, Kexin Yang 0002, Hang Zhang 0029, Stefano Merler, Alessandro Vespignani, Jiancheng Lv 0001, Tao Zhou 0001, Marco Ajelli
PLoS Comput. Biol.3
2020 Ancient-Modern Chinese Translation with a New Large Training Dataset
abstract
Ancient Chinese brings the wisdom and spirit culture of the Chinese nation. Automatic translation from ancient Chinese to modern Chinese helps to inherit and carry forward the quintessence of the ancients. However, the lack of large-scale parallel corpus limits the study of machine translation in ancient–modern Chinese. In this article, we propose an ancient–modern Chinese clause alignment approach based on the characteristics of these two languages. This method combines both lexical-based information and statistical-based information, which achieves 94.2 F1-score on our manual annotation Test set. We use this method to create a new large-scale ancient–modern Chinese parallel corpus that contains 1.24M bilingual pairs. To our best knowledge, this is the first large high-quality ancient–modern Chinese dataset. Furthermore, we analyzed and compared the performance of the SMT and various NMT models on this dataset and provided a strong baseline for this task.
Dayiheng Liu, Kexin Yang 0002, Qian Qu, Jiancheng Lv 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.2