Yingfa Chen

dblp:294/3236 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 36% Efficient and distributed learning · 34% Deep learning architectures and training · 11%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.722025
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity · ICML 2025
Cost-Optimal Grouped-Query Attention for Long-Context Modeling · EMNLP 2025
Machine learning › Efficient and distributed learning › model compression › sparsity
activation sparsity
0.912025
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity · ICML 2025
Machine learning › Deep learning architectures and training
feedforward neural network
0.912025
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity · ICML 2025
Natural language and speech › Language models and text generation
large language model
0.912025
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity · ICML 2025
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
long-context evaluation
0.812024
ınftyBench: Extending Long Context Evaluation Beyond 100K Tokens · ACL (1) 2024
Natural language and speech › Question answering and dialogue systems › human-computer dialogue
real-time conversation
0.812024
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models · EMNLP 2024
Natural language and speech › Language models and text generation › evaluation of language models
benchmark construction
0.712023
READIN: A Chinese Multi-Task Benchmark with Realistic and Diverse Input Noises · ACL (1) 2023
Natural language and speech › Speech recognition and synthesis › noise robustness
robustness to noisy input
0.712023
READIN: A Chinese Multi-Task Benchmark with Realistic and Diverse Input Noises · ACL (1) 2023
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling
0.312025
Cost-Optimal Grouped-Query Attention for Long-Context Modeling · EMNLP 2025
Natural language and speech › Language models and text generation
chinese language processing
0.212023
READIN: A Chinese Multi-Task Benchmark with Realistic and Diverse Input Noises · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

scaling law analysis · 0.9cost optimization · 0.9ReLU activation · 0.9FLOP analysis · 0.9benchmark construction · 0.8robust training · 0.7data augmentation · 0.7
YearPublicationVenuePosition
2025 Multi-Modal Multi-Granularity Tokenizer for Chu Bamboo Slips
abstract
This study presents a multi-modal multi-granularity tokenizer specifically designed for analyzing ancient Chinese scripts, focusing on the Chu bamboo slip (CBS) script used during the Spring and Autumn and Warring States period (771-256 BCE) in Ancient China. Considering the complex hierarchical structure of ancient Chinese scripts, where a single character may be a combination of multiple sub-characters, our tokenizer first adopts character detection to locate character boundaries. Then it conducts character recognition at both the character and sub-character levels. Moreover, to support the academic community, we assembled the first large-scale dataset of CBSs with over 100K annotated character image scans. On the part-of-speech tagging task built on our dataset, using our tokenizer gives a 5.5% relative improvement in F1-score compared to mainstream sub-word tokenizers. Our work not only aids in further investigations of the specific script but also has the potential to advance research on other forms of ancient Chinese scripts.
Yingfa Chen, Chenlong Hu, Shi Yu 0001, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
COLING1
2025 Cost-Optimal Grouped-Query Attention for Long-Context Modeling
abstract
Grouped-Query Attention (GQA) is a widely adopted strategy for reducing the computational cost of attention layers in large language models (LLMs).However, current GQA configurations are often suboptimal because they overlook how context length influences inference cost.Since inference cost grows with context length, the most cost-efficient GQA configuration should vary accordingly.In this work, we analyze the relationship among context length, model size, GQA configuration, and model loss, and introduce two innovations:(1) we decouple the total head size from the hidden size, enabling more flexible control over attention FLOPs; and (2) we jointly optimize the model size and the GQA configuration to arrive at a better allocation of inference resources between attention layers and other components.Our analysis reveals that commonly used GQA configurations are highly suboptimal for longcontext scenarios.Moreover, we propose a recipe for deriving cost-optimal GQA configurations.Our results show that for long-context scenarios, one should use fewer attention heads while scaling up the model size.Configurations selected by our recipe can reduce both memory usage and FLOPs by more than 50% compared to Llama-3's GQA, with no degradation in model capabilities.Our findings offer valuable insights for designing efficient longcontext LLMs. 1 Memory (GB)-57.8% -50.8%
Yingfa Chen, Zhen Leng Thai, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
EMNLP1
2025 Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
abstract
Activation sparsity denotes the existence of substantial weakly-contributed neurons within feed-forward networks of large language models (LLMs), providing wide potential benefits such as computation acceleration. However, existing works lack thorough quantitative studies on this useful property, in terms of both its measurement and influential factors. In this paper, we address three underexplored research questions: (1) How can activation sparsity be measured more accurately? (2) How is activation sparsity affected by the model architecture and training process? (3) How can we build a more sparsely activated and efficient LLM? Specifically, we develop a generalizable and performance-friendly metric, named CETT-PPL-1%, to measure activation sparsity. Based on CETT-PPL-1%, we quantitatively study the influence of various factors and observe several important phenomena, such as the convergent power-law relationship between sparsity and training data amount, the higher competence of ReLU activation than mainstream SiLU activation, the potential sparsity merit of a small width-depth ratio, and the scale insensitivity of activation sparsity. Finally, we provide implications for building sparse and effective LLMs, and demonstrate the reliability of our findings by training a 2.4B model with a sparsity ratio of 93.52%, showing 4.1$\times$ speedup compared with its dense version. The codes and checkpoints are available at https://github.com/thunlp/SparsingLaw/.
Yuqi Luo, Xu Han 0007, Yingfa Chen, Chaojun Xiao, Xiaojun Meng, Liqun Deng, Jiansheng Wei, Zhiyuan Liu 0001, Maosong Sun 0001
ICML4
2024 ınftyBench: Extending Long Context Evaluation Beyond 100K Tokens
abstract
Xinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Hao, Xu Han, Zhen Thai, Shuo Wang, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yingfa Chen, Shengding Hu, Zihang Xu, Moo Khai Hao, Xu Han 0007, Zhen Leng Thai, Shuo Wang 0013, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)2
2024 Robust and Scalable Model Editing for Large Language Models
abstract
Large language models (LLMs) can make predictions using parametric knowledge – knowledge encoded in the model weights – or contextual knowledge – knowledge presented in the context. In many scenarios, a desirable behavior is that LLMs give precedence to contextual knowledge when it conflicts with the parametric knowledge, and fall back to using their parametric knowledge when the context is irrelevant. This enables updating and correcting the model’s knowledge by in-context editing instead of retraining. Previous works have shown that LLMs are inclined to ignore contextual knowledge and fail to reliably fall back to parametric knowledge when presented with irrelevant context. In this work, we discover that, with proper prompting methods, instruction-finetuned LLMs can be highly controllable by contextual knowledge and robust to irrelevant context. Utilizing this feature, we propose EREN (Edit models by REading Notes) to improve the scalability and robustness of LLM editing. To better evaluate the robustness of model editors, we collect a new dataset, that contains irrelevant questions that are more challenging than the ones in existing datasets. Empirical results show that our method outperforms current state-of-the-art methods by a large margin. Unlike existing techniques, it can integrate knowledge from multiple edits, and correctly respond to syntactically similar but semantically unrelated inputs (and vice versa). The source code can be found at https://github.com/thunlp/EREN.
Yingfa Chen, Zhengyan Zhang, Xu Han 0007, Chaojun Xiao, Zhiyuan Liu 0001, Kuai Li, Maosong Sun 0001
LREC/COLING1
2024 Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
abstract
Xinrong Zhang, Yingfa Chen, Shengding Hu, Xu Han, Zihang Xu, Yuanwei Xu, Weilin Zhao, Maosong Sun, Zhiyuan Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yingfa Chen, Shengding Hu, Xu Han 0007, Zihang Xu, Yuanwei Xu, Weilin Zhao, Maosong Sun 0001, Zhiyuan Liu 0001
EMNLP2
2023 READIN: A Chinese Multi-Task Benchmark with Realistic and Diverse Input Noises
abstract
For many real-world applications, the usergenerated inputs usually contain various noises due to speech recognition errors caused by linguistic variations 1 or typographical errors (typos).Thus, it is crucial to test model performance on data with realistic input noises to ensure robustness and fairness.However, little study has been done to construct such benchmarks for Chinese, where various languagespecific input noises happen in the real world.In order to fill this important gap, we construct READIN: a Chinese multi-task benchmark with REalistic And Diverse Input Noises.READIN contains four diverse tasks and requests annotators to re-enter the original test data with two commonly used Chinese input methods: Pinyin input and speech input.We designed our annotation pipeline to maximize diversity, for example by instructing the annotators to use diverse input method editors (IMEs) for keyboard noises and recruiting speakers from diverse dialectical groups for speech noises.We experiment with a series of strong pretrained language models as well as robust training methods, we find that these models often suffer significant performance drops on READIN even with robustness methods like data augmentation.As the first large-scale attempt in creating a benchmark with noises geared towards user-generated inputs, we believe that READIN serves as an important complement to existing Chinese NLP benchmarks.The source code and dataset can be obtained from https://github.com/ thunlp/READIN.
Chenglei Si, Zhengyan Zhang, Yingfa Chen, Xiaozhi Wang, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)3
2023 Sub-Character Tokenization for Chinese Pretrained Language Models
abstract
Abstract Tokenization is fundamental to pretrained language models (PLMs). Existing tokenization methods for Chinese PLMs typically treat each character as an indivisible token. However, they ignore the unique feature of the Chinese writing system where additional linguistic information exists below the character level, i.e., at the sub-character level. To utilize such information, we propose sub-character (SubChar for short) tokenization. Specifically, we first encode the input text by converting each Chinese character into a short sequence based on its glyph or pronunciation, and then construct the vocabulary based on the encoded text with sub-word segmentation. Experimental results show that SubChar tokenizers have two main advantages over existing tokenizers: 1) They can tokenize inputs into much shorter sequences, thus improving the computational efficiency. 2) Pronunciation-based SubChar tokenizers can encode Chinese homophones into the same transliteration sequences and produce the same tokenization output, hence being robust to homophone typos. At the same time, models trained with SubChar tokenizers perform competitively on downstream tasks. We release our code and models at https://github.com/thunlp/SubCharTokenization to facilitate future work.
Chenglei Si, Zhengyan Zhang, Yingfa Chen, Fanchao Qi, Xiaozhi Wang, Zhiyuan Liu 0001, Yasheng Wang, Qun Liu 0001, Maosong Sun 0001
Trans. Assoc. Comput. Linguistics3