Yao Chen 0009

dblp:70/3621-9 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0003-4824-8531ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 81% Efficient and distributed learning · 19%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
AttnPO: Attention-Guided Process Supervision for Efficient Reasoning · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model reasoning
efficient reasoning
1.012026
AttnPO: Attention-Guided Process Supervision for Efficient Reasoning · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model reasoning
process supervision
1.012026
AttnPO: Attention-Guided Process Supervision for Efficient Reasoning · ACL (1) 2026
Natural language and speech › Language models and text generation › chain-of-thought reasoning
chain-of-thought distillation
0.912025
Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key Information · EMNLP 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key Information · EMNLP 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.312025
Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key Information · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

process supervision · 1.0attention guidance · 1.0mixture of layers · 0.9chain-of-thought distillation · 0.9attention alignment · 0.9
YearPublicationVenuePosition
2026 AttnPO: Attention-Guided Process Supervision for Efficient Reasoning
abstract
Shuaiyi Nie, Dingsiyu, Wenyuan Zhang, Linhao Yu, Tianmeng Yang, Yao Chen, Weichong Yin, Yu Sun, Hua Wu, Tingwen Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shuaiyi Nie, Siyu Ding, Wenyuan Zhang 0002, Linhao Yu, Tianmeng Yang, Yao Chen 0009, Weichong Yin, Hua Wu 0003, Tingwen Liu
ACL (1)6
2025 Improving Reasoning Capabilities in Small Models through Mixture-of-layers Distillation with Stepwise Attention on Key Information
abstract
The significant computational demands of large language models have increased interest in distilling reasoning abilities into smaller models via Chain-of-Thought (CoT) distillation.Current CoT distillation methods mainly focus on transferring teacher-generated rationales for complex reasoning to student models.However, they do not adequately explore teachers' dynamic attention toward critical information during reasoning.We find that language models exhibit progressive attention shifts towards key information during reasoning, which implies essential clues for drawing conclusions.Building on this observation and analysis, we introduce a novel CoT distillation framework that transfers the teacher's stepwise attention on key information to the student model.This establishes structured guidance for the student's progressive concentration on key information during reasoning.More importantly, we develop a Mixture of Layers module enabling dynamic alignment that adapts to different layers between the teacher and student.Our method achieves consistent performance improvements across multiple mathematical and commonsense reasoning datasets.To our knowledge, it is the first method to leverage stepwise attention within CoT distillation to improve small model reasoning. Input Question(a) A sample from the SVAMP dataset.The distilled student model fails to adequately utilize numerical information, leading to erroneous results, whereas the teacher model, during stepwise reasoning, effectively utilizes all numerical information to arrive at the correct final result.
Yao Chen 0009, Jiawei Sheng, Wenyuan Zhang 0002, Tingwen Liu
EMNLP1