Shaotian Yan

dblp:274/1197 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 56% Efficient and distributed learning · 10% Vision and language · 10%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.722025
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models · ICLR 2025
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning · ACL (1) 2025
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
attention intervention
0.912025
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models · ICLR 2025
Natural language and speech › Language models and text generation
hallucination detection
0.912025
Shallow Focus, Deep Fixes: Enhancing Shallow Layers Vision Attention Sinks to Alleviate Hallucination in LVLMs · EMNLP 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Shallow Focus, Deep Fixes: Enhancing Shallow Layers Vision Attention Sinks to Alleviate Hallucination in LVLMs · EMNLP 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning · ACL (1) 2025
Natural language and speech › Language models and text generation › prompting
chain-of-thought prompting
0.812024
Instance-adaptive Zero-shot Chain-of-Thought Prompting · NeurIPS 2024
Natural language and speech › Language models and text generation
prompting
0.812024
Instance-adaptive Zero-shot Chain-of-Thought Prompting · NeurIPS 2024
Data mining
clustering
0.612022
MPC: Multi-view Probabilistic Clustering · CVPR 2022
Data mining › clustering › multi-view clustering
incomplete multi-view clustering
0.612022
MPC: Multi-view Probabilistic Clustering · CVPR 2022
Data mining › clustering
multi-view clustering
0.612022
MPC: Multi-view Probabilistic Clustering · CVPR 2022
Computer vision › Segmentation and scene understanding
scene graph generation
0.412020
PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph Generation · ACM Multimedia 2020
Machine learning › Deep learning architectures and training
attention mechanism
0.312025
Shallow Focus, Deep Fixes: Enhancing Shallow Layers Vision Attention Sinks to Alleviate Hallucination in LVLMs · EMNLP 2025
Natural language and speech › Language models and text generation › in-context learning
few-shot prompting
0.312025
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models · ICLR 2025
Natural language and speech › Language models and text generation
in-context learning
0.312025
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models · ICLR 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.212024
Instance-adaptive Zero-shot Chain-of-Thought Prompting · NeurIPS 2024
Machine learning › Learning paradigms
class imbalance
0.112020
PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph Generation · ACM Multimedia 2020

Methods — techniques the papers use, named apart from their topics

representation finetuning · 0.9chain-of-thought prompting · 0.9attention sink analysis · 0.9attention manipulation · 0.9attention analysis · 0.9information flow analysis · 0.8probabilistic clustering · 0.6path propagation · 0.6graph-context-aware refinement · 0.6reweighting · 0.4predicate-correlation perception learning · 0.4graph encoder · 0.4
YearPublicationVenuePosition
2025 Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning
abstract
Chenxi Huang, Shaotian Yan, Liang Xie, Binbin Lin, Sinan Fan, Yue Xin, Deng Cai, Chen Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chenxi Huang 0004, Shaotian Yan, Liang Xie 0003, Binbin Lin 0001, Sinan Fan, Deng Cai 0001, Chen Shen 0003, Jieping Ye
ACL (1)2
2025 Shallow Focus, Deep Fixes: Enhancing Shallow Layers Vision Attention Sinks to Alleviate Hallucination in LVLMs
abstract
Xiaofeng Zhang, Yihao Quan, Chen Shen, Chaochen Gu, Xiaosong Yuan, Shaotian Yan, Jiawei Cao, Hao Cheng, Kaijie Wu, Jieping Ye. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Xiaofeng Zhang 0006, Yihao Quan, Chen Shen 0003, Chaochen Gu, Xiaosong Yuan, Shaotian Yan, Hao Cheng 0004, Kaijie Wu 0002, Jieping Ye
EMNLP6
2025 Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
abstract
Few-shot Chain-of-Thought (CoT) significantly enhances the reasoning capabilities of large language models (LLMs), functioning as a whole to guide these models in generating reasoning steps toward final answers. However, we observe that isolated segments, words, or tokens within CoT demonstrations can unexpectedly disrupt the generation process of LLMs. The model may overly concentrate on certain local information present in the demonstration, introducing irrelevant noise into the reasoning process and potentially leading to incorrect answers. In this paper, we investigate the underlying mechanism of CoT through dynamically tracing and manipulating the inner workings of LLMs at each output step, which demonstrates that tokens exhibiting specific attention characteristics are more likely to induce the model to take things out of context; these tokens directly attend to the hidden states tied with prediction, without substantial integration of non-local information. Building upon these insights, we propose a Few-shot Attention Intervention method (FAI) that dynamically analyzes the attention patterns of demonstrations to accurately identify these tokens and subsequently make targeted adjustments to the attention weights to effectively suppress their distracting effect on LLMs. Comprehensive experiments across multiple benchmarks demonstrate consistent improvements over baseline methods, with a remarkable 5.91\% improvement on the AQuA dataset, further highlighting the effectiveness of FAI.
Shaotian Yan, Chen Shen 0003, Wenxiao Wang 0001, Liang Xie 0003, Junjie Liu 0002, Jieping Ye
ICLR1
2025 From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
abstract
Xiaofeng Zhang, Yihao Quan, Chen Shen, Xiaosong Yuan, Shaotian Yan, Liang Xie, Wenxiao Wang, Chaochen Gu, Hao Tang, Jieping Ye. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xiaofeng Zhang 0006, Yihao Quan, Chen Shen 0003, Xiaosong Yuan, Shaotian Yan, Liang Xie 0003, Wenxiao Wang 0001, Chaochen Gu, Hao Tang 0005, Jieping Ye
NAACL (Long Papers)5
2024 Instance-adaptive Zero-shot Chain-of-Thought Prompting
abstract
Zero-shot Chain-of-Thought (CoT) prompting emerges as a simple and effective strategy for enhancing the performance of large language models (LLMs) in real-world reasoning tasks. Nonetheless, the efficacy of a singular, task-level prompt uniformly applied across the whole of instances is inherently limited since one prompt cannot be a good partner for all, a more appropriate approach should consider the interaction between the prompt and each instance meticulously. This work introduces an instance-adaptive prompting algorithm as an alternative zero-shot CoT reasoning scheme by adaptively differentiating good and bad prompts. Concretely, we first employ analysis on LLMs through the lens of information flow to detect the mechanism under zero-shot CoT reasoning, in which we discover that information flows from question to prompt and question to rationale jointly influence the reasoning results most. We notice that a better zero-shot CoT reasoning needs the prompt to obtain semantic information from the question then the rationale aggregates sufficient information from the question directly and via the prompt indirectly. On the contrary, lacking any of those would probably lead to a bad one. Stem from that, we further propose an instance-adaptive prompting strategy (IAP) for zero-shot CoT reasoning. Experiments conducted with LLaMA-2, LLaMA-3, and Qwen on math, logic, and commonsense reasoning tasks (e.g., GSM8K, MMLU, Causal Judgement) obtain consistent improvement, demonstrating that the instance-adaptive zero-shot CoT prompting performs better than other task-level methods with some curated prompts or sophisticated procedures, showing the significance of our findings in the zero-shot CoT reasoning mechanism.
Xiaosong Yuan, Chen Shen 0003, Shaotian Yan, Xiaofeng Zhang 0006, Liang Xie 0003, Wenxiao Wang 0001, Renchu Guan, Ying Wang 0009, Jieping Ye
NeurIPS3
2022 MPC: Multi-view Probabilistic Clustering
abstract
Despite the promising progress having been made, the two challenges of multi-view clustering (MVC) are still waiting for better solutions: i) Most existing methods are either not qualified or require additional steps for incomplete multi-view clustering and ii) noise or outliers might significantly degrade the overall clustering performance. In this paper, we propose a novel unified framework for incomplete and complete MVC named multi-view probabilistic clustering (MPC). MPC equivalently transforms multi-view pairwise posterior matching probability into composition of each view's individual distribution, which tolerates data missing and might extend to any number of views. Then graph-context-aware refinement with path propagation and co-neighbor propagation is used to refine pairwise probability, which alleviates the impact of noise and outliers. Finally, MPC also equivalently transforms probabilistic clustering's objective to avoid complete pairwise computation and adjusts clustering assignments by maximizing joint probability iteratively. Extensive experiments on multiple benchmarks for incomplete and complete MVC show that MPC significantly outperforms previous state-of-the-art methods in both effectiveness and efficiency.
Junjie Liu 0002, Junlong Liu, Shaotian Yan, Rongxin Jiang 0001, Xiang Tian 0002, Boxuan Gu, Yaowu Chen, Chen Shen 0003, Jianqiang Huang 0001
CVPR3
2020 PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph Generation
abstract
Today's scene graph generation (SGG) task is largely limited in realistic scenarios, mainly due to the extremely long-tailed bias of predicate annotation distribution. Thus, tackling the class imbalance trouble of SGG is critical and challenging. In this paper, we first discover that when predicate labels have strong correlation with each other, prevalent re-balancing strategies (e.g., re-sampling and re-weighting) will give rise to either over-fitting the tail data (e.g., bench sitting on sidewalk rather than on), or still suffering the adverse effect from the original uneven distribution (e.g., aggregating varied parked on/standing on/sitting on into on). We argue the principal reason is that re-balancing strategies are sensitive to the frequencies of predicates yet blind to their relatedness, which may play a more important role to promote the learning of predicate features. Therefore, we propose a novel Predicate-Correlation Perception Learning (PCPL for short) scheme to adaptively seek out appropriate loss weights by directly perceiving and utilizing the correlation among predicate classes. Moreover, our PCPL framework is further equipped with a graph encoder module to better extract context features. Extensive experiments on the benchmark VG150 dataset show that the proposed PCPL performs markedly better on tail classes while well-preserving the performance on head ones, which significantly outperforms previous state-of-the-art methods.
Shaotian Yan, Chen Shen 0003, Zhongming Jin 0001, Jianqiang Huang 0001, Rongxin Jiang 0001, Yaowu Chen, Xian-Sheng Hua 0001
ACM Multimedia1