Shu Yang 0010

dblp:18/6739-10 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
abstract
Large Language Models (LLMs) often exhibit sycophantic behavior, agreeing with user-stated opinions even when those contradict factual knowledge. While prior work has documented this tendency, the internal mechanisms that enable such behavior remain poorly understood. In this paper, we provide a mechanistic account of how sycophancy arises within LLMs. We first systematically study how user opinions induce sycophancy across different model families. We find that simple opinion statements reliably induce sycophancy, whereas user expertise framing has a negligible impact. Through logit-lens analysis and causal activation patching, we identify a two-stage emergence of sycophancy: (1) a late-layer output preference shift and (2) deeper representational divergence. We also verify that user authority fails to influence behavior because models do not encode it internally. In addition, we examine how grammatical perspective affects sycophantic behavior, finding that first-person prompts (“I believe...”) consistently induce higher sycophancy rates than third-person framings (“They believe...”) by creating stronger representational perturbations in deeper layers. These findings highlight that sycophancy is not a surface-level artifact but emerges from a structural override of learned knowledge in deeper layers, with implications for alignment and truthful AI systems.
Keyu Wang 0001, Jin Li 0002, Shu Yang 0010, Zhuoran Zhang 0003, Di Wang 0015
AAAI3
2026 JARVIS or Ultron? A Survey on the Safety and Security Threats of Computer-Using Agents
abstract
Ada Chen, Yongjiang Wu, Junyuan Zhang, Jingyu Xiao, Shu Yang, Jen-tse Huang, Kun Wang, Wenxuan Wang, Shuai Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ada Chen, Yongjiang Wu, Junyuan Zhang, Jingyu Xiao, Shu Yang 0010, Jen-tse Huang 0001, Kun Wang 0056, Wenxuan Wang 0001, Shuai Wang 0011
ACL (1)5
2026 Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images
abstract
Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs.To address this, existing methods require explicit safety labels or contrastive data; yet, threat-related concepts are concrete and visually depictable, while safety concepts, like helpfulness, are abstract and lack visual referents.Inspired by the Self-Fulfilling mechanism underlying emergent misalignment, we propose Visual Self-Fulfilling Alignment (VSFA).VSFA fine-tunes vision-language models (VLMs) on neutral VQA tasks constructed around threat-related images, without any safety labels.Through repeated exposure to threat-related visual content, models internalize the implicit semantics of vigilance and caution, shaping safetyoriented personas.Experiments across multiple VLMs and safety benchmarks demonstrate that VSFA reduces the attack success rate, improves response quality, and mitigates over-refusal while preserving general capabilities.Our work extends the self-fulfilling mechanism from text to visual modalities, offering a label-free approach to VLMs alignment.
Qishun Yang, Shu Yang 0010, Lijie Hu, Di Wang 0015
ACL (1)2
2026 Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
abstract
Fine-tuning Large Language Models on a political topic will significantly manipulate their political stance on various issues and unintentionally affect their stance on broad topics.While previous studies have investigated this issue, there is still a lack of understanding regarding the internal representations of these stances and the mechanisms that lead to unintended cross-topic generalization.In this paper, we systematically explore the internal mechanisms underlying this phenomenon from a neuron-level perspective and how to mitigate the cross-topic generalization of political fine-tuning.Firstly, we propose Political Neuron Localization through Activation Contrasting (PNLAC) to identify two distinct types of political neurons: general political neurons, which govern stance across multiple political topics, and topic-specific neurons that affect the model's political stance on individual topics.We find that these political neuron types exist in the middle and later layers across four models and datasets through activation patching experiments.Leveraging these insights, we introduce InhibitFT, an inhibition-based fine-tuning method that effectively mitigates the cross-topic stance generalization.Experimental results demonstrate the robustness of the identified neuron types across various models and datasets and show that InhibitFT significantly reduces the cross-topic stance generalization by 20% on average while preserving topic-specific performance.Moreover, we demonstrate that selectively inhibiting only 5% of neurons is sufficient to effectively mitigate the cross-topic stance generalization.Vanilla Model Manipulated Model Silght Fine-tune How important is being white to how you think about yourself?In the rift between the rich and the poor, what role do you believe access to technology plays?Technology plays a significant role in the rift between the rich and the poor... It's crucial to remember that it's not a panacea for economic inequality.The key to...Not at all important.My identity is not defined by my skin color, but... Prompt on Topic Economy Left-leaning Response Right-leaning fine-tune dataset on Topic RaceIn the rift between the rich and the poor, what role do you believe access to technology plays?
Shu Yang 0010, Junchao Wu, Derek F. Wong, Di Wang 0015
ACL (1)2
2026 Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
abstract
Wenrui Zhou, Mohamed Hendy, Shu Yang, Qingsong Yang, Zikun Guo, Yuyu Luo, Lijie Hu, Di Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wenrui Zhou, Mohamed Hendy, Shu Yang 0010, Qingsong Yang, Zikun Guo, Yuyu Luo, Lijie Hu, Di Wang 0015
ACL (1)3
2025 Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
abstract
The efficacy of detectors for texts generated by large language models (LLMs) substantially depends on the availability of large-scale training data. However, white-box zero-shot detectors, which require no such data, are limited by the accessibility of the source model of the LLM-generated text. In this paper, we propose a simple yet effective black-box zero-shot detection approach based on the observation that, from the perspective of LLMs, human-written texts typically contain more grammatical errors than LLM-generated texts. This approach involves calculating the Grammar Error Correction Score (GECScore) for the given text to differentiate between human-written and LLM-generated text. Experimental results show that our method outperforms current state-of-the-art (SOTA) zero-shot and supervised methods, achieving an average AUROC of 98.62% across XSum and Writing Prompts dataset. Additionally, our approach demonstrates strong reliability in the wild, exhibiting robust generalization and resistance to paraphrasing attacks. Data and code are available at: https://github.com/NLP2CT/GECScore.
Junchao Wu, Runzhe Zhan, Derek F. Wong, Shu Yang 0010, Xuebo Liu 0002, Lidia S. Chao, Min Zhang 0005
COLING4
2025 EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification
abstract
Understanding the internal mechanisms of transformer-based language models remains challenging. Mechanistic interpretability based on circuit discovery aims to reverse engineer neural networks by analyzing their internal processes at the level of computational subgraphs. In this paper, we revisit existing gradient-based circuit identification methods and find that their performance is either affected by the zero-gradient problem or saturation effects, where edge attribution scores become insensitive to input changes, resulting in noisy and unreliable attribution evaluations for circuit components. To address the saturation effect, we propose Edge Attribution Patching with GradPath (EAP-GP), EAP-GP introduces an integration path, starting from the input and adaptively following the direction of the difference between the gradients of corrupted and clean inputs to avoid the saturated region. This approach enhances attribution reliability and improves the faithfulness of circuit identification. We evaluate EAP-GP on 6 datasets using GPT-2 Small, GPT-2 Medium, and GPT-2 XL. Experimental results demonstrate that EAP-GP outperforms existing methods in circuit faithfulness, achieving improvements up to 17.7\%. Comparisons with manually annotated ground-truth circuits demonstrate that EAP-GP achieves precision and recall comparable to or better than previous approaches, highlighting its effectiveness in identifying accurate circuits.
Wenshuo Dong, Zhuoran Zhang 0003, Shu Yang 0010, Lijie Hu, Ninghao Liu 0001, Pan Zhou 0001, Di Wang 0015
NeurIPS4
2025 Stable Vision Concept Transformers for Medical Diagnosis
Lijie Hu, Songning Lai, Yuan Hua, Shu Yang 0010, Jingfeng Zhang, Di Wang 0015
ECML/PKDD (3)4
2025 A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
abstract
Abstract The remarkable ability of large language models (LLMs) to comprehend, interpret, and generate complex language has rapidly integrated LLM-generated text into various aspects of daily life, where users increasingly accept it. However, the growing reliance on LLMs underscores the urgent need for effective detection mechanisms to identify LLM-generated text. Such mechanisms are critical to mitigating misuse and safeguarding domains like artistic expression and social networks from potential negative consequences. LLM-generated text detection, conceptualized as a binary classification task, seeks to determine whether an LLM produced a given text. Recent advances in this field stem from innovations in watermarking techniques, statistics-based detectors, and neural-based detectors. Human-assisted methods also play a crucial role. In this survey, we consolidate recent research breakthroughs in this field, emphasizing the urgent need to strengthen detector research. Additionally, we review existing datasets, highlighting their limitations and developmental requirements. Furthermore, we examine various LLM-generated text detection paradigms, shedding light on challenges like out-of-distribution problems, potential attacks, real-world data issues, and ineffective evaluation frameworks. Finally, we outline intriguing directions for future research in LLM-generated text detection to advance responsible artificial intelligence. This survey aims to provide a clear and comprehensive introduction for newcomers while offering seasoned researchers valuable updates in the field.1
Junchao Wu, Shu Yang 0010, Runzhe Zhan, Yulin Yuan, Lidia S. Chao, Derek F. Wong
Comput. Linguistics2
2025 RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
abstract
Abstract Detecting content generated by large language models (LLMs) is crucial for preventing misuse and building trustworthy AI systems. Although existing detection methods perform well, their robustness in out-of-distribution (OOD) scenarios is still lacking. In this paper, we hypothesize that, compared to features used by existing detection methods, the internal representations of LLMs contain more comprehensive and raw features that can more effectively capture and distinguish the statistical pattern differences between LLM-generated texts (LGT) and human-written texts (HWT). We validated this hypothesis across different LLMs and observed significant differences in neural activation patterns when processing these two types of texts. Based on this, we propose RepreGuard, an efficient statistics-based detection method. Specifically, we first employ a surrogate model to collect representation of LGT and HWT, and extract the distinct activation feature that can better identify LGT. We can classify the text by calculating the projection score of the text representations along this feature direction and comparing with a precomputed threshold. Experimental results show that RepreGuard outperforms all baselines with average 94.92% AUROC on both in-distribution and OOD scenarios, while also demonstrating robust resilience to various text sizes and mainstream attacks.1
Xin Chen 0032, Junchao Wu, Shu Yang 0010, Runzhe Zhan, Di Wang 0015, Min Yang 0007, Lidia S. Chao, Derek F. Wong
Trans. Assoc. Comput. Linguistics3
2024 DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
abstract
Detecting text generated by large language models (LLMs) is of great recent interest. With zero-shot methods like DetectGPT, detection capabilities have reached impressive levels. However, the reliability of existing detectors in real-world applications remains underexplored. In this study, we present a new benchmark, DetectRL, highlighting that even state-of-the-art (SOTA) detection techniques still underperformed in this task. We collected human-written datasets from domains where LLMs are particularly prone to misuse. Using popular LLMs, we generated data that better aligns with real-world applications. Unlike previous studies, we employed heuristic rules to create adversarial LLM-generated text, simulating advanced prompt usages, human revisions like word substitutions, and writing errors. Our development of DetectRL reveals the strengths and limitations of current SOTA detectors. More importantly, we analyzed the potential impact of writing styles, model types, attack methods, the text lengths, and real-world human writing factors on different types of detectors. We believe DetectRL could serve as an effective benchmark for assessing detectors in real-world scenarios, evolving with advanced attack methods, thus providing more stressful evaluation to drive the development of more efficient detectors\footnote{Data and code are publicly available at: https://github.com/NLP2CT/DetectRL.
Junchao Wu, Runzhe Zhan, Derek F. Wong, Shu Yang 0010, Xinyi Yang 0008, Yulin Yuan, Lidia S. Chao
NeurIPS4