Yehan Yang

dblp:410/1744 · also YeHan Yang · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0003-3063-3896ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
2 papers
Digital forensics and information hiding · 54% Web and mobile security · 46%
Artificial intelligence
2 papers
Trustworthy machine learning · 59% Information extraction and text analysis · 41%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Digital forensics and information hiding › synthetic media detection
machine-generated text detection
1.012026
Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection · ACL (1) 2026
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring · NeurIPS 2025
Web and mobile security
harmful content detection
0.912025
From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring · NeurIPS 2025
Natural language and speech › Information extraction and text analysis
discourse analysis
0.312026
Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection · ACL (1) 2026
Natural language and speech › Information extraction and text analysis › discourse analysis › discourse parsing
rhetorical structure theory
0.312026
Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

rhetorical structure theory · 2.0elementary discourse unit features · 2.0token-level annotation · 1.7dual supervision · 1.7
YearPublicationVenuePosition
2026 Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection
abstract
The misuse of large language models (LLMs) requires precise detection of synthetic text.Existing works mainly follow binary or ternary classification settings, which can only distinguish pure human/LLM text or collaborative text at best.This remains insufficient for the nuanced regulation, as the LLM-polished human text and humanized LLM text often trigger different policy consequences.In this paper, we explore fine-grained LLM-generated text detection under a rigorous four-class setting.To handle such complexities, we propose RACE (Rhetorical Analysis for Creator-Editor Modeling), a fine-grained detection method that characterizes the distinct signatures of creator and editor.Specifically, RACE utilizes Rhetorical Structure Theory (RST) to construct a logic graph for the creator's foundation while extracting Elementary Discourse Unit (EDU)-level features for the editor's style.Experiments show that RACE outperforms 12 baselines in identifying fine-grained types with low false alarms, offering a policy-aligned solution for LLM regulation.
Yang Li 0196, Qiang Sheng 0001, Zhengjia Wang 0001, Yehan Yang, Danding Wang, Juan Cao 0001
ACL (1)4
2025 I-LEAD: A Digital-Intelligence-Powered Ecosystem for Innovation and Entrepreneurship Education
abstract
Generative artificial intelligence and large model agents are revolutionizing the higher education, influencing everything from talent development frameworks and teaching methodologies to knowledge acquisition processes and research paradigms. Meanwhile, in the context of digital transformation and technological innovation, data as a production factor and emerging productive forces are reshaping the requirements and demands for cultivating high-level interdisciplinary engineering talents. This paper draws on the joint educational achievements between Beijing University of Posts and Telecommunications and Queen Mary University of London, focusing on the coconstruction and sharing of experimental resources, as well as innovation-driven entrepreneurship education. It introduces a digital-intelligence-powered educational platform aimed at fostering internationally-minded, innovative, and outstanding talents. The platform's core philosophy, development strategy, functional modules, and technical framework are detailed. The wide recognition and interest among stakeholders further validate its potential to support the development of a cross-disciplinary, cross-professional, and cross-national ecosystem for innovation and entrepreneurship education. As a key outcome of the 20th Anniversary Development Conference of Joint Education of our two universities, I-LEAD is a platform designed to cultivate students' comprehensive innovative capabilities. Leveraging large language models (LLMs) and multi-agent technology, we have developed BUPT iMentor, an intelligent agent for innovation and entrepreneurship guidance; and BUPT EnPower, a cultivation assistant for personalized longlife learning. By LLMs with a robust knowledge based augmented generation and fine tuning, this tool effectively addresses common student challenges during innovation & entrepreneurship projects, such as idea generation and validation, access to relevant learning resources. I-LEAD provides comprehensive support, including customized course creation, problem-solving guidance, real-time interactive Q&A, and learning progress monitoring. This empowers students to independently plan their learning journeys and holistically enhance their academic and innovative skills. The effectiveness and practicality of the platform have been validated through a questionnaire survey. Therefore, we are extensively gathering feedback and continuously optimizing the platform's services. The teacher-student collaborative learning represents the future of higher education. This student-centered education system provides a platform for that.
Shuchang Liu 0005, Minghui Pan, Yehan Yang
EDUCON3
2025 From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
abstract
Though safety alignment has been applied to most large language models (LLMs), LLM service providers generally deploy a subsequent moderation as the external safety guardrail in real-world products. Existing moderators mainly practice a conventional full detection, which determines the harmfulness based on the complete LLM output, causing high service latency. Recent works pay more attention to partial detection where moderators oversee the generation midway and early stop the output if harmfulness is detected, but they directly apply moderators trained with the full detection paradigm to incomplete outputs, introducing a training-inference gap that lowers the performance. In this paper, we explore how to form a data-and-model solution that natively supports partial detection. For the data, we construct **FineHarm**, a dataset consisting of 29K prompt-response pairs with fine-grained token-level annotations to provide reasonable supervision for token-level training. Then, we propose the **Streaming Content Monitor (SCM)**, which is trained with dual supervision of response- and token-level labels and can follow the output stream of LLM to make a timely judgment of harmfulness. Experiments show that SCM gains 0.95+ in macro F1 score that is comparable to full-detection, by only seeing the first 18% of tokens in responses on average. Moreover, the SCM can serve as a pseudo-harmfulness annotator for improving safety alignment and lead to a higher harmlessness score than DPO.
Yang Li 0196, Qiang Sheng 0001, Yehan Yang, Xueyao Zhang, Juan Cao 0001
NeurIPS3