Eunhye Jeong

dblp:389/7248 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 50% Trustworthy machine learning · 25% Image recognition and object detection · 25%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
attention-based detection
0.912025
CLAWS: Creativity detection for LLM-generated solutions using Attention Window of Sections · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
creativity evaluation
0.912025
CLAWS: Creativity detection for LLM-generated solutions using Attention Window of Sections · NeurIPS 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
CLAWS: Creativity detection for LLM-generated solutions using Attention Window of Sections · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
CLAWS: Creativity detection for LLM-generated solutions using Attention Window of Sections · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

white-box detection · 0.9attention weight analysis · 0.9
YearPublicationVenuePosition
2025 CLAWS: Creativity detection for LLM-generated solutions using Attention Window of Sections
abstract
Recent advances in enhancing the reasoning ability of Large Language Models (LLMs) have been remarkably successful. LLMs trained with Reinforcement Learning (RL) for reasoning demonstrate strong performance in challenging tasks such as mathematics and coding, even with relatively small model sizes. However, despite these impressive improvements in task accuracy, the assessment of creativity in LLM generations has been largely overlooked in reasoning tasks, in contrast to writing tasks. The lack of research on creativity assessment in reasoning primarily stems from two challenges: (1) the difficulty of defining the range of creativity, and (2) the necessity of human evaluation in the assessment process. To address these challenges, we propose CLAWS, a novel method that defines and classifies mathematical solutions into Typical, Creative, and Hallucinated categories without human evaluation, by leveraging attention weights across prompt sections and output. CLAWS outperforms five existing white-box detection methods—Perplexity, Logit Entropy, Window Entropy, Hidden Score, and Attention Score—on five 7–8B math RL models (DeepSeek, Qwen, Mathstral, OpenMath2, and Oreal). We validate CLAWS on 4,545 math problems collected from 181 math contests (A(J)HSME, AMC, AIME). Our code is available at https://github.com/kkt94/CLAWS.
Keuntae Kim, Eunhye Jeong, Sehyeon Lee, Seohee Yoon
NeurIPS2