Shinwoo Park

dblp:331/1122 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 WaterMod: Modular Token-Rank Partitioning for Probability-Balanced LLM Watermarking
abstract
Large language models now draft news, legal analyses, and software code with human-level fluency. At the same time, regulations such as the EU AI Act mandate that each synthetic passage carry an imperceptible, machine-verifiable mark for provenance. Conventional logit-based watermarks satisfy this requirement by selecting a pseudorandom green vocabulary at every decoding step and boosting its logits, yet the random split can exclude the highest-probability token and thus erode fluency. WaterMod mitigates this limitation through a probability-aware modular rule. The vocabulary is first sorted in descending model probability; the resulting ranks are then partitioned by the residue rank mod k, which distributes adjacent—and therefore semantically similar—tokens across different classes. A fixed bias of small magnitude is applied to one selected class. In the zero-bit setting (k=2), an entropy-adaptive gate selects either the even or the odd parity as the green list. Because the top two ranks fall into different parities, this choice embeds a detectable signal while guaranteeing that at least one high-probability token remains available for sampling. In the multi-bit regime (k>2), the current payload digit d selects the color class whose ranks satisfy rank mod k = d. Biasing the logits of that class embeds exactly one base-k digit—equivalently log2(k) bits—per decoding step, thereby enabling fine-grained provenance tracing. The same modular arithmetic therefore supports both binary attribution and rich payloads. Experimental results demonstrate that WaterMod consistently attains strong watermark detection performance while maintaining generation quality in both zero-bit and multi-bit settings. This robustness holds across a range of tasks, including natural language generation, mathematical reasoning, and code synthesis.
Shinwoo Park, Hyeseon Ahn, Yo-Sub Han
AAAI1
2026 A Linguistics-Aware LLM Watermarking via Syntactic Predictability
abstract
As large language models (LLMs) continue to advance rapidly, reliable governance tools have become critical.Publicly verifiable watermarking is particularly essential for fostering a trustworthy AI ecosystem.A central challenge persists: balancing text quality against detection robustness.Recent studies have sought to navigate this trade-off by leveraging signals from model output distributions (e.g., token-level entropy); however, their reliance on these modelspecific signals presents a significant barrier to public verification, as the detection process requires access to the logits of the underlying model.We introduce STELA, a novel framework that aligns watermark strength with the linguistic degrees of freedom inherent in language.STELA dynamically modulates the signal using part-of-speech (POS) n-gram-modeled linguistic indeterminacy, weakening it in grammatically constrained contexts to preserve quality and strengthening it in contexts with greater linguistic flexibility to enhance detectability.Our detector operates without access to any model logits, thus facilitating publicly verifiable detection.Through extensive experiments on typologically diverse languages-analytic English, isolating Chinese, and agglutinative Korean-we show that STELA surpasses prior methods in detection robustness.
Shinwoo Park, Hyeseon An, Yo-Sub Han
ACL (1)1
2026 EnCur: Curriculum-based in-context learning with structural encoding for code time complexity prediction
Joonghyuk Hahn, Aditi, Seung-Yeop Baik, Shinwoo Park, Sang-Ki Ko, Yo-Sub Han
Expert Syst. Appl.4
2025 KatFishNet: Detecting LLM-Generated Korean Text through Linguistic Feature Analysis
abstract
The rapid advancement of large language models (LLMs) increases the difficulty of distinguishing between human-written and LLM-generated text. Detecting LLM-generated text is crucial for upholding academic integrity, preventing plagiarism, protecting copyrights, and ensuring ethical research practices. Most prior studies on detecting LLM-generated text focus primarily on English text. However, languages with distinct morphological and syntactic characteristics require specialized detection approaches. Their unique structures and usage patterns hinder the direct application of methods primarily designed for English. Among such languages, we focus on Korean, which has relatively flexible spacing rules, a rich morphological system, and less frequent comma usage compared to English. We introduce KatFish, the first benchmark dataset for detecting LLM-generated Korean text. The dataset consists of text written by humans and generated by four LLMs across three genres. By examining spacing patterns, part-of-speech diversity, and comma usage, we illuminate the linguistic differences between human-written and LLM-generated Korean text. Building on these observations, we propose KatFishNet, a detection method specifically designed for the Korean language. KatFishNet achieves an average of 19.78% higher AUC-ROC compared to the best-performing existing detection method. Our code and data are available at https://github.com/Shinwoo-Park/katfishnet.
Shinwoo Park, Shubin Kim, Do-Kyung Kim, Yo-Sub Han
ACL (1)1
2025 Mondrian: A Framework for Logical Abstract (Re)Structuring
abstract
The well-known rhetorical framework, ABT (And, But, Therefore), mirrors natural human cognition in structuring an argument's logical progression -apropos to academic communication.However, distilling the complexities of research into clear and concise prose requires careful sequencing of ideas and formulating clear connections between them.This presents a quiet inequitability for contributions from authors who struggle with English proficiency or academic writing conventions.We see this as impetus to introduce: Mondrian, a framework that identifies the key components of an abstract and reorients itself to properly reflect the ABT logical progression.The framework is composed of a deconstruction stage, reconstruction stage, and rephrasing.We introduce a novel metric for evaluating deviation from ABT structure, named EB-DTW, which accounts for both ordinality and a non-uniform distribution of importance in a sequence.Our overall approach aims to improve the comprehensibility of academic writing, particularly for non-native English speakers, along with a complementary metric.The effectiveness of Mondrian is tested with automatic metrics and extensive human evaluation, and demonstrated through impressive quantitative and qualitative results, with organization and overall coherence of an abstract improving by an average of 27.71% and 24.71%.
Elizabeth Orwig, Shinwoo Park, Hyundong Jin, Yo-Sub Han
EMNLP2
2025 Helical Structured Soft Growing Robot for Hazardous Gas Suction in Inaccessible Environments
abstract
Immediate removal of hazardous gases is critical for ensuring safety. Traditional methods, such as portable ventilation equipment, are difficult to use when hazardous gases are released in inaccessible environments. In this paper, we propose a novel mechanism that integrates an inflatable helical structure into a soft growing robot. The proposed mechanism is capable of performing suction through its inner channel after navigating complex environments, while maintaining the inherent advantages of the soft growing robot as it grows. The mechanism operates in two phases: a growing phase, in which the robot extends by eversion, and a suction phase, in which suction is performed through the inner channel of the robot. Experiments and demonstrations were conducted to evaluate the performance of the proposed mechanism. The experimental results confirmed the ability to maintain the passageway shape of the inner channel during suction operations and provided a design guideline. The demonstration validated that the mechanism can effectively navigate inaccessible environments and perform suction to remove hazardous gases.
Nam Gyun Kim, Dongoh Seo, Shinwoo Park, Jee-Hwan Ryu
ICRA4
2025 Advanced code time complexity prediction approach using contrastive learning
Shinwoo Park, Joonghyuk Hahn, Elizabeth Orwig, Sang-Ki Ko, Yo-Sub Han
Eng. Appl. Artif. Intell.1
2025 Detecting code paraphrased by large language models using coding style features
Shinwoo Park, Hyundong Jin, Jeong-Won Cha, Yo-Sub Han
Eng. Appl. Artif. Intell.1
2023 Contrastive Learning with Keyword-based Data Augmentation for Code Search and Code Question Answering
abstract
The semantic code search is to find code snippets from the collection of candidate code snippets with respect to a user query that describes functionality.Recent work on code search proposes data augmentation of queries for contrastive learning.This data augmentation approach modifies random words in queries.When a user web query for searching code snippet is too brief, the important word that represents the search intent of the query could be undesirably modified.A code snippet has informative components such as function name and documentation that describe its functionality.We propose to utilize these code components to identify important words and preserve them in the data augmentation step.We present Key-DAC (Keyword-based Data Augmentation for Contrastive learning) that identifies important words for code search from queries and code components based on term matching.KeyDAC augments query-code pairs while preserving keywords, and then leverages generated training instances for contrastive learning.We use KeyDAC to fine-tune various pre-trained language models and evaluate the performance of code search and code question answering via CoSQA and WebQueryTest.The experimental results confirm that KeyDAC substantially outperforms the current state-of-the-art performance, and achieves the new state-of-the-arts for both tasks.
Shinwoo Park, Youngwook Kim 0002, Yo-Sub Han
EACL1
2022 Generalizable Implicit Hate Speech Detection Using Contrastive Learning
abstract
Hate speech detection has gained increasing attention with the growing prevalence of hateful contents. When a text contains an obvious hate word or expression, it is fairly easy to detect it. However, it is challenging to identify implicit hate speech in nuance or context when there are insufficient lexical cues. Recently, there are several attempts to detect implicit hate speech leveraging pre-trained language models such as BERT and HateBERT. Fine-tuning on an implicit hate speech dataset shows satisfactory performance when evaluated on the test set of the dataset used for training. However, we empirically confirm that the performance drops at least 12.5%p in F1 score when tested on the dataset that is different from the one used for training. We tackle this cross-dataset underperforming problem using contrastive learning. Based on our observation of common underlying implications in various forms of hate posts, we propose a novel contrastive learning method, ImpCon, that pulls an implication and its corresponding posts close in representation space. We evaluate the effectiveness of ImpCon by running cross-dataset evaluation on three implicit hate speech benchmarks. The experimental results on cross-dataset show that ImpCon improves at most 9.10% on BERT, and 8.71% on HateBERT.
Youngwook Kim 0002, Shinwoo Park, Yo-Sub Han
COLING2