Joonghyuk Hahn

dblp:304/4027 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0009-0000-5890-4916ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Theory of computation · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 EnCur: Curriculum-based in-context learning with structural encoding for code time complexity prediction
Joonghyuk Hahn, Aditi, Seung-Yeop Baik, Shinwoo Park, Sang-Ki Ko, Yo-Sub Han
Expert Syst. Appl.1
2025 AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
abstract
Implicit hate speech involves subtle and indirect expressions of prejudice or hostility toward a group.Detecting it is challenging because it relies on nuanced context and implication rather than explicit offensive language.Current approaches rely on contrastive learning, which is shown to be effective on distinguishing hate and non-hate sentences.Humans, however, detect implicit hate speech by first identifying specific targets within the text and subsequently interpreting how these targets relate to their surrounding context.Motivated by this reasoning process, we propose Ample-Hate, a novel approach designed to mirror human inference for implicit hate detection.Am-pleHate identifies explicit targets using a pretrained Named Entity Recognition model and captures implicit target information via [CLS] tokens.It computes attention-based relationships between explicit, implicit targets and sentence context and then, directly injects these relational vectors into the final sentence representation.This amplifies the critical signals of target-context relations for determining implicit hate.Experiments demonstrate that Am-pleHate achieves state-of-the-art performance, outperforming contrastive learning baselines by an average of 82.14% and achieves faster convergence.Qualitative analyses further reveal that attention patterns produced by Am-pleHate closely align with human judgement, underscoring its interpretability and robustness.
Joonghyuk Hahn, Hyeseon Ahn, Yo-Sub Han
EMNLP2
2025 TCProF:Time-Complexity Prediction SSL Framework
abstract
Joonghyuk Hahn, Hyeseon Ahn, Jungin Kim, Soohan Lim, Yo-Sub Han. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Joonghyuk Hahn, Hyeseon Ahn, Jungin Kim, Soohan Lim, Yo-Sub Han
NAACL (Long Papers)1
2025 Advanced code time complexity prediction approach using contrastive learning
Shinwoo Park, Joonghyuk Hahn, Elizabeth Orwig, Sang-Ki Ko, Yo-Sub Han
Eng. Appl. Artif. Intell.2
2024 SuperST: Superficial Self-Training for Few-Shot Text Classification
abstract
In few-shot text classification, self-training is a popular tool in semi-supervised learning (SSL). It relies on pseudo-labels to expand data, which has demonstrated success. However, these pseudo-labels contain potential noise and provoke a risk of underfitting the decision boundary. While the pseudo-labeled data can indeed be noisy, fully acquiring this flawed data can result in the accumulation of further noise and eventually impacting the model performance. Consequently, self-training presents a challenge: mitigating the accumulation of noise in the pseudo-labels. Confronting this challenge, we introduce superficial learning, inspired by pedagogy’s focus on essential knowledge. Superficial learning in pedagogy is a learning scheme that only learns the material ‘at some extent’, not fully understanding the material. This approach is usually avoided in education but counter-intuitively in our context, we employ superficial learning to acquire only the necessary context from noisy data, effectively avoiding the noise. This concept serves as the foundation for SuperST, our self-training framework. SuperST applies superficial learning to the noisy data and fine-tuning to the less noisy data, creating an efficient learning cycle that prevents overfitting to the noise and spans the decision boundary effectively. Notably, SuperST improves the classifier accuracy for few-shot text classification by 18.5% at most and 8% in average, compared with the state-of-the-art SSL baselines. We substantiate our claim through empirical experiments and decision boundary analysis.
Ju Hyoung Lee, Joonghyuk Hahn, Jiho Park 0002, Yo-Sub Han
LREC/COLING2
2024 Universal Rewriting Rules for the Parikh Matrix Injectivity Problem
Ingyu Baek, Joonghyuk Hahn, Yo-Sub Han, Kai Salomaa
DLT2
2024 On the Decidability of Infix Inclusion Problem
Hyunjoon Cheon, Joonghyuk Hahn, Yo-Sub Han
Theory Comput. Syst.2
2023 ATHENA: Mathematical Reasoning with Thought Expansion
abstract
Solving math word problems depends on how to articulate the problems, the lens through which models view human linguistic expressions.Real-world settings count on such a method even more due to the diverse practices of the same mathematical operations.Earlier works constrain available thinking processes by limited prediction strategies without considering their significance in acquiring mathematical knowledge.We introduce Attention-based THought Expansion Network Architecture (ATHENA) to tackle the challenges of real-world practices by mimicking human thought expansion mechanisms in the form of neural network propagation.A thought expansion recurrently generates the candidates carrying the thoughts of possible math expressions driven from the previous step and yields reasonable thoughts by selecting the valid pathways to the goal.Our experiments show that ATHENA achieves a new state-of-the-art stage toward the ideal model that is compelling in variant questions even when the informativeness in training examples is restricted. 1 Context The school playground was originally [80] meters long and [40] meters wide.Later when the school is remodeled, the length is increased by [10] meters and the width is increased by [15] meters.Train on an example of a question-solution pair under the context above.Question How many square meters is the original playground area?Solution (80 × 40)Test on variant questions that share the context above.Q0 How many times the length of the original playground was the width?DeductReasoner ATHENA UnbiasedMWP (80 + 10) × (40 + 15) -(80 × 40) (X) UnbiasedMWP (80 + 10) × (40 -15) (X) UnbiasedMWP (1:N) 80 ÷ 40 (O) UnbiasedMWP (1:N) 80 ÷ 40 (O) Q1 How many square meters is the current playground area?DeductReasoner ATHENA UnbiasedMWP (80 + 10) × (40 + 15) -(80 × 40) (X) UnbiasedMWP (80 + 10) × (40 + 15) (O) UnbiasedMWP (1:N) 80 × 40 (X) UnbiasedMWP (1:N) (80 + 10) × (40 + 15) (O) Q2 How many square meters are increased by the current playground area compared to the original one?DeductReasoner ATHENA UnbiasedMWP (80 + 10) × (40 + 15) -(80 × 40) (O) UnbiasedMWP (80 + 10) × (40 + 15) -(80 × 40) (O) UnbiasedMWP (1:N) 80 × 40 (X) UnbiasedMWP (1:N) (80 + 10) × (40 + 15) -(80 × 40) (O) An example with a lexically similar context to that of above from the UnbiasedMWP Context The school basketball court was [20] meters long and [12] meters wide.After the renovation, the length is increased by [8] meters, and the width increases by [3] meters.Question How many square meters are increased?Solution (20 + 8) × (12 + 3) -(20 × 12)
JB. Kim, Hazel Kim, Joonghyuk Hahn, Yo-Sub Han
EMNLP3
2023 M-equivalence of Parikh Matrix over a Ternary Alphabet
Joonghyuk Hahn, Hyunjoon Cheon, Yo-Sub Han
CIAA1
2022 Boosting Code Summarization by Embedding Code Structures
abstract
Recent research on code summarization relies on the structural information from the abstract syntax tree (AST) of source codes. It is, however, questionable whether it is the most effective to use AST for expressing the structural information. We find that a program dependency graph (PDG) can represent the structure of a code more effectively. We propose PDG Boosting Module (PBM) that encodes PDG into graph embedding and the framework to implement the proposed PBM with the existing models. PBM achieves improvements of 6.67% (BLEU) and 7.47% (ROUGE) on average. We then analyze the experimental results, and examine how PBM helps the training of baseline models and its performance robustness. For the validation of robustness, we measure the performance of an out-of-domain benchmark dataset, and confirm its robustness. In addition, we apply a new evaluation measure, SBERT score, to evaluate the semantic performance. The models implemented with PBM improve the performance of SBERT score. This implies that they generate summaries that are semantically more similar to the reference summary.
Jikyoeng Son, Joonghyuk Hahn, Yo-Sub Han
COLING2
2022 On the Decidability of Infix Inclusion Problem
Hyunjoon Cheon, Joonghyuk Hahn, Yo-Sub Han
DLT2
2021 Most Pseudo-copy Languages Are Not Context-Free
Hyunjoon Cheon, Joonghyuk Hahn, Yo-Sub Han, Sang-Ki Ko
COCOON2