Seongtae Hong

dblp:62/9368 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0002-2073-7731ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 61% Transfer learning and domain adaptation · 39%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
cross-language information retrieval
1.012026
CLEAR: Cross-Lingual Enhancement in Retrieval via Reverse-training · ACL (1) 2026
Information retrieval › ranking
learning to rank
1.012026
Beyond Hard Negatives: The Importance of Score Distribution in Knowledge Distillation · SIGIR 2026
Information retrieval
score distribution
1.012026
Beyond Hard Negatives: The Importance of Score Distribution in Knowledge Distillation · SIGIR 2026
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.912025
Cross-Lingual Optimization for Language Transfer in Large Language Models · ACL (1) 2025
Natural language and speech › Language models and text generation › instruction following
instruction-following benchmark
0.912025
Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models · EMNLP 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models · EMNLP 2025
Information retrieval › retrieval models › neural retrieval
neural ranking model
0.312026
Beyond Hard Negatives: The Importance of Score Distribution in Knowledge Distillation · SIGIR 2026

Methods — techniques the papers use, named apart from their topics

reverse-training · 2.0knowledge distillation · 1.0hard negative mining · 1.0translation model · 0.9supervised fine-tuning · 0.9benchmark construction · 0.9
YearPublicationVenuePosition
2026 CLEAR: Cross-Lingual Enhancement in Retrieval via Reverse-training
abstract
Seungyoon Lee, Minhyuk Kim, Seongtae Hong, Youngjoon Jang, Dongsuk Oh, Heuiseok Lim. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Seungyoon Lee, Minhyuk Kim, Seongtae Hong, Youngjoon Jang 0002, Dongsuk Oh, Heuiseok Lim
ACL (1)3
2026 Beyond Hard Negatives: The Importance of Score Distribution in Knowledge Distillation
Youngjoon Jang 0002, Seongtae Hong, Hyeonseok Moon, Heuiseok Lim
SIGIR2
2025 Cross-Lingual Optimization for Language Transfer in Large Language Models
abstract
Adapting large language models to other languages typically employs supervised finetuning (SFT) as a standard approach.However, it often suffers from an overemphasis on English performance, a phenomenon that is especially pronounced in data-constrained environments.To overcome these challenges, we propose Cross-Lingual Optimization (CLO) that efficiently transfers an English-centric LLM to a target language while preserving its English capabilities.CLO utilizes publicly available English SFT data and a translation model to enable cross-lingual transfer.We conduct experiments using five models on six languages, each possessing varying levels of resource.Our results show that CLO consistently outperforms SFT in both acquiring target language proficiency and maintaining English performance.Remarkably, in low-resource languages, CLO with only 3,200 samples surpasses SFT with 6,400 samples, demonstrating that CLO can achieve better performance with less data.Furthermore, we find that SFT is particularly sensitive to data quantity in medium and lowresource languages, whereas CLO remains robust.Our comprehensive analysis emphasizes the limitations of SFT and incorporates additional training strategies in CLO to enhance efficiency.
Jungseob Lee, Seongtae Hong, Hyeonseok Moon, Heuiseok Lim
ACL (1)2
2025 MIGRATE: Cross-Lingual Adaptation of Domain-Specific LLMs through Code-Switching and Embedding Transfer
abstract
Large Language Models (LLMs) have rapidly advanced, with domain-specific expert models emerging to handle specialized tasks across various fields. However, the predominant focus on English-centric models demands extensive data, making it challenging to develop comparable models for middle and low-resource languages. To address this limitation, we introduce Migrate, a novel method that leverages open-source static embedding models and up to 3 million tokens of code-switching data to facilitate the seamless transfer of embeddings to target languages. Migrate enables effective cross-lingual adaptation without requiring large-scale domain-specific corpora in the target language, promoting the accessibility of expert LLMs to a diverse range of linguistic communities. Our experimental results demonstrate that Migrate significantly enhances model performance in target languages, outperforming baseline and existing cross-lingual transfer methods. This approach provides a practical and efficient solution for extending the capabilities of domain-specific expert models.
Seongtae Hong, Seungyoon Lee, Hyeonseok Moon, Heuiseok Lim
COLING1
2025 Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
abstract
Recent frontier-level LLMs have saturated many previously difficult benchmarks, leaving little room for further differentiation.This progress highlights the need for challenging benchmarks that provide objective verification.In this paper, we introduce MCBench, a benchmark designed to evaluate whether LLMs can execute string-matching NLP metrics by strictly following step-by-step instructions.Unlike prior benchmarks that depend on subjective judgments or general reasoning, MCBench offers an objective, deterministic and codeverifiable evaluation.This setup allows us to systematically test whether LLMs can maintain accurate step-by-step execution, including instruction adherence, numerical computation, and long-range consistency in handling intermediate results.To ensure objective evaluation of these abilities, we provide a parallel reference code that can evaluate the accuracy of LLM output.We provide three evaluative metrics and three benchmark variants designed to measure the detailed instruction understanding capability of LLMs.Our analyses show that MCBench serves as an effective and objective tool for evaluating the capabilities of cuttingedge LLMs.
Hyeonseok Moon, Seongtae Hong, Jaehyung Seo, Heuiseok Lim
EMNLP2
2017 Testing Robustness of UTAUT Model: An Invariance Analysis
abstract
In order to compare a research model accurately across different conditions, the model's measures must be invariant across them. In this study, the invariance of the UTAUT model's measures was tested along three dimensions: country, technology, and gender. Data were collected from two countries (Korea and the U.S.) for two technologies (Internet banking and MP3 players). The results show that although the UTAUT model is robust overall across different conditions, possible differences due to measurement non-invariance should be taken into account. The paper discusses implications of the study results and makes recommendations for future research.
Myung Soo Kang, Il Im, Seongtae Hong
J. Glob. Inf. Manag.3
2011 An international comparison of technology adoption: Testing the UTAUT model
Il Im, Seongtae Hong, Myung Soo Kang
Inf. Manag.2