Seung-Hoon Na

dblp:56/3784 · DBLP profile ↗
← Back
47ranked-venue papers
20as first author
10since 2021 · last 2026
0000-0002-4372-7125ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 22 · 14 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
YearPublicationVenuePosition
2026 GateLM: Jointly injecting knowledge graphs and texts for reasoning-enhanced language models on commonsense question answering
Jinwoo Min, Kun-Hui Lee, Roseline Nyange, Seung-Hoon Na
Expert Syst. Appl.4
2026 MECA: Modular editing via customized expert networks and adaptors in large language models
Roseline Nyange, Shanbao Qiao, Seung-Hoon Na
Expert Syst. Appl.3
2026 OrthoEdit: Principled and Stable Knowledge Editing via Orthogonal Subspace Projection
abstract
Abstract Large language models (LLMs) encode extensive factual knowledge through pretraining, yet often require targeted updates to correct errors, incorporate new information, or revise outdated facts. Recent approaches to knowledge editing, such as projection-based constraints, parameter pruning, and regularization, have proven effective in improving editing accuracy and stability. However, these methods often fail to maintain a clear separation between new edits and existing knowledge, leading to interference and degradation over time. We propose OrthoEdit, a principled framework for stable and scalable knowledge editing that ensures each parameter update is orthogonal to both pre-existing and previously edited knowledge, remains strictly non-interfering and preserving the integrity of prior edits. OrthoEdit enables exact subspace control through three coordinated steps: progressive null space refinement, principal subspace extraction, and orthogonal projection. This yields compact and well-aligned updates that systematically satisfy all accumulated constraints. Comprehensive experiments across diverse models and benchmarks demonstrate that OrthoEdit consistently enhances editing accuracy and robustness while preserving general capabilities—even through extended sequences of batched edits. Our code is available at https://github.com/JoveReCode/OrthoEdit.git.
Shanbao Qiao, Xuebing Liu, Akshat Gupta, Seung-Hoon Na
Trans. Assoc. Comput. Linguistics4
2025 Wasserstein Distance Constraint and Parameter Sparsification for Batched and Iterative Knowledge Editing
abstract
Model knowledge editing has become a widely researched topic because it enables efficient and rapid injection of new knowledge into language models or the correction of erroneous or outdated knowledge. Existing model knowledge editing methods typically categorized into single-instance sequential editing and massive one-time editing. However, in practical applications, the batched and iterative editing manner better aligns with model updating patterns. In this work, we explored the performance of parameter-update-based models in a new batched iterative editing benchmark. Our findings show that with an increase in the number of editing iterations, the accumulation of updated parameters leads to a greater change in the distribution of model parameters, making it more challenging to maintain editing performance and model stability. To address this degradation issue, we propose two methods: the Wasserstein distance constraint and update parameter sparsification, where the Wasserstein distance constraint optimizes the transition of parameter distribution before and after the editing, and update parameter sparsification significantly reduces the number of update parameters, thereby alleviating the issue of instability in the parameter distribution caused by the accumulation of update parameters through iterations. Our methods can be generally applied to different parameter-update-based knowledge editing models. Experiments on the zsRE and CounterFact datasets demonstrate that our methods can improve editing performance and enhance the later-stage stability of batched iterative editing across different models.
Shanbao Qiao, Xuebing Liu, Seung-Hoon Na
AAAI3
2024 RADCoT: Retrieval-Augmented Distillation to Specialization Models for Generating Chain-of-Thoughts in Query Expansion
abstract
Large language models (LLMs) have demonstrated superior performance to that of small language models (SLM) in information retrieval for various subtasks including dense retrieval, reranking, query expansion, and pseudo-document generation. However, the parameter sizes of LLMs are extremely large, making it expensive to operate LLMs stably for providing LLM-based retrieval services. Recently, retrieval-augmented language models have been widely employed to significantly reduce the parameter size by retrieving relevant knowledge from large-scale corpora and exploiting the resulting “in-context” knowledge as additional model input, thereby substantially reducing the burden of internalizing and retaining world knowledge in model parameters. Armed by the retrieval-augmented language models, we present a retrieval-augmented model specialization that distills the capability of LLMs to generate the chain-of-thoughts (CoT) for query expansion – that is, injects the LLM’s capability to generate CoT into a retrieval-augmented SLM – referred to as RADCoT. Experimental results on the MS-MARCO, TREC DL 19, 20 datasets show that RADCoT yields consistent improvements over distillation without retrieval, achieving comparable performance to that of the query expansion method using LLM-based CoTs. Our code is publicly available at https://github.com/ZIZUN/RADCoT.
Eunhwan Park, Dong Hyeon Jeon, Inho Kang, Seung-Hoon Na
LREC/COLING5
2023 RINK: Reader-Inherited Evidence Reranker for Table-and-Text Open Domain Question Answering
abstract
Most approaches used in open-domain question answering on hybrid data that comprises both tabular-and-textual contents are based on a Retrieval-Reader pipeline in which the retrieval module finds relevant “heterogenous” evidence for a given question and the reader module generates an answer from the retrieved evidence. In this paper, we present a Retriever-Reranker-Reader framework by newly proposing a Reader-INherited evidence reranKer (RINK) where a reranker module is designed by finetuning the reader’s neural architecture based on a simple prompting method. Our underlying assumption of reusing the reader’s module for the reranker is that the reader’s ability to generating an answer from evidence contains the knowledge required for the reranking, because the reranker needs to “read” in-depth a question and evidences more carefully and elaborately than a baseline retriever. Furthermore, we present a simple and effective pretraining method by extensively deploying the commonly used data augmentation methods of cell corruption and cell reordering based on the pretraining tasks - tabular-and-textual entailment and cross-modal masked language modeling. Experimental results on OTT-QA, a large-scale table-and-text open-domain question answering dataset, show that the proposed RINK armed with our pretraining procedure makes improvements over the baseline reranking method and leads to state-of-the-art performance.
Eunhwan Park, Daeryong Seo, Seonhoon Kim, Inho Kang, Seung-Hoon Na
AAAI6
2023 ExplainMeetSum: A Dataset for Explainable Meeting Summarization Aligned with Human Intent
abstract
To enhance the explainability of meeting summarization, we construct a new dataset called "ExplainMeetSum," an augmented version of QMSum, by newly annotating evidence sentences that faithfully "explain" a summary.Using ExplainMeetSum, we propose a novel multiple extractor guided summarization, namely Multi-DYLE, which extensively generalizes DYLE to enable using a supervised extractor based on human-aligned extractive oracles.We further present an explainabilityaware task, named "Explainable Evidence Extraction" (E3), which aims to automatically detect all evidence sentences that support a given summary.Experimental results on the QMSum dataset show that the proposed Multi-DYLE outperforms DYLE with gains of up to 3.13 in the ROUGE-1 score.We further present the initial results on the E3 task, under the settings using separate and joint evaluation metrics. 1
Hyun Kim 0003, Minsoo Cho, Seung-Hoon Na
ACL (1)3
2023 Feature structure distillation with Centered Kernel Alignment in BERT transferring
Hee-Jun Jung, Seung-Hoon Na, Kangil Kim
Expert Syst. Appl.3
2022 SISER: Semantic-Infused Selective Graph Reasoning for Fact Verification
abstract
This study proposes Semantic-Infused SElective Graph Reasoning (SISER) for fact verification, which newly presents semantic-level graph reasoning and injects its reasoning-enhanced representation into other types of graph-based and sequence-based reasoning methods. SISER combines three reasoning types: 1) semantic-level graph reasoning, which uses a semantic graph from evidence sentences, whose nodes are elements of a triple – <Subject, Verb, Object>, 2) “semantic-infused” sentence-level “selective” graph reasoning, which combine semantic-level and sentence-level representations and perform graph reasoning in a selective manner using the node selection mechanism, and 3) sequence reasoning, which concatenates all evidence sentences and performs attention-based reasoning. Experiment results on a large-scale dataset for Fact Extraction and VERification (FEVER) show that SISER outperforms the previous graph-based approaches and achieves state-of-the-art performance.
Eunhwan Park, Jong-Hyeon Lee, Dong Hyeon Jeon, Seonhoon Kim, Inho Kang, Seung-Hoon Na
COLING6
2022 Frustratingly Easy System Combination for Grammatical Error Correction
abstract
In this paper, we formulate system combination for grammatical error correction (GEC) as a simple machine learning task: binary classification.We demonstrate that with the right problem formulation, a simple logistic regression algorithm can be highly effective for combining GEC models.Our method successfully increases the F 0.5 score from the highest base GEC system by 4.2 points on the CoNLL-2014 test set and 7.2 points on the BEA-2019 test set.Furthermore, our method outperforms the state of the art by 4.0 points on the BEA-2019 test set, 1.2 points on the CoNLL-2014 test set with original annotation, and 3.4 points on the CoNLL-2014 test set with alternative annotation.We also show that our system combination generates better corrections with higher F 0.5 scores than the conventional ensemble.1
Muhammad Reza Qorib, Seung-Hoon Na, Hwee Tou Ng
NAACL-HLT2
2020 Word Reordering for Translation into Korean Sign Language Using Syntactically-guided Classification
abstract
Machine translation aims to break the language barrier that prevents communication with others and increase access to information. Deaf people face huge language barriers in their daily lives, including access to digital and spoken information. There are very few digital resources for sign language processing. In this article, we present a transfer-based machine translation system for translating Korean-to-Korean Sign Language (KSL) glosses, mainly composed of (1) dictionary-based lexical transfer and (2) a hybrid syntactic transfer based on a data-driven model. In particular, we formulate complicated word reordering problems in syntactic transfer as multi-class classification tasks and propose “syntactically guided” data-driven syntactic transfer. The core part of our study is a neural classification model for reordering order-important constituent pairs with a reordering task that is newly designed for Korean-to-KSL translation. The experiment results evaluated on news transcript data show that the proposed system achieves a BLEU score of 0.512 and a RIBES score of 0.425, significantly improving upon the baseline system performance.
Hun-Young Jung, Jong-Hyeok Lee, Eunju Min, Seung-Hoon Na
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2020 Uniformly Interpolated Balancing for Robust Prediction in Translation Quality Estimation: A Case Study of English-Korean Translation
abstract
There has been growing interest among researchers in quality estimation (QE), which attempts to automatically predict the quality of machine translation (MT) outputs. Most existing works on QE are based on supervised approaches using quality-annotated training data. However, QE training data quality scores readily become imbalanced or skewed : QE data are mostly composed of high translation quality sentence pairs but the data lack low translation quality sentence pairs. The use of imbalanced data with an induced quality estimator tends to produce biased translation quality scores with “high” translation quality scores assigned even to poorly translated sentences. To address the data imbalance, this article proposes a simple, efficient procedure called uniformly interpolated balancing to construct more balanced QE training data by inserting greater uniformness to training data. The proposed uniformly interpolated balancing procedure is based on the preparation of two different types of manually annotated QE data: (1) default skewed data and (2) near-uniform data . First, we obtain default skewed data in a naive manner without considering the imbalance by manually annotating qualities on MT outputs. Second, we obtain near-uniform data in a selective manner by manually annotating a subset only, which is selected from the automatically quality-estimated sentence pairs. Finally, we create uniformly interpolated balanced data by combining these two types of data, where one half originates from the default skewed data and the other half originates from the near-uniform data. We expect that uniformly interpolated balancing reflects the intrinsic skewness of the true quality distribution and manages the imbalance problem. Experimental results on an English-Korean quality estimation task show that the proposed uniformly interpolated balancing leads to robustness on both skewed and uniformly distributed quality test sets when compared to the test sets of other non-balanced datasets.
Hyun Kim 0003, Seung-Hoon Na
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2019 Improving LSTM CRFs using character-based compositions for Korean named entity recognition
Seung-Hoon Na, Hyun Kim 0003, Jinwoo Min, Kangil Kim
Comput. Speech Lang.1
2019 Multi-task Stack Propagation for Neural Quality Estimation
abstract
Quality estimation is an important task in machine translation that has attracted increased interest in recent years. A key problem in translation-quality estimation is the lack of a sufficient amount of the quality annotated training data. To address this shortcoming, the Predictor-Estimator was proposed recently by introducing “word prediction” as an additional pre-subtask that predicts a current target word with consideration of surrounding source and target contexts, resulting in a two-stage neural model composed of a predictor and an estimator . However, the original Predictor-Estimator is not trained on a continuous stacking model but instead in a cascaded manner that separately trains the predictor from the estimator. In addition, the Predictor-Estimator is trained based on single-task learning only, which uses target-specific quality-estimation data without using other training data that are available from other-level quality-estimation tasks. In this article, we thus propose a multi-task stack propagation , which extensively applies stack propagation to fully train the Predictor-Estimator on a continuous stacking architecture and multi-task learning to enhance the training data from related other-level quality-estimation tasks. Experimental results on WMT17 quality-estimation datasets show that the Predictor-Estimator trained with multi-task stack propagation provides statistically significant improvements over the baseline models. In particular, under an ensemble setting, the proposed multi-task stack propagation leads to state-of-the-art performance at all the sentence/word/phrase levels for WMT17 quality estimation tasks.
Hyun Kim 0003, Jong-Hyeok Lee, Seung-Hoon Na
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2019 Transition-Based Korean Dependency Parsing Using Hybrid Word Representations of Syllables and Morphemes with LSTMs
abstract
Recently, neural approaches for transition-based dependency parsing have become one of the state-of-the art methods for performing dependency parsing tasks in many languages. In neural transition-based parsing, a parser state representation is first computed from the configuration of a stack and a buffer, which is then fed into a feed-forward neural network model that predicts the next transition action. Given that words are basic elements of a stack and buffer, a parser state representation is considerably affected by how a word representation is defined. In particular, word representation issues become more critical in morphologically rich languages such as Korean, as the set of potential words is not bound but introduce the second-order vocabulary complexity, called the phrase vocabulary complexity due to the agglutinative characteristics of the language. In this article, we propose a hybrid word representation that combines two compositional word representations, each of which is derived from representations of syllables and morphemes , respectively. Our underlying assumption for this hybrid word representation is that, because both syllables and morphemes are two common ways of decomposing Korean words, it is expected that their effects in inducing word representation are complementary to one another. Experimental results carried on Sejong and SPMRL 2014 datasets show that our proposed hybrid word representation leads to the state-of-the-art performance.
Seung-Hoon Na, Jianri Li, Jong-Hun Shin, Kangil Kim
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2018 Verbosity normalized pseudo-relevance feedback in information retrieval
Seung-Hoon Na, Kangil Kim
Inf. Process. Manag.1
2017 Center-shared sliding ensemble of neural networks for syntax analysis of natural language
Kangil Kim, Yun Jin, Seung-Hoon Na, Young Kil Kim
Expert Syst. Appl.3
2017 Predictor-Estimator: Neural Quality Estimation Based on Target Word Prediction for Machine Translation
abstract
Recently, quality estimation has been attracting increasing interest from machine translation researchers, aiming at finding a good estimator for the “quality” of machine translation output. The common approach for quality estimation is to treat the problem as a supervised regression/classification task using a quality-annotated noisy parallel corpus, called quality estimation data , as training data. However, the available size of quality estimation data remains small, due to the too-expensive cost of creating such data. In addition, most conventional quality estimation approaches rely on manually designed features to model nonlinear relationships between feature vectors and corresponding quality labels. To overcome these problems, this article proposes a novel neural network architecture for quality estimation task—called the predictor-estimator —that considers word prediction as an additional pre-task. The major component of the proposed neural architecture is a word prediction model based on a modified neural machine translation model—a probabilistic model for predicting a target word conditioned on all the other source and target contexts. The underlying assumption is that the word prediction model is highly related to quality estimation models and is therefore able to transfer useful knowledge to quality estimation tasks. Our proposed quality estimation method sequentially trains the following two types of neural models: (1) Predictor : a neural word prediction model trained from parallel corpora and (2) Estimator : a neural quality estimation model trained from quality estimation data. To transfer word a prediction task to a quality estimation task, we generate quality estimation feature vectors from the word prediction model and feed them into the quality estimation model. The experimental results on WMT15 and 16 quality estimation datasets show that our proposed method has great potential in the various sub-challenges.
Hyun Kim 0003, Hun-Young Jung, Hong-Seok Kwon, Jong-Hyeok Lee, Seung-Hoon Na
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2015 Conditional Random Fields for Korean Morpheme Segmentation and POS Tagging
abstract
There has been recent interest in statistical approaches to Korean morphological analysis. However, previous studies have been based mostly on generative models, including a hidden Markov model (HMM), without utilizing discriminative models such as a conditional random field (CRF). We present a two-stage discriminative approach based on CRFs for Korean morphological analysis. Similar to methods used for Chinese, we perform two disambiguation procedures based on CRFs: (1) morpheme segmentation and (2) POS tagging. In morpheme segmentation, an input sentence is segmented into sequences of morphemes, where a morpheme unit is either atomic or compound. In the POS tagging procedure, each morpheme (atomic or compound) is assigned a POS tag. Once POS tagging is complete, we carry out a post-processing of the compound morphemes, where each compound morpheme is further decomposed into atomic morphemes, which is based on pre-analyzed patterns and generalized HMMs obtained from the given tagged corpus. Experimental results show the promise of our proposed method.
Seung-Hoon Na
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2015 Two-Stage Document Length Normalization for Information Retrieval
abstract
The standard approach for term frequency normalization is based only on the document length. However, it does not distinguish the verbosity from the scope, these being the two main factors determining the document length. Because the verbosity and scope have largely different effects on the increase in term frequency, the standard approach can easily suffer from insufficient or excessive penalization depending on the specific type of long document. To overcome these problems, this article proposes two-stage normalization by performing verbosity and scope normalization separately, and by employing different penalization functions. In verbosity normalization, each document is prenormalized by dividing the term frequency by the verbosity of the document. In scope normalization, an existing retrieval model is applied in a straightforward manner to the prenormalized document, finally leading us to formulate our proposed verbosity normalized (VN) retrieval model. Experimental results carried out on standard TREC collections demonstrate that the VN model leads to marginal but statistically significant improvements over standard retrieval models.
Seung-Hoon Na
ACM Trans. Inf. Syst.1
2014 Partial-update dimensionality reduction for accumulating co-occurrence events
Seung-Hoon Na, Jong-Hyeok Lee
Pattern Recognit. Lett.1
2013 Refining sentence similarity with discourse information in dialog system
Sangkeun Jung, Seung-Hoon Na
INTERSPEECH2
2013 Probabilistic co-relevance for query-sensitive similarity measurement in information retrieval
Seung-Hoon Na
Inf. Process. Manag.1
2012 Utilizing local evidence for blog feed search
Yeha Lee, Seung-Hoon Na, Jong-Hyeok Lee
Inf. Retr.2
2012 Memory-restricted latent semantic analysis to accumulate term-document co-occurrence events
Seung-Hoon Na, Jong-Hyeok Lee
Pattern Recognit. Lett.1
2011 Enriching document representation via translation for improved monolingual information retrieval
abstract
Word ambiguity and vocabulary mismatch are critical problems in information retrieval. To deal with these problems, this paper proposes the use of translated words to enrich document representation, going beyond the words in the original source language to represent a document. In our approach, each original document is automatically translated into an auxiliary language, and the resulting translated document serves as a semantically enhanced representation for supplementing the original bag of words. The core of our translation representation is the expected term frequency of a word in a translated document, which is calculated by averaging the term frequencies over all possible translations, rather than focusing on the 1-best translation only. To achieve better efficiency of translation, we do not rely on full-fledged machine translation, but instead use monotonic translation by removing the time-consuming reordering component. Experiments carried out on standard TREC test collections show that our proposed translation representation leads to statistically significant improvements over using only the original language of the document collection.
Seung-Hoon Na, Hwee Tou Ng
SIGIR1
2010 RankSVR: can preference data help regression?
abstract
In some regression applications (e.g., an automatic movie scoring system), a large number of ranking data is available in addition to the original regression data. This paper studies whether and how the ranking data can improve the accuracy of regression task. In particular, this paper first proposes an extension of SVR (Support Vector Regression), RankSVR, which incorporates ranking constraints in the learning of regression function. Second, this paper proposes novel sampling methods for RankSVR, which selectively choose samples of ranking data for training of regression functions in order to maximize the performance of RankSVR. While it is relatively easier to acquire ranking data than regression data, incorporating all the ranking data in the learning of regression doest not always generate the best output. Moreoever, adding too many ranking constraints into the regression problem substantially lengthens the training time. Our proposed sampling methods find the ranking samples that maximize the regression performance. Experimental results on synthetic and real data sets show that, when the ranking data is additionally available, RankSVR significantly performs better than SVR by utilizing ranking constraints in the learning of regression, and also show that our sampling methods improve the RankSVR performance better than the random sampling.
Hwanjo Yu, Sungchul Kim, Seung-Hoon Na
CIKM3
2009 An improved feedback approach using relevant local posts for blog feed retrieval
abstract
Blog feed search aims to identify a blog feed with a recurring interest in a given topic. In this paper, we investigate the "pseudo-relevance feedback" for blog feed search task, where its unit of relevance judgment is not based on a blog post but a blog feed (the collection of all its constituent posts). This paper focuses on two characteristics of feed search task, blog feed's topical diversity and multifaceted property of query. We propose a novel feed-level selection of local posts which uses only highly relevant local posts in each top-ranked feed, in order to capture the correct and diverse relevant information to a given topic. Experimental results show that the proposed approach outperforms traditional feedback approaches. Especially, the proposed approach gives 2% further increase of nDCG over the best performing result of TREC '08 Blog Distillation Task.
Yeha Lee, Seung-Hoon Na, Jong-Hyeok Lee
CIKM2
2009 Improving Opinion Retrieval Based on Query-Specific Sentiment Lexicon
Seung-Hoon Na, Yeha Lee, Sang-Hyob Nam, Jong-Hyeok Lee
ECIR1
2009 DiffPost: Filtering Non-relevant Content Based on Content Difference between Two Consecutive Blog Posts
Sang-Hyob Nam, Seung-Hoon Na, Yeha Lee, Jong-Hyeok Lee
ECIR2
2009 A 2-poisson model for probabilistic coreference of named entities for improved text retrieval
abstract
Text retrieval queries frequently contain named entities. The standard approach of term frequency weighting does not work well when estimating the term frequency of a named entity, since anaphoric expressions (like he, she, the movie, etc) are frequently used to refer to named entities in a document, and the use of anaphoric expressions causes the term frequency of named entities to be underestimated. In this paper, we propose a novel 2-Poisson model to estimate the frequency of anaphoric expressions of a named entity, without explicitly resolving the anaphoric expressions. Our key assumption is that the frequency of anaphoric expressions is distributed over named entities in a document according to the probabilities of whether the document is elite for the named entities. This assumption leads us to formulate our proposed Co-referentially Enhanced Entity Frequency (CEEF). Experimental results on the text collection of TREC Blog Track show that CEEF achieves significant and consistent improvements over state-of-the-art retrieval methods using standard term frequency estimation. In particular, we achieve a 3% increase of MAP over the best performing run of TREC 2008 Blog Track.
Seung-Hoon Na, Hwee Tou Ng
SIGIR1
2009 On co-authorship for author disambiguation
In-Su Kang, Seung-Hoon Na, Hanmin Jung, Pyung Kim, Won-Kyung Sung, Jong-Hyeok Lee
Inf. Process. Manag.2
2008 Improving Term Frequency Normalization for Multi-topical Documents and Application to Language Modeling Approaches
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee
ECIR1
2008 Structural Re-ranking with Cluster-Based Retrieval
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee
ECIR1
2008 Revisit of Nearest Neighbor Test for Direct Evaluation of Inter-document Similarities
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee
ECIR1
2008 Query-Based Inter-document Similarity Using Probabilistic Co-relevance Model
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee
ECIR1
2008 Automatic Extraction of English-Chinese Transliteration Pairs using Dynamic Window and Tokenizer
Chengguo Jin, Seung-Hoon Na, Dong-Il Kim, Jong-Hyeok Lee
IJCNLP2
2008 Search Result Clustering Using Label Language Model
Yeha Lee, Seung-Hoon Na, Jong-Hyeok Lee
IJCNLP2
2008 Exploiting proximity feature in bigram language model for information retrieval
abstract
Language modeling approaches have been effectively dealing with the dependency among query terms based on N-gram such as bigram or trigram models. However, bigram language models suffer from adjacency-sparseness problem which means that dependent terms are not always adjacent in documents, but can be far from each other, sometimes with distance of a few sentences in a document. To resolve the adjacency-sparseness problem, this paper proposes a new type of bigram language model by explicitly incorporating the proximity feature between two adjacent terms in a query. Experimental results on three test collections show that the proposed bigram language model significantly improves previous bigram model as well as Tao's approach, the state-of-art method for proximity-based method.
Seung-Hoon Na, Jungi Kim, In-Su Kang, Jong-Hyeok Lee
SIGIR1
2007 Cluster-based patent retrieval
In-Su Kang, Seung-Hoon Na, Jungi Kim, Jong-Hyeok Lee
Inf. Process. Manag.2
2007 Parsimonious translation models for information retrieval
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee
Inf. Process. Manag.1
2007 Adaptive document clustering based on query-based similarity
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee
Inf. Process. Manag.1
2007 An empirical study of query expansion and cluster-based retrieval in language modeling approach
Seung-Hoon Na, In-Su Kang, Ji-Eun Roh, Jong-Hyeok Lee
Inf. Process. Manag.1
2006 Collection-based compound noun segmentation for Korean information retrieval
In-Su Kang, Seung-Hoon Na, Jong-Hyeok Lee
Inf. Retr.2
2004 Influence of WSD on Cross-Language Information Retrieval
In-Su Kang, Seung-Hoon Na, Jong-Hyeok Lee
IJCNLP2
2004 Improving Relevance Feedback in Language Modeling Approach: Maximum a Posteriori Probability Criterion and Three-Component Mixture Model
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee
IJCNLP1
2004 Lightweight Natural Language Database Interfaces
In-Su Kang, Seung-Hoon Na, Jong-Hyeok Lee, Gijoo Yang
NLDB2