VLDB 2026 Research / reviewers in the wild / expert
Seung-Hoon Na
dblp:56/3784
· DBLP profile ↗
47ranked-venue papers
20as first author
10since 2021 · last 2026
0000-0002-4372-7125ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 22 · 14 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GateLM: Jointly injecting knowledge graphs and texts for reasoning-enhanced language models on commonsense question answering
Jinwoo Min, Kun-Hui Lee, Roseline Nyange, Seung-Hoon Na |
Expert Syst. Appl. | 4 |
| 2026 | MECA: Modular editing via customized expert networks and adaptors in large language models
Roseline Nyange, Shanbao Qiao, Seung-Hoon Na |
Expert Syst. Appl. | 3 |
| 2026 | OrthoEdit: Principled and Stable Knowledge Editing via Orthogonal Subspace ProjectionabstractAbstract Large language models (LLMs) encode extensive factual knowledge through pretraining, yet often require targeted updates to correct errors, incorporate new information, or revise outdated facts. Recent approaches to knowledge editing, such as projection-based constraints, parameter pruning, and regularization, have proven effective in improving editing accuracy and stability. However, these methods often fail to maintain a clear separation between new edits and existing knowledge, leading to interference and degradation over time. We propose OrthoEdit, a principled framework for stable and scalable knowledge editing that ensures each parameter update is orthogonal to both pre-existing and previously edited knowledge, remains strictly non-interfering and preserving the integrity of prior edits. OrthoEdit enables exact subspace control through three coordinated steps: progressive null space refinement, principal subspace extraction, and orthogonal projection. This yields compact and well-aligned updates that systematically satisfy all accumulated constraints. Comprehensive experiments across diverse models and benchmarks demonstrate that OrthoEdit consistently enhances editing accuracy and robustness while preserving general capabilities—even through extended sequences of batched edits. Our code is available at https://github.com/JoveReCode/OrthoEdit.git. Shanbao Qiao, Xuebing Liu, Akshat Gupta, Seung-Hoon Na |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Wasserstein Distance Constraint and Parameter Sparsification for Batched and Iterative Knowledge EditingabstractModel knowledge editing has become a widely researched topic because it enables efficient and rapid injection of new knowledge into language models or the correction of erroneous or outdated knowledge. Existing model knowledge editing methods typically categorized into single-instance sequential editing and massive one-time editing. However, in practical applications, the batched and iterative editing manner better aligns with model updating patterns. In this work, we explored the performance of parameter-update-based models in a new batched iterative editing benchmark. Our findings show that with an increase in the number of editing iterations, the accumulation of updated parameters leads to a greater change in the distribution of model parameters, making it more challenging to maintain editing performance and model stability. To address this degradation issue, we propose two methods: the Wasserstein distance constraint and update parameter sparsification, where the Wasserstein distance constraint optimizes the transition of parameter distribution before and after the editing, and update parameter sparsification significantly reduces the number of update parameters, thereby alleviating the issue of instability in the parameter distribution caused by the accumulation of update parameters through iterations. Our methods can be generally applied to different parameter-update-based knowledge editing models. Experiments on the zsRE and CounterFact datasets demonstrate that our methods can improve editing performance and enhance the later-stage stability of batched iterative editing across different models. Shanbao Qiao, Xuebing Liu, Seung-Hoon Na |
AAAI | 3 |
| 2024 | RADCoT: Retrieval-Augmented Distillation to Specialization Models for Generating Chain-of-Thoughts in Query ExpansionabstractLarge language models (LLMs) have demonstrated superior performance to that of small language models (SLM) in information retrieval for various subtasks including dense retrieval, reranking, query expansion, and pseudo-document generation. However, the parameter sizes of LLMs are extremely large, making it expensive to operate LLMs stably for providing LLM-based retrieval services. Recently, retrieval-augmented language models have been widely employed to significantly reduce the parameter size by retrieving relevant knowledge from large-scale corpora and exploiting the resulting “in-context” knowledge as additional model input, thereby substantially reducing the burden of internalizing and retaining world knowledge in model parameters. Armed by the retrieval-augmented language models, we present a retrieval-augmented model specialization that distills the capability of LLMs to generate the chain-of-thoughts (CoT) for query expansion – that is, injects the LLM’s capability to generate CoT into a retrieval-augmented SLM – referred to as RADCoT. Experimental results on the MS-MARCO, TREC DL 19, 20 datasets show that RADCoT yields consistent improvements over distillation without retrieval, achieving comparable performance to that of the query expansion method using LLM-based CoTs. Our code is publicly available at https://github.com/ZIZUN/RADCoT. Eunhwan Park, Dong Hyeon Jeon, Inho Kang, Seung-Hoon Na |
LREC/COLING | 5 |
| 2023 | RINK: Reader-Inherited Evidence Reranker for Table-and-Text Open Domain Question AnsweringabstractMost approaches used in open-domain question answering on hybrid data that comprises both tabular-and-textual contents are based on a Retrieval-Reader pipeline in which the retrieval module finds relevant “heterogenous” evidence for a given question and the reader module generates an answer from the retrieved evidence. In this paper, we present a Retriever-Reranker-Reader framework by newly proposing a Reader-INherited evidence reranKer (RINK) where a reranker module is designed by finetuning the reader’s neural architecture based on a simple prompting method. Our underlying assumption of reusing the reader’s module for the reranker is that the reader’s ability to generating an answer from evidence contains the knowledge required for the reranking, because the reranker needs to “read” in-depth a question and evidences more carefully and elaborately than a baseline retriever. Furthermore, we present a simple and effective pretraining method by extensively deploying the commonly used data augmentation methods of cell corruption and cell reordering based on the pretraining tasks - tabular-and-textual entailment and cross-modal masked language modeling. Experimental results on OTT-QA, a large-scale table-and-text open-domain question answering dataset, show that the proposed RINK armed with our pretraining procedure makes improvements over the baseline reranking method and leads to state-of-the-art performance. Eunhwan Park, Daeryong Seo, Seonhoon Kim, Inho Kang, Seung-Hoon Na |
AAAI | 6 |
| 2023 | ExplainMeetSum: A Dataset for Explainable Meeting Summarization Aligned with Human IntentabstractTo enhance the explainability of meeting summarization, we construct a new dataset called "ExplainMeetSum," an augmented version of QMSum, by newly annotating evidence sentences that faithfully "explain" a summary.Using ExplainMeetSum, we propose a novel multiple extractor guided summarization, namely Multi-DYLE, which extensively generalizes DYLE to enable using a supervised extractor based on human-aligned extractive oracles.We further present an explainabilityaware task, named "Explainable Evidence Extraction" (E3), which aims to automatically detect all evidence sentences that support a given summary.Experimental results on the QMSum dataset show that the proposed Multi-DYLE outperforms DYLE with gains of up to 3.13 in the ROUGE-1 score.We further present the initial results on the E3 task, under the settings using separate and joint evaluation metrics. 1 Hyun Kim 0003, Minsoo Cho, Seung-Hoon Na |
ACL (1) | 3 |
| 2023 | Feature structure distillation with Centered Kernel Alignment in BERT transferring
Hee-Jun Jung, Seung-Hoon Na, Kangil Kim |
Expert Syst. Appl. | 3 |
| 2022 | SISER: Semantic-Infused Selective Graph Reasoning for Fact VerificationabstractThis study proposes Semantic-Infused SElective Graph Reasoning (SISER) for fact verification, which newly presents semantic-level graph reasoning and injects its reasoning-enhanced representation into other types of graph-based and sequence-based reasoning methods. SISER combines three reasoning types: 1) semantic-level graph reasoning, which uses a semantic graph from evidence sentences, whose nodes are elements of a triple – <Subject, Verb, Object>, 2) “semantic-infused” sentence-level “selective” graph reasoning, which combine semantic-level and sentence-level representations and perform graph reasoning in a selective manner using the node selection mechanism, and 3) sequence reasoning, which concatenates all evidence sentences and performs attention-based reasoning. Experiment results on a large-scale dataset for Fact Extraction and VERification (FEVER) show that SISER outperforms the previous graph-based approaches and achieves state-of-the-art performance. Eunhwan Park, Jong-Hyeon Lee, Dong Hyeon Jeon, Seonhoon Kim, Inho Kang, Seung-Hoon Na |
COLING | 6 |
| 2022 | Frustratingly Easy System Combination for Grammatical Error CorrectionabstractIn this paper, we formulate system combination for grammatical error correction (GEC) as a simple machine learning task: binary classification.We demonstrate that with the right problem formulation, a simple logistic regression algorithm can be highly effective for combining GEC models.Our method successfully increases the F 0.5 score from the highest base GEC system by 4.2 points on the CoNLL-2014 test set and 7.2 points on the BEA-2019 test set.Furthermore, our method outperforms the state of the art by 4.0 points on the BEA-2019 test set, 1.2 points on the CoNLL-2014 test set with original annotation, and 3.4 points on the CoNLL-2014 test set with alternative annotation.We also show that our system combination generates better corrections with higher F 0.5 scores than the conventional ensemble.1 Muhammad Reza Qorib, Seung-Hoon Na, Hwee Tou Ng |
NAACL-HLT | 2 |
| 2020 | Word Reordering for Translation into Korean Sign Language Using Syntactically-guided ClassificationabstractMachine translation aims to break the language barrier that prevents communication with others and increase access to information. Deaf people face huge language barriers in their daily lives, including access to digital and spoken information. There are very few digital resources for sign language processing. In this article, we present a transfer-based machine translation system for translating Korean-to-Korean Sign Language (KSL) glosses, mainly composed of (1) dictionary-based lexical transfer and (2) a hybrid syntactic transfer based on a data-driven model. In particular, we formulate complicated word reordering problems in syntactic transfer as multi-class classification tasks and propose “syntactically guided” data-driven syntactic transfer. The core part of our study is a neural classification model for reordering order-important constituent pairs with a reordering task that is newly designed for Korean-to-KSL translation. The experiment results evaluated on news transcript data show that the proposed system achieves a BLEU score of 0.512 and a RIBES score of 0.425, significantly improving upon the baseline system performance. Hun-Young Jung, Jong-Hyeok Lee, Eunju Min, Seung-Hoon Na |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2020 | Uniformly Interpolated Balancing for Robust Prediction in Translation Quality Estimation: A Case Study of English-Korean TranslationabstractThere has been growing interest among researchers in quality estimation (QE), which attempts to automatically predict the quality of machine translation (MT) outputs. Most existing works on QE are based on supervised approaches using quality-annotated training data. However, QE training data quality scores readily become imbalanced or skewed : QE data are mostly composed of high translation quality sentence pairs but the data lack low translation quality sentence pairs. The use of imbalanced data with an induced quality estimator tends to produce biased translation quality scores with “high” translation quality scores assigned even to poorly translated sentences. To address the data imbalance, this article proposes a simple, efficient procedure called uniformly interpolated balancing to construct more balanced QE training data by inserting greater uniformness to training data. The proposed uniformly interpolated balancing procedure is based on the preparation of two different types of manually annotated QE data: (1) default skewed data and (2) near-uniform data . First, we obtain default skewed data in a naive manner without considering the imbalance by manually annotating qualities on MT outputs. Second, we obtain near-uniform data in a selective manner by manually annotating a subset only, which is selected from the automatically quality-estimated sentence pairs. Finally, we create uniformly interpolated balanced data by combining these two types of data, where one half originates from the default skewed data and the other half originates from the near-uniform data. We expect that uniformly interpolated balancing reflects the intrinsic skewness of the true quality distribution and manages the imbalance problem. Experimental results on an English-Korean quality estimation task show that the proposed uniformly interpolated balancing leads to robustness on both skewed and uniformly distributed quality test sets when compared to the test sets of other non-balanced datasets. Hyun Kim 0003, Seung-Hoon Na |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2019 | Improving LSTM CRFs using character-based compositions for Korean named entity recognition
Seung-Hoon Na, Hyun Kim 0003, Jinwoo Min, Kangil Kim |
Comput. Speech Lang. | 1 |
| 2019 | Multi-task Stack Propagation for Neural Quality EstimationabstractQuality estimation is an important task in machine translation that has attracted increased interest in recent years. A key problem in translation-quality estimation is the lack of a sufficient amount of the quality annotated training data. To address this shortcoming, the Predictor-Estimator was proposed recently by introducing “word prediction” as an additional pre-subtask that predicts a current target word with consideration of surrounding source and target contexts, resulting in a two-stage neural model composed of a predictor and an estimator . However, the original Predictor-Estimator is not trained on a continuous stacking model but instead in a cascaded manner that separately trains the predictor from the estimator. In addition, the Predictor-Estimator is trained based on single-task learning only, which uses target-specific quality-estimation data without using other training data that are available from other-level quality-estimation tasks. In this article, we thus propose a multi-task stack propagation , which extensively applies stack propagation to fully train the Predictor-Estimator on a continuous stacking architecture and multi-task learning to enhance the training data from related other-level quality-estimation tasks. Experimental results on WMT17 quality-estimation datasets show that the Predictor-Estimator trained with multi-task stack propagation provides statistically significant improvements over the baseline models. In particular, under an ensemble setting, the proposed multi-task stack propagation leads to state-of-the-art performance at all the sentence/word/phrase levels for WMT17 quality estimation tasks. Hyun Kim 0003, Jong-Hyeok Lee, Seung-Hoon Na |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2019 | Transition-Based Korean Dependency Parsing Using Hybrid Word Representations of Syllables and Morphemes with LSTMsabstractRecently, neural approaches for transition-based dependency parsing have become one of the state-of-the art methods for performing dependency parsing tasks in many languages. In neural transition-based parsing, a parser state representation is first computed from the configuration of a stack and a buffer, which is then fed into a feed-forward neural network model that predicts the next transition action. Given that words are basic elements of a stack and buffer, a parser state representation is considerably affected by how a word representation is defined. In particular, word representation issues become more critical in morphologically rich languages such as Korean, as the set of potential words is not bound but introduce the second-order vocabulary complexity, called the phrase vocabulary complexity due to the agglutinative characteristics of the language. In this article, we propose a hybrid word representation that combines two compositional word representations, each of which is derived from representations of syllables and morphemes , respectively. Our underlying assumption for this hybrid word representation is that, because both syllables and morphemes are two common ways of decomposing Korean words, it is expected that their effects in inducing word representation are complementary to one another. Experimental results carried on Sejong and SPMRL 2014 datasets show that our proposed hybrid word representation leads to the state-of-the-art performance. Seung-Hoon Na, Jianri Li, Jong-Hun Shin, Kangil Kim |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2018 | Verbosity normalized pseudo-relevance feedback in information retrieval
Seung-Hoon Na, Kangil Kim |
Inf. Process. Manag. | 1 |
| 2017 | Center-shared sliding ensemble of neural networks for syntax analysis of natural language
Kangil Kim, Yun Jin, Seung-Hoon Na, Young Kil Kim |
Expert Syst. Appl. | 3 |
| 2017 | Predictor-Estimator: Neural Quality Estimation Based on Target Word Prediction for Machine TranslationabstractRecently, quality estimation has been attracting increasing interest from machine translation researchers, aiming at finding a good estimator for the “quality” of machine translation output. The common approach for quality estimation is to treat the problem as a supervised regression/classification task using a quality-annotated noisy parallel corpus, called quality estimation data , as training data. However, the available size of quality estimation data remains small, due to the too-expensive cost of creating such data. In addition, most conventional quality estimation approaches rely on manually designed features to model nonlinear relationships between feature vectors and corresponding quality labels. To overcome these problems, this article proposes a novel neural network architecture for quality estimation task—called the predictor-estimator —that considers word prediction as an additional pre-task. The major component of the proposed neural architecture is a word prediction model based on a modified neural machine translation model—a probabilistic model for predicting a target word conditioned on all the other source and target contexts. The underlying assumption is that the word prediction model is highly related to quality estimation models and is therefore able to transfer useful knowledge to quality estimation tasks. Our proposed quality estimation method sequentially trains the following two types of neural models: (1) Predictor : a neural word prediction model trained from parallel corpora and (2) Estimator : a neural quality estimation model trained from quality estimation data. To transfer word a prediction task to a quality estimation task, we generate quality estimation feature vectors from the word prediction model and feed them into the quality estimation model. The experimental results on WMT15 and 16 quality estimation datasets show that our proposed method has great potential in the various sub-challenges. Hyun Kim 0003, Hun-Young Jung, Hong-Seok Kwon, Jong-Hyeok Lee, Seung-Hoon Na |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2015 | Conditional Random Fields for Korean Morpheme Segmentation and POS TaggingabstractThere has been recent interest in statistical approaches to Korean morphological analysis. However, previous studies have been based mostly on generative models, including a hidden Markov model (HMM), without utilizing discriminative models such as a conditional random field (CRF). We present a two-stage discriminative approach based on CRFs for Korean morphological analysis. Similar to methods used for Chinese, we perform two disambiguation procedures based on CRFs: (1) morpheme segmentation and (2) POS tagging. In morpheme segmentation, an input sentence is segmented into sequences of morphemes, where a morpheme unit is either atomic or compound. In the POS tagging procedure, each morpheme (atomic or compound) is assigned a POS tag. Once POS tagging is complete, we carry out a post-processing of the compound morphemes, where each compound morpheme is further decomposed into atomic morphemes, which is based on pre-analyzed patterns and generalized HMMs obtained from the given tagged corpus. Experimental results show the promise of our proposed method. Seung-Hoon Na |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2015 | Two-Stage Document Length Normalization for Information RetrievalabstractThe standard approach for term frequency normalization is based only on the document length. However, it does not distinguish the verbosity from the scope, these being the two main factors determining the document length. Because the verbosity and scope have largely different effects on the increase in term frequency, the standard approach can easily suffer from insufficient or excessive penalization depending on the specific type of long document. To overcome these problems, this article proposes two-stage normalization by performing verbosity and scope normalization separately, and by employing different penalization functions. In verbosity normalization, each document is prenormalized by dividing the term frequency by the verbosity of the document. In scope normalization, an existing retrieval model is applied in a straightforward manner to the prenormalized document, finally leading us to formulate our proposed verbosity normalized (VN) retrieval model. Experimental results carried out on standard TREC collections demonstrate that the VN model leads to marginal but statistically significant improvements over standard retrieval models. Seung-Hoon Na |
ACM Trans. Inf. Syst. | 1 |
| 2014 | Partial-update dimensionality reduction for accumulating co-occurrence events
Seung-Hoon Na, Jong-Hyeok Lee |
Pattern Recognit. Lett. | 1 |
| 2013 | Refining sentence similarity with discourse information in dialog system
Sangkeun Jung, Seung-Hoon Na |
INTERSPEECH | 2 |
| 2013 | Probabilistic co-relevance for query-sensitive similarity measurement in information retrieval
Seung-Hoon Na |
Inf. Process. Manag. | 1 |
| 2012 | Utilizing local evidence for blog feed search
Yeha Lee, Seung-Hoon Na, Jong-Hyeok Lee |
Inf. Retr. | 2 |
| 2012 | Memory-restricted latent semantic analysis to accumulate term-document co-occurrence events
Seung-Hoon Na, Jong-Hyeok Lee |
Pattern Recognit. Lett. | 1 |
| 2011 | Enriching document representation via translation for improved monolingual information retrievalabstractWord ambiguity and vocabulary mismatch are critical problems in information retrieval. To deal with these problems, this paper proposes the use of translated words to enrich document representation, going beyond the words in the original source language to represent a document. In our approach, each original document is automatically translated into an auxiliary language, and the resulting translated document serves as a semantically enhanced representation for supplementing the original bag of words. The core of our translation representation is the expected term frequency of a word in a translated document, which is calculated by averaging the term frequencies over all possible translations, rather than focusing on the 1-best translation only. To achieve better efficiency of translation, we do not rely on full-fledged machine translation, but instead use monotonic translation by removing the time-consuming reordering component. Experiments carried out on standard TREC test collections show that our proposed translation representation leads to statistically significant improvements over using only the original language of the document collection. Seung-Hoon Na, Hwee Tou Ng |
SIGIR | 1 |
| 2010 | RankSVR: can preference data help regression?abstractIn some regression applications (e.g., an automatic movie scoring system), a large number of ranking data is available in addition to the original regression data. This paper studies whether and how the ranking data can improve the accuracy of regression task. In particular, this paper first proposes an extension of SVR (Support Vector Regression), RankSVR, which incorporates ranking constraints in the learning of regression function. Second, this paper proposes novel sampling methods for RankSVR, which selectively choose samples of ranking data for training of regression functions in order to maximize the performance of RankSVR. While it is relatively easier to acquire ranking data than regression data, incorporating all the ranking data in the learning of regression doest not always generate the best output. Moreoever, adding too many ranking constraints into the regression problem substantially lengthens the training time. Our proposed sampling methods find the ranking samples that maximize the regression performance. Experimental results on synthetic and real data sets show that, when the ranking data is additionally available, RankSVR significantly performs better than SVR by utilizing ranking constraints in the learning of regression, and also show that our sampling methods improve the RankSVR performance better than the random sampling. Hwanjo Yu, Sungchul Kim, Seung-Hoon Na |
CIKM | 3 |
| 2009 | An improved feedback approach using relevant local posts for blog feed retrievalabstractBlog feed search aims to identify a blog feed with a recurring interest in a given topic. In this paper, we investigate the "pseudo-relevance feedback" for blog feed search task, where its unit of relevance judgment is not based on a blog post but a blog feed (the collection of all its constituent posts). This paper focuses on two characteristics of feed search task, blog feed's topical diversity and multifaceted property of query. We propose a novel feed-level selection of local posts which uses only highly relevant local posts in each top-ranked feed, in order to capture the correct and diverse relevant information to a given topic. Experimental results show that the proposed approach outperforms traditional feedback approaches. Especially, the proposed approach gives 2% further increase of nDCG over the best performing result of TREC '08 Blog Distillation Task. Yeha Lee, Seung-Hoon Na, Jong-Hyeok Lee |
CIKM | 2 |
| 2009 | Improving Opinion Retrieval Based on Query-Specific Sentiment Lexicon
Seung-Hoon Na, Yeha Lee, Sang-Hyob Nam, Jong-Hyeok Lee |
ECIR | 1 |
| 2009 | DiffPost: Filtering Non-relevant Content Based on Content Difference between Two Consecutive Blog Posts
Sang-Hyob Nam, Seung-Hoon Na, Yeha Lee, Jong-Hyeok Lee |
ECIR | 2 |
| 2009 | A 2-poisson model for probabilistic coreference of named entities for improved text retrievalabstractText retrieval queries frequently contain named entities. The standard approach of term frequency weighting does not work well when estimating the term frequency of a named entity, since anaphoric expressions (like he, she, the movie, etc) are frequently used to refer to named entities in a document, and the use of anaphoric expressions causes the term frequency of named entities to be underestimated. In this paper, we propose a novel 2-Poisson model to estimate the frequency of anaphoric expressions of a named entity, without explicitly resolving the anaphoric expressions. Our key assumption is that the frequency of anaphoric expressions is distributed over named entities in a document according to the probabilities of whether the document is elite for the named entities. This assumption leads us to formulate our proposed Co-referentially Enhanced Entity Frequency (CEEF). Experimental results on the text collection of TREC Blog Track show that CEEF achieves significant and consistent improvements over state-of-the-art retrieval methods using standard term frequency estimation. In particular, we achieve a 3% increase of MAP over the best performing run of TREC 2008 Blog Track. Seung-Hoon Na, Hwee Tou Ng |
SIGIR | 1 |
| 2009 | On co-authorship for author disambiguation
In-Su Kang, Seung-Hoon Na, Hanmin Jung, Pyung Kim, Won-Kyung Sung, Jong-Hyeok Lee |
Inf. Process. Manag. | 2 |
| 2008 | Improving Term Frequency Normalization for Multi-topical Documents and Application to Language Modeling Approaches
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee |
ECIR | 1 |
| 2008 | Structural Re-ranking with Cluster-Based Retrieval
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee |
ECIR | 1 |
| 2008 | Revisit of Nearest Neighbor Test for Direct Evaluation of Inter-document Similarities
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee |
ECIR | 1 |
| 2008 | Query-Based Inter-document Similarity Using Probabilistic Co-relevance Model
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee |
ECIR | 1 |
| 2008 | Automatic Extraction of English-Chinese Transliteration Pairs using Dynamic Window and Tokenizer
Chengguo Jin, Seung-Hoon Na, Dong-Il Kim, Jong-Hyeok Lee |
IJCNLP | 2 |
| 2008 | Search Result Clustering Using Label Language Model
Yeha Lee, Seung-Hoon Na, Jong-Hyeok Lee |
IJCNLP | 2 |
| 2008 | Exploiting proximity feature in bigram language model for information retrievalabstractLanguage modeling approaches have been effectively dealing with the dependency among query terms based on N-gram such as bigram or trigram models. However, bigram language models suffer from adjacency-sparseness problem which means that dependent terms are not always adjacent in documents, but can be far from each other, sometimes with distance of a few sentences in a document. To resolve the adjacency-sparseness problem, this paper proposes a new type of bigram language model by explicitly incorporating the proximity feature between two adjacent terms in a query. Experimental results on three test collections show that the proposed bigram language model significantly improves previous bigram model as well as Tao's approach, the state-of-art method for proximity-based method. Seung-Hoon Na, Jungi Kim, In-Su Kang, Jong-Hyeok Lee |
SIGIR | 1 |
| 2007 | Cluster-based patent retrieval
In-Su Kang, Seung-Hoon Na, Jungi Kim, Jong-Hyeok Lee |
Inf. Process. Manag. | 2 |
| 2007 | Parsimonious translation models for information retrieval
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee |
Inf. Process. Manag. | 1 |
| 2007 | Adaptive document clustering based on query-based similarity
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee |
Inf. Process. Manag. | 1 |
| 2007 | An empirical study of query expansion and cluster-based retrieval in language modeling approach
Seung-Hoon Na, In-Su Kang, Ji-Eun Roh, Jong-Hyeok Lee |
Inf. Process. Manag. | 1 |
| 2006 | Collection-based compound noun segmentation for Korean information retrieval
In-Su Kang, Seung-Hoon Na, Jong-Hyeok Lee |
Inf. Retr. | 2 |
| 2004 | Influence of WSD on Cross-Language Information Retrieval
In-Su Kang, Seung-Hoon Na, Jong-Hyeok Lee |
IJCNLP | 2 |
| 2004 | Improving Relevance Feedback in Language Modeling Approach: Maximum a Posteriori Probability Criterion and Three-Component Mixture Model
Seung-Hoon Na, In-Su Kang, Jong-Hyeok Lee |
IJCNLP | 1 |
| 2004 | Lightweight Natural Language Database Interfaces
In-Su Kang, Seung-Hoon Na, Jong-Hyeok Lee, Gijoo Yang |
NLDB | 2 |