Hae-Chang Rim

dblp:61/4583 · DBLP profile ↗
← Back
67ranked-venue papers
1as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 1 first-authorDatabases, data management, data science and information retrieval · 16Graphics, computer vision, multimedia, augmented reality and games · 3Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
12 papers
Information retrieval · 76% Data mining · 18% Web and social media mining · 6%
Artificial intelligence
11 papers
Question answering and dialogue systems · 37% Information extraction and text analysis · 37% Probabilistic and Bayesian machine learning · 13%

Topics — the 30 heaviest of 47, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval models
0.232010
High precision opinion retrieval using sentiment-relevance flows · SIGIR 2010
Word or Phrase? Learning Which Unit to Stress for Information Retrieval · ACL/IJCNLP 2009
Latent semantic indexing model for boolean query formulation · SIGIR 2000
Information retrieval
ranking
0.222010
Achieving high accuracy retrieval using intra-document term ranking · SIGIR 2010
Question Utility: A Novel Static Ranking of Question Search · AAAI 2008
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging
0.242009
Probabilistic Modeling of Korean Morphology · IEEE Trans. Speech Audio Process. 2009
Self-Organizing Markov Models and Their Application to Part-of-Speech Tagging · ACL 2003
Hidden Markov Model-Based Korean Part-of-Speech Tagging Considering High Agglutinativity, Word-Spacing, and Lexical Correlativity · ACL 2000
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering
knowledge-based question answering
0.212014
Joint Relational Embeddings for Knowledge-based Question Answering · EMNLP 2014
Information retrieval › ranking
graph-based ranking
0.112012
Finding interesting posts in Twitter based on retweet graph analysis · SIGIR 2012
Information retrieval › web search › web information retrieval
social media retrieval
0.112012
Finding interesting posts in Twitter based on retweet graph analysis · SIGIR 2012
Web and social media mining
social network analysis
0.112012
Finding interesting posts in Twitter based on retweet graph analysis · SIGIR 2012
Data mining › text mining
information extraction and text analysis
0.112010
Contextual video advertising system using scene information inferred from video scripts · SIGIR 2010
Information retrieval › document retrieval
opinion retrieval
0.112010
High precision opinion retrieval using sentiment-relevance flows · SIGIR 2010
Information retrieval
topic relevance
0.112010
High precision opinion retrieval using sentiment-relevance flows · SIGIR 2010
Data mining › text mining › text classification
naive bayes text classification
0.122006
Some Effective Techniques for Naive Bayes Text Classification · IEEE Trans. Knowl. Data Eng. 2006
A new method of parameter estimation for multinomial naive bayes text classifiers · SIGIR 2002
Data mining › text mining
text classification
0.122006
Some Effective Techniques for Naive Bayes Text Classification · IEEE Trans. Knowl. Data Eng. 2006
A new method of parameter estimation for multinomial naive bayes text classifiers · SIGIR 2002
Natural language and speech › Information extraction and text analysis
morphological analysis
0.112009
Probabilistic Modeling of Korean Morphology · IEEE Trans. Speech Audio Process. 2009
Information retrieval › text analysis
keyword extraction
0.112009
Finding advertising keywords on video scripts · SIGIR 2009
Information retrieval › retrieval models
term weighting
0.112009
Word or Phrase? Learning Which Unit to Stress for Information Retrieval · ACL/IJCNLP 2009
Information retrieval
text analysis
0.112009
Finding advertising keywords on video scripts · SIGIR 2009
Information retrieval
query understanding
0.112008
Bridging Lexical Gaps between Queries and Questions on Large Online Q&A Collections with Compact Translation Models · EMNLP 2008
Natural language and speech › Question answering and dialogue systems › open-ended question answering
definitional question answering
0.112006
Probabilistic model for definitional question answering · SIGIR 2006
Natural language and speech › Question answering and dialogue systems
domain-specific question answering
0.112006
K-QARD: A Practical Korean Question Answering Framework for Restricted Domain · ACL 2006
Data mining › predictive modeling
classification
0.112006
Some Effective Techniques for Naive Bayes Text Classification · IEEE Trans. Knowl. Data Eng. 2006
Data mining › dimensionality reduction › feature selection
feature weighting
0.112006
Some Effective Techniques for Naive Bayes Text Classification · IEEE Trans. Knowl. Data Eng. 2006
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge base
0.112014
Joint Relational Embeddings for Knowledge-based Question Answering · EMNLP 2014
Information retrieval
query processing
0.012004
Information retrieval using word senses: root sense tagging approach · SIGIR 2004
Information retrieval › search engines
semantic search
0.012004
Information retrieval using word senses: root sense tagging approach · SIGIR 2004
Information retrieval › query understanding
word sense disambiguation
0.012004
Information retrieval using word senses: root sense tagging approach · SIGIR 2004
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
markov chain
0.012003
Self-Organizing Markov Models and Their Application to Part-of-Speech Tagging · ACL 2003
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › markov chain
variable memory markov models
0.012003
Self-Organizing Markov Models and Their Application to Part-of-Speech Tagging · ACL 2003
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model
0.022000
Part-of-Speech Tagging Based on Hidden Markov Model Assuming Joint Independence · ACL 2000
Hidden Markov Model-Based Korean Part-of-Speech Tagging Considering High Agglutinativity, Word-Spacing, and Lexical Correlativity · ACL 2000
Information retrieval › online advertising › sponsored search
ad retrieval
0.012010
Contextual video advertising system using scene information inferred from video scripts · SIGIR 2010
Information retrieval
evaluation
0.012010
High precision opinion retrieval using sentiment-relevance flows · SIGIR 2010

Methods — techniques the papers use, named apart from their topics

relational embedding · 0.2latent space mapping · 0.2compact translation models · 0.2HITS algorithm variant · 0.1sentiment-relevance flow · 0.1script-based inference · 0.1unit selection for retrieval · 0.1scene-based features · 0.1probabilistic modeling · 0.1learning-based feature selection · 0.1static ranking · 0.1naive bayes · 0.1language modeling · 0.1feature weighting · 0.1back-off smoothing · 0.1hidden markov model · 0.1syllable-based word recognition · 0.0statistical information acquisition · 0.0
YearPublicationVenuePosition
2015 Knowledge-based question answering using the semantic embedding space
Do-Gil Lee, So-Young Park, Hae-Chang Rim
Expert Syst. Appl.4
2014 Joint Relational Embeddings for Knowledge-based Question Answering
abstract
Transforming a natural language (NL) question into a corresponding logical form (LF) is central to the knowledge-based question answering (KB-QA) task.Unlike most previous methods that achieve this goal based on mappings between lexicalized phrases and logical predicates, this paper goes one step further and proposes a novel embedding-based approach that maps NL-questions into LFs for KB-QA by leveraging semantic associations between lexical representations and KBproperties in the latent space.Experimental results demonstrate that our proposed method outperforms three KB-QA baseline methods on two publicly released QA data sets.
Nan Duan 0001, Ming Zhou 0001, Hae-Chang Rim
EMNLP4
2014 Identifying interesting Twitter contents using topical analysis
Hae-Chang Rim
Expert Syst. Appl.2
2014 Discovering High-Quality Threaded Discussions in Online Forums
Jung-Tae Lee, Hae-Chang Rim
J. Comput. Sci. Technol.3
2014 Multiple categorizations of products: cognitive modeling of customers through social media data mining
Gil-Young Song, Youngjoon Cheon, Kihwang Lee, Heui-Seok Lim, Kyung-Yong Chung, Hae-Chang Rim
Pers. Ubiquitous Comput.6
2012 Assessing Writing Fluency of non-English-speaking Student for Automated Essay Scoring - How to Automatically Evaluate the Fluency in English Essay
Min-Jeong Kim, Hyoung-Gyu Lee, Hae-Chang Rim
CSEDU (2)4
2012 Finding interesting posts in Twitter based on retweet graph analysis
abstract
Millions of posts are being generated in real-time by users in social networking services, such as Twitter. However, a considerable number of those posts are mundane posts that are of interest to the authors and possibly their friends only. This paper investigates the problem of automatically discovering valuable posts that may be of potential interest to a wider audience. Specifically, we model the structure of Twitter as a graph consisting of users and posts as nodes and retweet relations between the nodes as edges. We propose a variant of the HITS algorithm for producing a static ranking of posts. Experimental results on real world data demonstrate that our method can achieve better performance than several baseline methods.
Jung-Tae Lee, Seung-Wook Lee, Hae-Chang Rim
SIGIR4
2012 A new generative opinion retrieval model integrating multiple ranking factors
Seung-Wook Lee, Young-In Song, Jung-Tae Lee, Kyoung-Soo Han, Hae-Chang Rim
J. Intell. Inf. Syst.5
2012 Content-based mobile spam classification using stylistically motivated features
Dae-Neung Sohn, Jung-Tae Lee, Kyoung-Soo Han, Hae-Chang Rim
Pattern Recognit. Lett.4
2011 Phrase Segmentation Model using Collocation and Translational Entropy
Hyoung-Gyu Lee, Joo-Young Lee, Min-Jeong Kim, Hae-Chang Rim, Joong-Hwi Shin, Young-Sook Hwang
MTSummit4
2011 Sentence-based relevance flow analysis for high accuracy retrieval
abstract
Traditional ranking models for information retrieval lack the ability to make a clear distinction between relevant and nonrelevant documents at top ranks if both have similar bag-of-words representations with regard to a user query. We aim to go beyond the bag-of-words approach to document ranking in a new perspective, by representing each document as a sequence of sentences. We begin with an assumption that relevant documents are distinguishable from nonrelevant ones by sequential patterns of relevance degrees of sentences to a query. We introduce the notion of relevance flow, which refers to a stream of sentence-query relevance within a document. We then present a framework to learn a function for ranking documents effectively based on various features extracted from their relevance flows and leverage the output to enhance existing retrieval models. We validate the effectiveness of our approach by performing a number of retrieval experiments on three standard test collections, each comprising a different type of document: news articles, medical references, and blog posts. Experimental results demonstrate that the proposed approach can improve the retrieval performance at the top ranks significantly as compared with the state-of-the-art retrieval models regardless of document type.
Jung-Tae Lee, Jangwon Seo, Jiwoon Jeon, Hae-Chang Rim
J. Assoc. Inf. Sci. Technol.4
2010 An Empirical Study on Web Mining of Parallel Data
Gumwon Hong, Chi-Ho Li, Ming Zhou 0001, Hae-Chang Rim
COLING4
2010 Identifying Idiomatic Expressions Using Phrase Alignments in Bilingual Parallel Corpus
Hyoung-Gyu Lee, Min-Jeong Kim, Gumwon Hong, Sang-Bum Kim, Young-Sook Hwang, Hae-Chang Rim
PRICAI6
2010 High precision opinion retrieval using sentiment-relevance flows
abstract
Opinion retrieval involves the measuring of opinion score of a document about the given topic. We propose a new method, namely sentiment-relevance flow, that naturally unifies the topic relevance and the opinionated nature of a document. Experiments conducted over a large-scaled Web corpus show that the proposed approach improves performance of opinion retrieval in terms of precision at top ranks.
Seung-Wook Lee, Jung-Tae Lee, Young-In Song, Hae-Chang Rim
SIGIR4
2010 Achieving high accuracy retrieval using intra-document term ranking
abstract
Most traditional ranking models roughly score the relevance of a given document by observing simple term statistics, such as the occurrence of query terms within the document or within the collection. Intuitively, the relative importance of query terms with regard to other individual non-query terms in a document can also be exploited to promote the ranks of documents in which the query is dedicated as the main topic. In this paper, we introduce a simple technique named intra-document term ranking, which involves ranking all the terms in a document according to their relative importance within that particular document. We demonstrate that the information regarding the rank positions of given query terms within the intra-document term ranking can be useful for enhancing the precision of top-retrieved results by traditional ranking models. Experiments are conducted on three standard TREC test collections.
Hyun-Wook Woo, Jung-Tae Lee, Seung-Wook Lee, Young-In Song, Hae-Chang Rim
SIGIR5
2010 Contextual video advertising system using scene information inferred from video scripts
abstract
With the rise of digital video consumptions, contextual video advertising demands have been increasing in recent years. This paper presents a novel video advertising system that selects relevant text ads for a given video scene by automatically identifying the situation of the scene. The situation information of video scenes is inferred from available video scripts. Experimental results show that the use of the situation information enhances the accuracy of ad retrieval for video scenes. The proposed system represents one of the pioneer video advertising systems using contextual information obtained from video scripts.
Bong-Jun Yi, Jung-Tae Lee, Hyun-Wook Woo, Hae-Chang Rim
SIGIR4
2009 Word or Phrase? Learning Which Unit to Stress for Information Retrieval
Young-In Song, Jung-Tae Lee, Hae-Chang Rim
ACL/IJCNLP3
2009 Finding advertising keywords on video scripts
abstract
A key to success to contextual in-video advertising is finding advertising keywords on video contents effectively, but there has been little literature in the area so far. This paper presents some preliminary results of our learning-based system that finds relevant advertising keywords on particular scene of video contents using their scripts. The system is trained with not only features proven useful in earlier studies but novel features that reflect the situation of a targeted scene. Experimental results show that the new features are potentially helpful for enhancing the accuracy of keyword extraction for contextual in-video advertising.
Jung-Tae Lee, Hyungdong Lee, Hee-Seon Park, Young-In Song, Hae-Chang Rim
SIGIR5
2009 Probabilistic Modeling of Korean Morphology
abstract
This paper proposes new probabilistic models for analyzing Korean morphology. In order to take advantage of the characteristics of Korean morphology, the proposed models are based on three linguistic units: eojeol (a Korean spacing unit), morpheme, and syllable. Unlike previous approaches that are based on rules and dictionaries, the probabilistic approach proposed in this study can automatically acquire complete linguistic knowledge from part-of-speech (POS) tagged corpora. In addition, this approach, without any system modification, is easily applicable to other corpora with different tag sets and annotation guidelines. The three different models and their combinations are evaluated on three corpora over a wide range of conditions. The eojeol-unit and syllable-unit models compensate for the weaknesses of the morpheme-unit model. The eojeol-unit model performed efficiently, and improved the precision. The syllable-unit model improved in precision as well, showing a particularly robust performance in treating unknown words. The proposed approach is also proven to outperform the previous approaches.
Do-Gil Lee, Hae-Chang Rim
IEEE Trans. Speech Audio Process.2
2008 Question Utility: A Novel Static Ranking of Question Search
Young-In Song, Chin-Yew Lin, Yunbo Cao, Hae-Chang Rim
AAAI4
2008 Semantic Dependency Parsing using N-best Semantic Role Sequences and Roleset Information
Joo-Young Lee, Hancheol Cho, Hae-Chang Rim
CoNLL3
2008 Bridging Lexical Gaps between Queries and Questions on Large Online Q&A Collections with Compact Translation Models
Jung-Tae Lee, Sang-Bum Kim, Young-In Song, Hae-Chang Rim
EMNLP4
2008 Combining Local and Global Resources for Constructing an Error-Minimized Opinion Word Dictionary
Jung-Tae Lee, Young-In Song, Hae-Chang Rim
PRICAI4
2008 A novel retrieval approach reflecting variability of syntactic phrase representation
Young-In Song, Kyoung-Soo Han, Sang-Bum Kim, So-Young Park, Hae-Chang Rim
J. Intell. Inf. Syst.5
2007 Building a Large-Scale Commonsense Knowledge Base by Converting an Existing One in a Different Language
Yuchul Jung, Joo-Young Lee, Sung-Hyon Myaeng, Hae-Chang Rim
CICLing6
2007 Answer extraction and ranking strategies for definitional question answering using linguistic features and definition terminology
Kyoung-Soo Han, Young-In Song, Sang-Bum Kim, Hae-Chang Rim
Inf. Process. Manag.4
2006 K-QARD: A Practical Korean Question Answering Framework for Restricted Domain
abstract
We present a Korean question answering framework for restricted domains, called K-QARD. K-QARD is developed to achieve domain portability and robustness, and the framework is successfully applied to build question answering systems for several domains.
Young-In Song, Hoo-Jung Chung, Kyoung-Soo Han, Joo-Young Lee, Hae-Chang Rim, Jae-Won Lee
ACL5
2006 Probabilistic model for definitional question answering
abstract
This paper proposes a probabilistic model for definitional question answering (QA) that reflects the characteristics of the definitional question. The intention of the definitional question is to request the definition about the question target. Therefore, an answer for the definitional question should contain the content relevant to the topic of the target, and have a representation form of the definition style. Modeling the problem of definitional QA from both the topic and definition viewpoints, the proposed probabilistic model converts the task of answering the definitional questions into that of estimating the three language models: topic language model, definition language model, and general language model. The proposed model systematically combines several evidences in a probabilistic framework. Experimental results show that a definitional QA system based on the proposed probabilistic model is comparable to state-of-the-art systems.
Kyoung-Soo Han, Young-In Song, Hae-Chang Rim
SIGIR3
2006 ME-based biomedical named entity recognition using lexical knowledge
abstract
In this paper, we present a two-phase biomedical NE-recognition method based on a ME model: we first recognize biomedical terms and then assign appropriate semantic classes to the recognized terms. In the two-phase NE-recognition method, the performance of the term-recognition phase is very important, because the semantic classification is performed on the region identified at the recognition phase. In this study, in order to improve the performance of term recognition, we try to incorporate lexical knowledge into pre- and postprocessing of the term-recognition phase. In the preprocessing step, we use domain-salient words as lexical knowledge obtained by corpus comparison. In the postprocessing step, we utilize χ 2 -based collocations gained from Medline corpus. In addition, we use morphological patterns extracted from the training data as features for learning the ME-based classifiers. Experimental results show that the performance of NE-recognition can be improved by utilizing such lexical knowledge.
Kyung-Mi Park, Hae-Chang Rim, Young-Sook Hwang
ACM Trans. Asian Lang. Inf. Process.3
2006 Some Effective Techniques for Naive Bayes Text Classification
abstract
While naive Bayes is quite effective in various data mining tasks, it shows a disappointing result in the automatic text classification problem. Based on the observation of naive Bayes for the natural language text, we found a serious problem in the parameter estimation process, which causes poor results in text classification domain. In this paper, we propose two empirical heuristics: per-document text normalization and feature weighting method. While these are somewhat ad hoc methods, our proposed naive Bayes text classifier performs very well in the standard benchmark collections, competing with state-of-the-art text classifiers based on a highly complex learning method such as SVM.
Sang-Bum Kim, Kyoung-Soo Han, Hae-Chang Rim, Sung-Hyon Myaeng
IEEE Trans. Knowl. Data Eng.3
2005 Maximum Entropy Based Semantic Role Labeling
Kyung-Mi Park, Hae-Chang Rim
CoNLL2
2005 Two-Phase Biomedical Named Entity Recognition Using A Hybrid Method
Juntae Yoon, Kyung-Mi Park, Hae-Chang Rim
IJCNLP4
2005 Word Sense Disambiguation by Relative Selection
Hee-Cheol Seo, Hae-Chang Rim, Myung-Gil Jang
IJCNLP2
2005 Feature-Based Korean Grammar Utilizing Learned Constraint Rules
abstract
In this paper, we propose a feature-based Korean grammar utilizing the learned constraint rules in order to improve parsing efficiency. The proposed grammar consists of feature structures, feature operations, and constraint rules; and it has the following characteristics. First, a feature structure includes several features to express useful linguistic information for Korean parsing. Second, a feature operation generating a new feature structure is restricted to the binary-branching form which can deal with Korean properties such as variable word order and constituent ellipsis. Third, constraint rules improve efficiency by preventing feature operations from generating spurious feature structures. Moreover, these rules are learned from a Korean treebank by a decision tree learning algorithm. The experimental results show that the feature-based Korean grammar can reduce the number of candidates by a third of candidates at most and it runs 1.5 ∼ 2 times faster than a CFG on a statistical parser.
So-Young Park, Yong-Jae Kwak, Hae-Chang Rim, Heui-Seok Lim
Comput. Intell.3
2005 Improving query translation in English-Korean cross-language information retrieval
Hee-Cheol Seo, Sang-Bum Kim, Hae-Chang Rim, Sung-Hyon Myaeng
Inf. Process. Manag.3
2004 Unsupervised Event Extraction from Biomedical Text Based on Event and Pattern Information
Hong-Woo Chun, Young-Sook Hwang, Hae-Chang Rim
CICLing3
2004 Unlexicalized Dependency Parser for Variable Word Order Languages Based on Local Contextual Pattern
Hoo-Jung Chung, Hae-Chang Rim
CICLing2
2004 Comparative Analysis of Term Distributions in a Sentence and in a Document for Sentence Retrieval
Kyoung-Soo Han, Hae-Chang Rim
CICLing2
2004 Recomputation of Class Relevance Scores for Improving Text Classification
Sang-Bum Kim, Hae-Chang Rim
CICLing2
2004 Probabilistic Shift-Reduce Parsing Model Using Rich Contextual Information
Yong-Jae Kwak, So-Young Park, Joon-Ho Lim, Hae-Chang Rim
CICLing4
2004 Towards Language-Independent Sentence Boundary Detection
Do-Gil Lee, Hae-Chang Rim
CICLing2
2004 A Semi-automatic Tree Annotating Workbench for Building a Korean Treebank
Joon-Ho Lim, So-Young Park, Yong-Jae Kwak, Hae-Chang Rim
CICLing4
2004 Evaluation of Feature Combination for Effective Structural Disambiguation
So-Young Park, Yong-Jae Kwak, Joon-Ho Lim, Hae-Chang Rim
CICLing4
2004 Word Sense Disambiguation Based on Weight Distribution Model with Multiword Expression
Hee-Cheol Seo, Young-Sook Hwang, Hae-Chang Rim
CICLing3
2004 A Term Weighting Method Based on Lexical Chain for Automatic Summarization
Young-In Song, Kyoung-Soo Han, Hae-Chang Rim
CICLing3
2004 Semantic Role Labeling using Maximum Entropy Model
Joon-Ho Lim, Young-Sook Hwang, So-Young Park, Hae-Chang Rim
CoNLL4
2004 Two-Phase Semantic Role Labeling based on Support Vector Machines
Kyung-Mi Park, Young-Sook Hwang, Hae-Chang Rim
CoNLL3
2004 Unsupervised Event Extraction from Biomedical Literature Using Co-occurrence Information and Basic Patterns
Hong-Woo Chun, Young-Sook Hwang, Hae-Chang Rim
IJCNLP3
2004 Syllable-based probabilistic morphological analysis model of Korean
abstract
In this paper, we present a syllable-based probabilistic morphological analysis model of Korean. While the pre-viousmorphological analyzers that regardmorpheme as a processing unit, the model exploits syllable as a process-ing unit in order to endure the unknown word problem. Actually, it does not use any morpheme dictionary. In contract to the previous systems that depend on manually constructed linguistic knowledge, the proposed system can fully automatically acquire the linguistic knowledge from annotated corpora. Besides, without any modifica-tion, the system can be applied to other corpus having a different tagset and annotation guidelines. We describe the model and present experimental results on two cor-pora. 1.
Do-Gil Lee, Hae-Chang Rim
INTERSPEECH2
2004 Partially lexicalized parsing model utilizing rich features
So-Young Park, Yong-Jae Kwak, Joon-Ho Lim, Hae-Chang Rim, Soo-Hong Kim
INTERSPEECH4
2004 Information retrieval using word senses: root sense tagging approach
abstract
Information retrieval using word senses is emerging as a good research challenge on semantic information retrieval. In this paper, we propose a new method using word senses in information retrieval: root sense tagging method. This method assigns coarse-grained word senses defined in WordNet to query terms and document terms by unsupervised way using co-occurrence information constructed automatically. Our sense tagger is crude, but performs consistent disambiguation by considering only the single most informative word as evidence to disambiguate the target word. We also allow multiple-sense assignment to alleviate the problem caused by incorrect disambiguation.Experimental results on a large-scale TREC collection show that our approach to improve retrieval effectiveness is successful, while most of the previous work failed to improve performances even on small text collection. Our method also shows promising results when is combined with pseudo relevance feedback and state-of-the-art retrieval function such as BM25.
Sang-Bum Kim, Hee-Cheol Seo, Hae-Chang Rim
SIGIR3
2004 Unsupervised word sense disambiguation using WordNet relatives
Hee-Cheol Seo, Hoo-Jung Chung, Hae-Chang Rim, Sung-Hyon Myaeng, Soo-Hong Kim
Comput. Speech Lang.3
2004 Biomedical named entity recognition using two-phase model based on SVMs
Ki-Joong Lee, Young-Sook Hwang, Hae-Chang Rim
J. Biomed. Informatics4
2003 Self-Organizing Markov Models and Their Application to Part-of-Speech Tagging
abstract
This paper presents a method to develop a class of variable memory Markov models that have higher memory capacity than traditional (uniform memory) Markov models. The structure of the variable memory models is induced from a manually annotated corpus through a decision tree learning algorithm. A series of comparative experiments show the resulting models outperform uniform memory Markov models in a part-of-speech tagging task.
Jin-Dong Kim, Hae-Chang Rim, Jun'ichi Tsujii
ACL2
2003 A Syllable Based Word Recognition Model for Korean Noun Extraction
abstract
Noun extraction is very important for many NLP applications such as information retrieval, automatic text classification, and information extraction. Most of the previous Korean noun extraction systems use a morphological analyzer or a Part-of-Speech (POS) tagger. Therefore, they require much of the linguistic knowledge such as morpheme dictionaries and rules (e.g. morphosyntactic rules and morphological rules).This paper proposes a new noun extraction method that uses the syllable based word recognition model. It finds the most probable syllable-tag sequence of the input sentence by using automatically acquired statistical information from the POS tagged corpus and extracts nouns by detecting word boundaries. Furthermore, it does not require any labor for constructing and maintaining linguistic knowledge. We have performed various experiments with a wide range of variables influencing the performance. The experimental results show that without morphological analysis or POS tagging, the proposed method achieves comparable performance with the previous methods.
Do-Gil Lee, Hae-Chang Rim, Heui-Seok Lim
ACL2
2002 Effective Methods for Improving Naive Bayes Text Classifiers
Sang-Bum Kim, Hae-Chang Rim, Dongsuk Yook, Heui-Seok Lim
PRICAI2
2002 A new method of parameter estimation for multinomial naive bayes text classifiers
abstract
Multinomial naive Bayes classifiers have been widely used for the probabilistic text classification. However, their parameter estimation method sometimes generates inappropriate probabilities. In this paper, we propose a topic document model approach for naive Bayes text classification, where their parameters are estimated with an expectation from the training documents. Experiments are conducted on Reuters 21578 and 20 Newsgroup collection, and our proposed approach obtained a significant improvement in performace over the conventional approach.
Sang-Bum Kim, Hae-Chang Rim, Heui-Seok Lim
SIGIR2
2000 Part-of-Speech Tagging Based on Hidden Markov Model Assuming Joint Independence
abstract
In this paper we present part-of-speech taggers based on hidden Markov models, which adopt a less strict Markov assumption to consider rich contexts. In models whose parameters are very specific like lexicalized ones, sparse-data problem is very serious and also conditional probabilities tend to be estimated unreliably. To overcome data-sparseness, a simplified version of the well-known back-off smoothing method is used. To mitigate unreliable estimation problem, our models assume joint independence instead of conditional independence because joint probabilities have the same degree of estimation reliability. In experiments for the Brown corpus, models with rich contexts achieve relatively high accuracy and some models assuming joint independence show better results than the corresponding HMMs.
Sang-Zoo Lee, Jun'ichi Tsujii, Hae-Chang Rim
ACL3
2000 Hidden Markov Model-Based Korean Part-of-Speech Tagging Considering High Agglutinativity, Word-Spacing, and Lexical Correlativity
abstract
&% ' )( * '+ -, .! !/ 10 2 3 4 4 5 6 74 8 ) 7 9 !: ; 74 < 64 )0 >= ?4 4 ) ;@ -% 9 ; 7 A% 9 B C % ED F HG -'+ -, 7 / ! ; 74 -F H " 7 JI K ML C / N " 8 A% 9 -% ED O QP R !S 9T VU XW MW ZY [ '[ ]\ _9ba Mc R NT ed gf h d 5i kj lc m nd od p 7q Vr \ _s ;t *u q vt w _x Xy B&x zw _s *u 3\ t e{ |w E^} V{ q & q 9{ | q s * ;{ |t B w x A w E v Ew w E9 Ew 3 B _ 9 5 9& ' _w B ' w E v Ew E B E E E 9 E\ r \ _^ q nq t * * ;{ |{ o 7{ V V t ew E v _w 9 ¡\ _ r ¢ P 9[ ET ¤£ ¥h !P R !S §¦ d g¨ p 7q nr \ s ;t eu q vt ©w _x Xª 5w Eu r &t q s ©} n{ q 9
Sang-Zoo Lee, Jun'ichi Tsujii, Hae-Chang Rim
ACL3
2000 Lexicalized Hidden Markov Models for Part-of-Speech Tagging
Sang-Zoo Lee, Jun'ichi Tsujii, Hae-Chang Rim
COLING3
2000 KCAT: A Korean Corpus Annotating Tool Minimizing Human Intervention
Won-Ho Ryu, Jin-Dong Kim, Hae-Chang Rim, Heui-Seok Lim
COLING3
2000 Latent semantic indexing model for boolean query formulation
abstract
A new model named Boolean Latent Semantic Indexing model based on the Singular Value Decomposition and Boolean query formulation is introduced. While the Singular Value Decomposition alleviates the problems of lexical matching in the traditional information retrieval model, Boolean query formulation can help users to make precise representation of their information search needs. Retrieval experiments on a number of test collections seem to show that the proposed model achieves substantial performance gains over the Latent Semantic Indexing model.
Dae-Ho Baek, Heui-Seok Lim, Hae-Chang Rim
SIGIR3
1999 HMM Specialization with Selective Lexicalization
Jin-Dong Kim, Sang-Zoo Lee, Hae-Chang Rim
EMNLP3
1999 Resolving Ambiguious Segmentation of Korean Compound Nouns Using Statistics and Rules
abstract
Korean compound nouns may be written as a sequence of characters without blanks between unit nouns. For Korean processing systems, Korean compound nouns have to be first segmented into a sequence of unit nouns. However, the segmentation task is difficult because a sequence of characters may be ambiguously segmented to several sequences of appropriate unit nouns. Moreover, this task is not trivial because Korean compound nouns may include many unknown unit nouns. This paper proposes a new method for KCNS (Korean Compound Noun Segmentation) and reports on the appliccation of such a segmentationtechnique to enhance the performance of an information retrieval system. According to our method, compound nouns are first segmented by using a dictionary and structure patterns. If they are ambiguously segmented, we resolve the ambiguities by using statistical information and a preference rule. Moreover, we employ three kinds of heuristics in order to segment compound nouns with unknown unit nouns. To evaluate KCNS, we use three kinds of data from various domains. Experimental results show that the precision of KCNS's output is approximately 96% on average, regardless of domains. The effectiveness of using the segmented unit nouns provided by KCNS for indexing is proved by improving retrieval performance of our information retrieval system.
Bo-Hyun Yun, Yong-Jae Kwak, Hae-Chang Rim
Comput. Intell.3
1999 Alleviating syntactic term mismatches in Korean text retrieval
Bo-Hyun Yun, Yong-Jae Kwak, Hae-Chang Rim
Inf. Process. Manag.3
1998 Intriguing Aspects of Oriental Languages
abstract
This paper includes a description of 3 affiliated oriental languages: Chinese, Japanese, and Korean. It includes a description of the origins of these 3 languages and the inter-relationship among them. Drawn from the viewpoints of several experienced researchers in the field of OCR (Optical Character Recognition) and computational linguistics, it attempts to bring out the intriguing aspects of these 3 ideographic languages, including the formation and composition of pictograms, special features, learning, understanding, contextual information, and recognition of characters and words, and their relations to poetic expressions and pattern recognition techniques. Numerous references are given and comments on future trends are also presented.
Ching Y. Suen, Shunji Mori, Hae-Chang Rim, Patrick Shen-Pei Wang
Int. J. Pattern Recognit. Artif. Intell.3
1990 Transforming Syntactic Graphs into Semantic Graphs
abstract
In this paper, we present a computational method for transforming a syntactic graph, which represents all syntactic interpretations of a sentence, into a semantic graph which filters out certain interpretations, but also incorporates any remaining ambiguities. We argue that the resulting ambiguous graph, supported by an exclusion matrix, is a useful data structure for question answering and other semantic processing. Our research is based on the principle that ambiguity is an inherent aspect of natural language communication.
Hae-Chang Rim, Jungyun Seo, Robert F. Simmons
ACL1