VLDB 2026 Research / reviewers in the wild / expert
Sang-goo Lee
dblp:67/5511 · also Sang-Goo Lee
· DBLP profile ↗
76ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-0063-0083ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 33 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10 · 4 since 2021Software engineering, systems software and programming languages · 5 · 3 first-authorSystems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CUB: Benchmarking Context Utilisation Techniques for Language ModelsabstractIncorporating external knowledge is crucial for knowledge-intensive tasks, such as question answering and fact checking. However, language models (LMs) may ignore relevant information that contradicts outdated parametric memory or be distracted by irrelevant contexts. While many context utilisation manipulation techniques (CMTs) have recently been proposed to alleviate these issues, few have seen systematic comparison. In this paper, we develop CUB (Context Utilisation Benchmark) - the first comprehensive benchmark designed to help diagnose CMTs under diverse noisy context conditions within retrieval-augmented generation (RAG). With this benchmark, we conduct the most extensive evaluation to date of seven state-of-the-art methods, representative of the main categories of CMTs, across three diverse datasets and tasks, applied to 11 LMs. Our findings expose critical gaps in current CMT evaluation practices, demonstrating the need for holistic testing. We reveal that most existing CMTs struggle to handle the full spectrum of context types encountered in real-world RAG scenarios. We also find that many CMTs display inflated performance on simple synthesised datasets, compared to more realistic datasets with naturally occurring samples. Lovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee, Richard Johansson, Hyunsoo Cho, Isabelle Augenstein |
ACL (1) | 4 |
| 2025 | When to Speak, When to Abstain: Contrastive Decoding with AbstentionabstractLarge Language Models (LLMs) demonstrate exceptional performance across diverse tasks by leveraging pre-trained (i.e., parametric) and external (i.e., contextual) knowledge. While substantial efforts have been made to enhance the utilization of both forms of knowledge, situations in which models lack relevant information remain underexplored. To investigate this challenge, we first present a controlled testbed featuring four distinct knowledge access scenarios, including the aforementioned edge case, revealing that conventional LLM usage exhibits insufficient robustness in handling all instances. Addressing this limitation, we propose Contrastive Decoding with Abstention (CDA), a novel training-free decoding method that allows LLMs to generate responses when relevant knowledge is available and to abstain otherwise. CDA estimates the relevance of both knowledge sources for a given input, adaptively deciding which type of information to prioritize and which to exclude. Through extensive experiments, we demonstrate that CDA can effectively perform accurate generation and abstention simultaneously, enhancing reliability and preserving user trust. Hyuhng Joon Kim, Youna Kim, Sang-goo Lee, Taeuk Kim |
ACL (1) | 3 |
| 2024 | Complete the Feature Space: Diffusion-Based Fictional ID Generation for Face Recognition
Myeong-Yeon Yi, Naeun Ko, Yonghyun Jeong, Sang-goo Lee, Seunggyu Chang |
BMVC | 5 |
| 2024 | Aligning Language Models to Explicitly Handle AmbiguityabstractHyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim, Choonghyun Park, Kang Min Yoo, Sang-goo Lee, Taeuk Kim. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Hyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim, Choonghyun Park, Kang Min Yoo, Sang-goo Lee, Taeuk Kim |
EMNLP | 7 |
| 2024 | Proxy-based Item Representation for Attribute and Context-aware RecommendationabstractNeural network approaches in recommender systems have shown remarkable success by representing a large set of items as a learnable vector embedding table. However, infrequent items may suffer from inadequate training opportunities, making it difficult to learn meaningful representations. We examine that in attribute and context-aware settings, the poorly learned embeddings of infrequent items impair the recommendation accuracy. To address such an issue, we propose a proxy-based item representation that allows each item to be expressed as a weighted sum of learnable proxy embeddings. Here, the proxy weight is determined by the attributes and context of each item and may incorporate bias terms in case of frequent items to further reflect collaborative signals. The proxy-based method calculates the item representations compositionally, ensuring each representation resides inside a well-trained simplex and, thus, acquires guaranteed quality. Additionally, that the proxy embeddings are shared across all items allows the infrequent items to borrow training signals of frequent items in a unified model structure and end-to-end manner. Our proposed method is a plug-and-play model that can replace the item encoding layer of any neural network-based recommendation model, while consistently improving the recommendation performance with much smaller parameter usage. Experiments conducted on real-world recommendation benchmark datasets demonstrate that our proposed model outperforms state-of-the-art models in terms of recommendation accuracy by up to 17% while using only 10% of the parameters. Jinseok Seol, Minseok Gang, Sang-goo Lee, Jaehui Park |
WSDM | 3 |
| 2023 | Prompt-Augmented Linear Probing: Scaling beyond the Limit of Few-Shot In-Context LearnersabstractThrough in-context learning (ICL), large-scale language models are effective few-shot learners without additional model fine-tuning. However, the ICL performance does not scale well with the number of available training sample as it is limited by the inherent input length constraint of the underlying language model. Meanwhile, many studies have revealed that language models are also powerful feature extractors, allowing them to be utilized in a black-box manner and enabling the linear probing paradigm, where lightweight discriminators are trained on top of the pre-extracted input representations. This paper proposes prompt-augmented linear probing (PALP), a hybrid of linear probing and ICL, which leverages the best of both worlds. PALP inherits the scalability of linear probing and the capability of enforcing language models to derive more meaningful representations via tailoring input into a more conceivable form. Throughout in-depth investigations on various datasets, we verified that PALP significantly closes the gap between ICL in the data-hungry scenario and fine-tuning in the data-abundant scenario with little training overhead, potentially making PALP a strong alternative in a black-box scenario. Hyunsoo Cho, Hyuhng Joon Kim, Junyeob Kim, Sang-Woo Lee 0001, Sang-goo Lee, Kang Min Yoo, Taeuk Kim |
AAAI | 5 |
| 2023 | CELDA: Leveraging Black-box Language Model as Enhanced Classifier without LabelsabstractUtilizing language models (LMs) without internal access is becoming an attractive paradigm in the field of NLP as many cutting-edge LMs are released through APIs and boast a massive scale.The de-facto method in this type of black-box scenario is known as prompting, which has shown progressive performance enhancements in situations where data labels are scarce or unavailable.Despite their efficacy, they still fall short in comparison to fully supervised counterparts and are generally brittle to slight modifications.In this paper, we propose Clustering-Enhanced Linear Discriminative Analysis (CELDA), a novel approach that improves the text classification accuracy with a very weak-supervision signal (i.e., name of the labels).Our framework draws a precise decision boundary without accessing weights or gradients of the LM model or data labels.The core ideas of CELDA are twofold: (1) extracting a refined pseudo-labeled dataset from an unlabeled dataset, and (2) training a lightweight and robust model on the top of LM, which learns an accurate decision boundary from an extracted noisy dataset.Throughout in-depth investigations on various datasets, we demonstrated that CELDA reaches new state-of-theart in weakly-supervised text classification and narrows the gap with a fully-supervised model.Additionally, our proposed methodology can be applied universally to any LM and has the potential to scale to larger models, making it a more viable option for utilizing large LMs. Hyunsoo Cho, Youna Kim, Sang-goo Lee |
ACL (1) | 3 |
| 2023 | Multi-scale Contrastive Learning for Complex Scene GenerationabstractRecent advances in Generative Adversarial Networks (GANs) have enabled photo-realistic synthesis of single object images. Yet, modeling more complex distributions, such as scenes with multiple objects, remains challenging. The difficulty stems from the incalculable variety of scene configurations which contain multiple objects of different categories placed at various locations. In this paper, we aim to alleviate the difficulty by enhancing the discriminative ability of the discriminator through a locally defined self-supervised pretext task. To this end, we design a discriminator to leverage multi-scale local feedback that guides the generator to better model local semantic structures in the scene. Then, we require the discriminator to carry out pixel-level contrastive learning at multiple scales to enhance discriminative capability on local regions. Experimental results on several challenging scene datasets show that our method improves the synthesis quality by a substantial margin compared to state-of-the-art baselines. Hanbit Lee, Youna Kim, Sang-goo Lee |
WACV | 3 |
| 2022 | Ground-Truth Labels Matter: A Deeper Look into Input-Label DemonstrationsabstractKang Min Yoo, Junyeob Kim, Hyuhng Joon Kim, Hyunsoo Cho, Hwiyeol Jo, Sang-Woo Lee, Sang-goo Lee, Taeuk Kim. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Kang Min Yoo, Junyeob Kim, Hyuhng Joon Kim, Hyunsoo Cho, Hwiyeol Jo, Sang-Woo Lee 0001, Sang-goo Lee, Taeuk Kim |
EMNLP | 7 |
| 2022 | Exploiting Session Information in BERT-based Session-aware Sequential RecommendationabstractIn recommendation systems, utilizing the user interaction history as sequential information has resulted in great performance improvement. However, in many online services, user interactions are commonly grouped by sessions that presumably share preferences, which requires a different approach from ordinary sequence representation techniques. To this end, sequence representation models with a hierarchical structure or various viewpoints have been developed but with a rather complex network structure. In this paper, we propose three methods to improve recommendation performance by exploiting session information while minimizing additional parameters in a BERT-based sequential recommendation model: using session tokens, adding session segment embeddings, and a time-aware self-attention. We demonstrate the feasibility of the proposed methods through experiments on widely used recommendation datasets. Jinseok Seol, Youngrok Ko, Sang-goo Lee |
SIGIR | 3 |
| 2021 | Self-Guided Contrastive Learning for BERT Sentence RepresentationsabstractTaeuk Kim, Kang Min Yoo, Sang-goo Lee. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Taeuk Kim, Kang Min Yoo, Sang-goo Lee |
ACL/IJCNLP (1) | 3 |
| 2021 | Masked Contrastive Learning for Anomaly DetectionabstractDetecting anomalies is one fundamental aspect of a safety-critical software system, however, it remains a long-standing problem. Numerous branches of works have been proposed to alleviate the complication and have shown promising results. In particular, self-supervised learning based methods are spurring interest due to their capability of learning diverse representations without additional labels. Among self-supervised learning tactics, contrastive learning is one specific framework showing pronounced results in various fields including anomaly detection. However, the primary objective of contrastive learning is to learn task-agnostic features without any labels, which is not entirely suited to discern anomalies. In this paper, we propose a task-specific variant of contrastive learning named masked contrastive learning, which is more befitted for anomaly detection. Moreover, we propose a new inference method dubbed self-ensemble inference that further boosts performance by leveraging the ability learned through auxiliary self-supervision tasks. By combining our models, we can outperform previous state-of-the-art methods by a significant margin on various benchmark datasets. Hyunsoo Cho, Jinseok Seol, Sang-goo Lee |
IJCAI | 3 |
| 2020 | Variational Hierarchical Dialog Autoencoder for Dialog State Tracking Data AugmentationabstractRecent works have shown that generative data augmentation, where synthetic samples generated from deep generative models complement the training dataset, benefit NLP tasks.In this work, we extend this approach to the task of dialog state tracking for goaloriented dialogs.Due to the inherent hierarchical structure of goal-oriented dialogs over utterances and related annotations, the deep generative model must be capable of capturing the coherence among different hierarchies and types of dialog features.We propose the Variational Hierarchical Dialog Autoencoder (VHDA) for modeling the complete aspects of goal-oriented dialogs, including linguistic features and underlying structured annotations, namely speaker information, dialog acts, and goals.The proposed architecture is designed to model each aspect of goal-oriented dialogs using inter-connected latent variables and learns to generate coherent goal-oriented dialogs from the latent spaces.To overcome training issues that arise from training complex variational models, we propose appropriate training strategies.Experiments on various dialog datasets show that our model improves the downstream dialog trackers' robustness via generative data augmentation.We also discover additional benefits of our unified approach to modeling goal-oriented dialogsdialog response generation and user simulation, where our model outperforms previous strong baselines. Kang Min Yoo, Hanbit Lee, Franck Dernoncourt, Trung Bui, Walter Chang, Sang-goo Lee |
EMNLP (1) | 6 |
| 2020 | Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar Induction
Taeuk Kim, Jihun Choi 0002, Daniel Edmiston, Sang-goo Lee |
ICLR | 4 |
| 2020 | Learning Context Using Segment-Level LSTM for Neural Sequence LabelingabstractThis article introduces an approach that learns segment-level context for sequence labeling in natural language processing (NLP). Previous approaches limit their basic unit to a word for feature extraction because sequence labeling is a tokenlevel task in which labels are annotated word-by-word. However, the text segment is an ultimate unit for labeling, and we are easily able to obtain segment information from annotated labels in a IOB/IOBES format. Most neural sequence labeling models expand their learning capacity by employing additional layers, such as a character-level layer, or jointly training NLP tasks with common knowledge. The architecture of our model is based on the charLSTM-BiLSTM-CRF model, and we extend the model with an additional segment-level layer called segLSTM. We therefore suggest a sequence labeling algorithm called charLSTM-BiLSTMCRF-segLSTMsLM which employs an additional segment-level long short-term memory (LSTM) that trains features by learning adjacent context in a segment. We demonstrate the performance of our model on four sequence labeling datasets, namely, Peen Tree Bank, CoNLL 2000, CoNLL 2003, and OntoNotes 5.0. Experimental results show that our model performs better than state-of-theart variants of BiLSTM-CRF. In particular, the proposed model enhances the performance of tasks for finding appropriate labels of multiple token segments. Youhyun Shin, Sang-goo Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Dynamic Compositionality in Recursive Neural Networks with Structure-Aware Tag RepresentationsabstractMost existing recursive neural network (RvNN) architectures utilize only the structure of parse trees, ignoring syntactic tags which are provided as by-products of parsing. We present a novel RvNN architecture that can provide dynamic compositionality by considering comprehensive syntactic information derived from both the structure and linguistic tags. Specifically, we introduce a structure-aware tag representation constructed by a separate tag-level tree-LSTM. With this, we can control the composition function of the existing wordlevel tree-LSTM by augmenting the representation as a supplementary input to the gate functions of the tree-LSTM. In extensive experiments, we show that models built upon the proposed architecture obtain superior or competitive performance on several sentence-level tasks such as sentiment analysis and natural language inference when compared against previous tree-structured models and other sophisticated neural models. Taeuk Kim, Jihun Choi 0002, Daniel Edmiston, Sanghwan Bae, Sang-goo Lee |
AAAI | 5 |
| 2019 | Data Augmentation for Spoken Language Understanding via Joint Variational GenerationabstractData scarcity is one of the main obstacles of domain adaptation in spoken language understanding (SLU) due to the high cost of creating manually tagged SLU datasets. Recent works in neural text generative models, particularly latent variable models such as variational autoencoder (VAE), have shown promising results in regards to generating plausible and natural sentences. In this paper, we propose a novel generative architecture which leverages the generative power of latent variable models to jointly synthesize fully annotated utterances. Our experiments show that existing SLU models trained on the additional synthetic examples achieve performance gains. Our approach not only helps alleviate the data scarcity issue in the SLU task for many datasets but also indiscriminately improves language understanding performances for various SLU models, supported by extensive experiments and rigorous statistical testing. Kang Min Yoo, Youhyun Shin, Sang-goo Lee |
AAAI | 3 |
| 2019 | A Cross-Sentence Latent Variable Model for Semi-Supervised Text Sequence MatchingabstractWe present a latent variable model for predicting the relationship between a pair of text sequences.Unlike previous auto-encodingbased approaches that consider each sequence separately, our proposed framework utilizes both sequences within a single model by generating a sequence that has a given relationship with a source sequence.We further extend the cross-sentence generating framework to facilitate semi-supervised training.We also define novel semantic constraints that lead the decoder network to generate semantically plausible and diverse sequences.We demonstrate the effectiveness of the proposed model from quantitative and qualitative experiments, while achieving state-of-the-art results on semi-supervised natural language inference and paraphrase identification. Jihun Choi 0002, Taeuk Kim, Sang-goo Lee |
ACL (1) | 3 |
| 2019 | Cell-aware Stacked LSTMs for Modeling SentencesabstractWe propose a method of stacking multiple long short-term memory (LSTM) layers for modeling sentences. In contrast to the conventional stacked LSTMs where only hidden states are fed as input to the next layer, the suggested architecture accepts both hidden and memory cell states of the preceding layer and fuses information from the left and the lower context using the soft gating mechanism of LSTMs. Thus the architecture modulates the amount of information to be delivered not only in horizontal recurrence but also in vertical connections, from which useful features extracted from lower layers are effectively conveyed to upper layers. We dub this architecture Cell-aware Stacked LSTM (CAS-LSTM) and show from experiments that our models bring significant performance gain over the standard LSTMs on benchmark datasets for natural language inference, paraphrase detection, sentiment classification, and machine translation. We also conduct extensive qualitative analysis to understand the internal behavior of the suggested approach. Jihun Choi 0002, Taeuk Kim, Sang-goo Lee |
ACML | 3 |
| 2019 | Don't Just Scratch the Surface: Enhancing Word Representations for Korean with HanjaabstractKang Min Yoo, Taeuk Kim, Sang-goo Lee. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Kang Min Yoo, Taeuk Kim, Sang-goo Lee |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Fashion Attributes-to-Image Synthesis Using Attention-Based Generative Adversarial NetworkabstractIn this paper, we present a method to generate fashion product images those are consistent with a given set of fashion attributes. Since distinct fashion attributes are related to different local sub-regions of a product image, we propose to use generative adversarial network with attentional discriminator. The attribute-attended loss signal from discriminator leads generator to generate more consistent images with given attributes. In addition, we present a generator based on Product-of-Gaussian to encode the composition of fashion attributes in effective way. To verify the proposed model whether it generates consistent image, an oracle attribute classifier is trained and judge the consistency of given attributes and the generated images. Our model significantly outperforms the baseline model in terms of correctness measured by the pre-trained oracle classifier. We show not only qualitative performance but also synthesized images with various combinations of attributes, so we can compare them with baseline model. Hanbit Lee, Sang-goo Lee |
WACV | 2 |
| 2019 | Utterance Generation With Variational Auto-Encoder for Slot Filling in Spoken Language UnderstandingabstractSlot filling must be trained using human-labeled data that are expensive and only a limited amount of labeled utterances are readily available for learning. Data generation methods can help increase the size of the dataset and make variations to the training dataset by means of emerging new instances. We propose a novel labeled utterance generation algorithm to augment training data. Our hypothesis is that words in an utterance can be separated into the two parts, namely, slot values that are instances of slot types and the accompanying contexts. Our model aims to generate utterances that are diverse combinations of slot values and contexts that can appear together. To create various utterances containing a given condition, our deep generative model uses a conditional variational auto-encoder architecture. We conduct experiments on various slot filling datasets, specifically airline travel information systems (ATIS), Snips, and MIT Corpus. A quantitative analysis shows that the application of data augmentation using the proposed model improves the F1 score for slot filling. We also demonstrate that our labeled utterance generation model yields more desirable utterances. Youhyun Shin, Kang Min Yoo, Sang-goo Lee |
IEEE Signal Process. Lett. | 3 |
| 2018 | Learning to Compose Task-Specific Tree StructuresabstractFor years, recursive neural networks (RvNNs) have been shown to be suitable for representing text into fixed-length vectors and achieved good performance on several natural language processing tasks. However, the main drawback of RvNNs is that they require structured input, which makes data preparation and model implementation hard. In this paper, we propose Gumbel Tree-LSTM, a novel tree-structured long short-term memory architecture that learns how to compose task-specific tree structures only from plain text data efficiently. Our model uses Straight-Through Gumbel-Softmax estimator to decide the parent node among candidates dynamically and to calculate gradients of the discrete decision. We evaluate the proposed model on natural language inference and sentiment analysis, and show that our model outperforms or is at least comparable to previous models. We also find that our model converges significantly faster than other models. Jihun Choi 0002, Kang Min Yoo, Sang-goo Lee |
AAAI | 3 |
| 2018 | Automatic Generation of Multiple-Choice Fill-in-the-Blank Question Using Document Embedding
Junghyuk Park, Hyunsoo Cho, Sang-goo Lee |
AIED (2) | 3 |
| 2018 | Slot Filling with Delexicalized Sentence GenerationabstractWe introduce a novel approach that jointly learns slot filling and delexicalized sentence generation. There have been recent attempts to tackle slot filling as a type of sequence labeling problem, with encoder-decoder attention framework. We further improve the framework by training the model to generate delexicalized sentences, in which words according to slot values are replaced with slot labels. Slot filling with delexicalization shows better results compared to models having a single learning objective of filling slots. The proposed method achieves state-of-the-art slot filling performance on ATIS dataset. We experiment different variants of our model and find that delexicalization encourages generalization by sharing weights among the words with same labels and helps the model to further leverage certain linguistic features. Youhyun Shin, Kang Min Yoo, Sang-goo Lee |
INTERSPEECH | 3 |
| 2017 | Partition-Based Clustering with Sliding Windows for Data Streams
Jonghem Youn, Jihun Choi 0002, Junho Shim, Sang-goo Lee |
DASFAA (2) | 4 |
| 2016 | A grapheme-level approach for constructing a Korean morphological analyzer without linguistic knowledgeabstractMorphological analysis is an essential step for processing the Korean language, due to highly agglutinative properties of the language. In this paper, we propose a novel approach for constructing a Korean morphological analyzer that can capture linguistic properties using graphemes as basic processing units. Since our model does not utilize prior linguistic knowledge, the model can be applied to other training corpora with ease. Our model performs morphological analysis through two consecutive sequence labeling tasks: lexical form recovery and part-of-speech tagging. In the lexical form recovery step, morphological changes of an input sentence are restored to the original form. Then in the part-of-speech step, corresponding part-of-speech tags are attached to the recovered form. Experimental results show that our model outperforms previous models which are constructed without prior knowledge. Jihun Choi 0002, Jonghem Youn, Sang-goo Lee |
IEEE BigData | 3 |
| 2016 | Quote Recommendation in Dialogue using Deep Neural NetworkabstractQuotes, or quotations, are well known phrases or sentences that we use for various purposes such as emphasis, elaboration, and humor. In this paper, we introduce a task of recommending quotes which are suitable for given dialogue context and we present a deep learning recommender system which combines recurrent neural network and convolutional neural network in order to learn semantic representation of each utterance and construct a sequence model for the dialog thread. We collected a large set of twitter dialogues with quote occurrences in order to evaluate proposed recommender system. Experimental results show that our approach outperforms not only the other state-of-the-art algorithms in quote recommendation task, but also other neural network based methods built for similar tasks. Hanbit Lee, Yeonchan Ahn, Haejun Lee, Seungdo Ha, Sang-goo Lee |
SIGIR | 5 |
| 2016 | Handling data skew in join algorithms using MapReduce
Jaeseok Myung, Junho Shim, Jongheum Yeon, Sang-goo Lee |
Expert Syst. Appl. | 4 |
| 2016 | Fast and scalable vector similarity joins with MapReduce
Byoungju Yang, Hyun Joon Kim, Junho Shim, Sang-goo Lee |
J. Intell. Inf. Syst. | 5 |
| 2015 | A Fast k-Nearest Neighbor Search Using Query-Specific Signature Selectionabstractk-nearest neighbor (k-NN) search aims at finding k points nearest to a query point in a given dataset. k-NN search is important in various applications, but it becomes extremely expensive in a high-dimensional large dataset. To address this performance issue, locality-sensitive hashing (LSH) is suggested as a method of probabilistic dimension reduction while preserving the relative distances between points. However, the performance of existing LSH schemes is still inconsistent, requiring a large amount of search time in some datasets while the k-NN approximation accuracy is low. In this paper, we target on improving the performance of k-NN search and achieving a consistent k-NN search that performs well in various datasets. In this regard, we propose a novel LSH scheme called Signature Selection LSH (S2LSH). First, we generate a highly diversified signature pool containing signature regions of various sizes and shapes. Then, for a given query point, we rank signature regions of the query and select points in the highly ranked signature regions as k-NN candidates of the query. Extensive experiments show that our approach consistently outperforms the state-of-the-art LSH schemes. Youngki Park, Heasoo Hwang, Sang-goo Lee |
CIKM | 3 |
| 2015 | Constructing compact and effective graphs for recommender systems via node and edge aggregations
Sankeun Lee 0001, Minsuk Kahng, Sang-goo Lee |
Expert Syst. Appl. | 3 |
| 2015 | Reversed CF: A fast collaborative filtering algorithm using a k-nearest neighbor graph
Youngki Park, Sungchan Park, Woosung Jung, Sang-goo Lee |
Expert Syst. Appl. | 4 |
| 2014 | Greedy Filtering: A Scalable Algorithm for K-Nearest Neighbor Graph Construction
Youngki Park, Sungchan Park, Sang-goo Lee, Woosung Jung |
DASFAA (1) | 3 |
| 2014 | A graph-theoretic approach to optimize keyword queries in relational databases
Jaehui Park, Sang-goo Lee |
Knowl. Inf. Syst. | 2 |
| 2013 | A heterogeneous graph-based recommendation simulatorabstractHeterogeneous graph-based recommendation frameworks have flexibility in that they can incorporate various recommendation algorithms and various kinds of information to produce better results. In this demonstration, we present a heterogeneous graph-based recommendation simulator which enables participants to experience the flexibility of a heterogeneous graph-based recommendation method. With our system, participants can simulate various recommendation semantics by expressing the semantics via meaningful paths like User → Movie → User → Movie. The simulator then returns the recommendation results on the fly based on the user-customized semantics using a fast Monte Carlo algorithm. Yeonchan Ahn, Sungchan Park, Sankeun Lee 0001, Sang-goo Lee |
RecSys | 4 |
| 2013 | Editorial - DASFAA2012
Sang-goo Lee, Xiaofang Zhou 0001 |
Data Knowl. Eng. | 1 |
| 2013 | Tridex: A lightweight triple index for relational database-based Semantic Web data management
Seungseok Kang, Junho Shim, Sang-goo Lee |
Expert Syst. Appl. | 3 |
| 2013 | PathRank: Ranking nodes on a heterogeneous graph for flexible hybrid recommender systems
Sankeun Lee 0001, Sungchan Park, Minsuk Kahng, Sang-goo Lee |
Expert Syst. Appl. | 4 |
| 2013 | Pragmatic correlation analysis for probabilistic ranking over relational data
Jaehui Park, Sang-goo Lee |
Expert Syst. Appl. | 2 |
| 2013 | Exploiting inter-operation parallelism for matrix chain multiplication using MapReduce
Jaeseok Myung, Sang-goo Lee |
J. Supercomput. | 2 |
| 2012 | PathRank: a novel node ranking measure on a heterogeneous graph for recommender systemsabstractIn this paper, we present a novel random-walk based node ranking measure, PathRank, which is defined on a heterogeneous graph by extending the Personalized PageRank algorithm. Not only can our proposed measure exploit the semantics behind the different types of nodes and edges in a heterogeneous graph, but also it can emulate various recommendation semantics such as collaborative filtering, content-based filtering, and their combinations. The experimental results show that PathRank can produce more various and effective recommendation results compared to existing approaches. Sankeun Lee 0001, Sungchan Park, Minsuk Kahng, Sang-goo Lee |
CIKM | 4 |
| 2012 | Exploiting paths for entity search in RDF graphsabstractThe field of entity search using Semantic Web (RDF) data has gained more interest recently. In this paper, we propose a probabilistic entity retrieval model for RDF graphs using paths in the graph. Unlike previous work which assumes that all descriptions of an entity are directly linked to the entity node, we assume that an entity can be described with any node that can be reached from the entity node by following paths in the RDF graph. Our retrieval model simulates the generation process of query terms from an entity node by traversing the graph. We evaluate our approach using a standard evaluation framework for entity search. Minsuk Kahng, Sang-goo Lee |
SIGIR | 2 |
| 2011 | Exploiting Correlation to Rank Database Query Results
Jaehui Park, Sang-goo Lee |
DASFAA (2) | 2 |
| 2011 | Random walk based entity ranking on graph for multidimensional recommendationabstractIn many applications, flexibility of recommendation, which is the capability of handling multiple dimensions and various recommendation types, is very important. In this paper, we focus on the flexibility of recommendation and propose a graph-based multidimensional recommendation method. We consider the problem as an entity ranking problem on the graph which is constructed using an implicit feedback dataset (e.g. music listening log), and we adapt Personalized PageRank algorithm to rank entities according to a given query that is represented as a set of entities in the graph. Our model has advantages in that not only can it support the flexibility, but also it can take advantage of exploiting indirect relationships in the graph so that it can perform competitively with the other existing recommendation methods without suffering from the sparsity problem. Sankeun Lee 0001, Sang-il Song, Minsuk Kahng, Sang-goo Lee |
RecSys | 5 |
| 2011 | Keyword search in relational databases
Jaehui Park, Sang-goo Lee |
Knowl. Inf. Syst. | 2 |
| 2010 | Entity-Event Lifelog Ontology Model (EELOM) for LifeLog Ontology Schema DefinitionabstractA set of lifelogs is a dataset that describes a person's life. A high quality and large set of lifelogs is expected to be useful for many applications. Only by integrating the logs that are already available from various devices and creating semantic relationships among them, we can build a useful lifelog ontology that can support various applications. In this paper, we propose Entity-Event Lifelog Ontology Model (EELOM) for designing lifelog ontology that can be used for integrating various available data sources. Sankeun Lee 0001, Gihyun Gong, Sang-goo Lee |
APWeb | 3 |
| 2010 | Applying Taxonomic Knowledge and Semantic Collaborative Filtering to Personalized Search: A Bayesian Belief Network Based ApproachabstractKeyword-based search exploits the exact match between the index terms of a query and documents. Thus, some documents, although they are relevant to the given query, may not be returned to users unless the documents include the index terms of the query. Some search engines use the authority of documents, which is derived from the links of documents, to help keyword-based search provide more accurate search results. However, unlike the Web documents, if the links between documents do not exist, it is difficult to exploit the authority for ranking documents. In this paper, our goals are to derive the implicit authority of documents that do not have explicit links through semantic collaborative filtering (SCF), and to retrieve documents that are semantically related to the given query. To achieve these goals, we represent users' preferences, queries and documents with their corresponding concepts by extending a Bayesian belief network. It is because the Bayesian belief network provides a clear formalism for mapping the users' preferences, queries and documents to their corresponding concepts. The concepts are extracted from a taxonomic knowledgebase such as the Open Directory Project Web directory. In our experiment, we have shown that the extended Bayesian belief network using taxonomic knowledge outperforms the conventional approaches for personalized search. Jae-Won Lee, Han-Joon Kim, Sang-goo Lee |
APWeb | 3 |
| 2010 | A General Maturity Model and Reference Architecture for SaaS Service
Seungseok Kang, Jaeseok Myung, Jongheum Yeon, Seong-wook Ha, Taehyung Cho, Ji-man Chung, Sang-goo Lee |
DASFAA (2) | 7 |
| 2010 | An Efficient Similarity Join Algorithm with Cosine Similarity Predicate
Jaehui Park, Junho Shim, Sang-goo Lee |
DEXA (2) | 4 |
| 2010 | Ranking Objects Based on Attribute Value Correlation
Jaehui Park, Sang-goo Lee |
DEXA (2) | 2 |
| 2010 | Si-Fi: interactive similar item finder
Inbeom Hwang, Minsuk Kahng, Sung Eun Park, Jinwook Seo, Sang-goo Lee |
SIGIR | 5 |
| 2010 | Conceptual collaborative filtering recommendation: A probabilistic learning approach
Jae-Won Lee, Han-Joon Kim, Sang-goo Lee |
Neurocomputing | 3 |
| 2008 | Exploiting Attribute-Wise Distribution of Keywords and Category Dependent Attributes for E-Catalog Classification
Young-gon Kim, Taehee Lee 0001, Sang-goo Lee, Jong-Heung Park |
ICIC (1) | 3 |
| 2007 | CST-Trees: Cache Sensitive T-Trees
Ig-hoon Lee, Junho Shim, Sang-goo Lee, Jonghoon Chun |
DASFAA | 3 |
| 2006 | A Snappy B+-Trees Index Reconstruction for Main-Memory Storage Systems
Ig-hoon Lee, Junho Shim, Sang-goo Lee |
ICCSA (1) | 3 |
| 2005 | A Practical Ontology for Product Information Management
Dongkyu Kim, Sang-goo Lee, Jonghoon Chun, Zoonky Lee, Heungsun Park |
iiWAS | 2 |
| 2004 | Using Relational Database Constraints to Design Materialized Views in Data Warehouses
Taehee Lee 0001, Jae-Young Chang, Sang-goo Lee |
APWeb | 3 |
| 2004 | Ontological Approaches to Enterprise Applications
Dongkyu Kim, Yuan-Chi Chang, Juhnyoung Lee, Sang-goo Lee |
ER | 4 |
| 2004 | An Intelligent Information System for Organizing Online Text Documents
Han-Joon Kim, Sang-goo Lee |
Knowl. Inf. Syst. | 2 |
| 2003 | The Knowledge Modeling for Chronic Urticaria Assessment in Clinical Decision Support System with PDA
Miyoung Kwak, Seung-Bin Han, Gyochang Kim, Jinwook Choi, Jonghoon Chun, Kangsun Lee, Heeseung Bom, Sang-goo Lee |
AMIA | 8 |
| 2003 | Building topic hierarchy based on fuzzy relations
Han-Joon Kim, Sang-goo Lee |
Neurocomputing | 2 |
| 2001 | A Method for Processing Boolean Queries Using a Result Cache
Jae-Heon Cheong, Sang-goo Lee, Jonghoon Chun |
DEXA | 2 |
| 2001 | Application of Information Technology: A DBMS-based Medical Teleconferencing SystemabstractThis article presents the design of a medical teleconferencing system that is integrated with a multimedia patient database and incorporates easy-to-use tools and functions to effectively support collaborative work between physicians in remote locations. The design provides a virtual workspace that allows physicians to collectively view various kinds of patient data. By integrating the teleconferencing function into this workspace, physicians are able to conduct conferences using the same interface and have real-time access to the database during conference sessions. The authors have implemented a prototype based on this design. The prototype uses a high-speed network test bed and a manually created substitute for the integrated patient database. Jonghoon Chun, Han-Joon Kim, Sang-goo Lee, Jinwook Choi, Hanik Cho |
J. Am. Medical Informatics Assoc. | 3 |
| 2000 | Implementation of Mobile Computing System in Clinical Environment: MobileNurse™
Sookyung Hyun, Jinwook Choi, Jonghoon Chun, Sang-goo Lee, Daihee Kim |
AMIA | 4 |
| 2000 | A Semi-Supervised Document Clustering Technique for Information OrganizationabstractThis paper discusses a new type of semi-supervised docu-ment clustering that uses partial supervision to partition a large set of documents. Most clustering methods organizes documents into groups based only on similarity measures. Unfortunately, the traditional approaches to document clus-tering are often unable to correctly discern structural details hidden within the document corpus because their algorithms inherently strongly depend on the document themselves and their similarity to each other. In this paper, we attempt to isolate more semantically coherent clusters by employing the domain-specific knowledge provided by a document analyst. By using external human knowledge to guide the clustering mechanism with some flexibility when creating the clusters, clustering efficiency can be considerably enhanced. As a ba-sic clustering strategy, we use a variant of complete-linkage agglomerative hierarchical clustering, and develop the con-cepts (or seeds) of requested clusters by exploiting user-relevance feedback. Although the proposed method is slow when applied to large document collection, it yields higher quality clusters. Through experiments using the Reuters-21578 corpus, we show that the proposed method outper-forms unsupervised clustering method. Han-Joon Kim, Sang-goo Lee |
CIKM | 2 |
| 2000 | Identifying relevant constraints for semantic query optimization
Sang-goo Lee, Lawrence J. Henschen, Jonghun Chun, Taehee Lee 0001 |
Inf. Softw. Technol. | 1 |
| 1999 | A New Flash Memory Management for Flash Storage SystemabstractProposes a new way of managing flash memory space for flash memory-specific file systems based on a log-structured file system. Flash memory has attractive features such as non-volatility and fast I/O speed, but it also suffers from an inability to update in place, and limited usage cycles. These drawbacks require many changes to conventional storage (file) management techniques. Our focus is on lowering the cleaning cost and evenly utilizing flash memory cells while maintaining a balance between these two often-conflicting goals. The cleaning efficiency is enhanced by dynamically separating cold data and non-cold data. The second goal, cycle leveling, is achieved to the degree where the maximum difference between erase cycles is below the error range of the hardware. Simulation results show that the proposed method has a significant benefit over naive methods: a maximum of 35% reduction in the cleaning cost with evenly-spread writes across segments. Han-Joon Kim, Sang-goo Lee |
COMPSAC | 2 |
| 1999 | Extended Conditions for Answering an Aggregate Query Using Materialized Views
Jae-Young Chang, Sang-goo Lee |
Inf. Process. Lett. | 2 |
| 1998 | Query Reformulation Using Materialized Views in Data Warehousing EnvironmentabstractMaterialized views offer opportunities for significant performance gain in query evaluation by providing fast access to pi-e-computed data.The question of when and how to use a materialized view in processing a given query is a difficult one attracting a significant amount of research.In previous works, only materialized views whose relations are contained in those of a query have been used and, as a result, certain potentially useful materialized views were excluded from consideration.Proposed in this paper are new ways of utilizing materialized views in answering a query with aggregation operations; Views including relations not referred to in the given query are utilized.We identify the conditions where a materialized view can be used in reformulating a query.Also presented are algorithms to find the most efficient reformulated query.The proposed conditions and corresponding algorithms provide significant and practical performance improvements to the data warehousing environment.' This work was supported in Jae-Young Chang, Sang-goo Lee |
DOLAP | 2 |
| 1997 | A Telemedicine System Using B-ISDN
Jinwook Choi, Young-Han Kim 0002, Sang-goo Lee, Jaeok Lee, Byung Hee Oh, Hanik Cho |
AMIA | 3 |
| 1997 | An optimization of disjunctive queries: union-pushdownabstractMost previous works on query optimization techniques deal with conjunctive queries only because the queries with disjunctive predicates are complex to optimize. Hence, for disjunctive queries, query optimizers based on these techniques generate plans using rather simple methods such as CNF- and DNF-based optimization. However, the plans generated by these methods perform extremely poorly for certain types of queries. The authors propose new query optimization method, union-pushdown, for disjunctive queries. This method is composed of four phases, and each phase utilizes some advantageous techniques of CNF- and DNF-based methods. They analyze the performance of the union-pushdown plan against those of conventional plans and show that union-pushdown can be applied to various disjunctive query types without performance degradation. Jae-Young Chang, Sang-goo Lee |
COMPSAC | 2 |
| 1996 | A Logic Database System with Extended FunctionalityabstractWe present the architecture and design details of a logic database system that extends functionality in a number of ways. The system supports indefinite (non Horn) facts in extensional database (EDB) and existential quantification in intensional database (IDB) rules. Semantic query optimization is integrated as part of the query processing module. The system also deals with efficient rule management for IDB and integrity constraints (IC). The proposed system supports the following database: the EDB consists of a set of positive ground formula, not necessarily Horn; the IDB includes rules that are Horn and may have existential quantifiers. The IC includes range restricted Horn clauses with no existential quantification. Sang-goo Lee, Dong-Hoon Choi |
COMPSAC | 1 |
| 1995 | On Foundations of Constraint Optimization
Sang-goo Lee |
DASFAA | 1 |
| 1991 | Semantic Query Reformulation in Deductive DatabasesabstractA method is proposed for identifying relevant integrity constraints (ICs) for queries involving joins/unions of base relations and defined relations by use of graphs. The method does not rely on heavy preprocessing or redundancy. To effectively select those ICs that are relevant to a given query, the relationship between the predicates in the query is identified using an AND/OR tree where an AND mode represents a join operation and an OR node represents a union operation. Ways of collecting ICs are described that are not directly related to the query but can be useful in query optimization.> Sang-goo Lee, Lawrence J. Henschen, Ghassan Z. Qadah |
ICDE | 1 |
| 1990 | Semantic and structural query reformulation for efficient manipulation of very large knowledge basesabstractThe authors present a framework for a knowledge base system that supports complex objects and two-level rules. By assuming that the portion of a rule base that is related to a query is small enough to fit in main memory, the bottleneck of the inference stage is not in unifying or managing complex objects but in identifying relevant rules for the query. However, efficient storage and manipulation of complex objects is critical in the physical database access stage where the fact base consists of large number of general objects. Consequently, the system has been divided into two virtually independent stages. An obvious, application of a two-level rule base is in semantic query optimization, where the integrity constraints will be the semantic rules and application of restrictions from them is optional. By supplying a number of special system predicates, the two-level rule base can be used to control the activities of the knowledge base. The two levels of rules naturally map to rules and meta rules in artificial intelligence applications.> Sang-goo Lee |
COMPSAC | 1 |