EDBT 2026 Demo / reviewers in the wild / expert
SangKeun Lee 0001
dblp:73/3458-1
· DBLP profile ↗
77ranked-venue papers
12as first author
25since 2021 · last 2026
0000-0002-6249-8217ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 2 first-author · 23 since 2021Databases, data management, data science and information retrieval · 32 · 9 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-authorSystems, architecture and hardware · 6 · 2 first-author · 1 since 2021Computer networks · 5Theory of computation · 3Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters' Lack of KnowledgeabstractSimulating human-like Theory of Mind (ToM) has been a longstanding problem in natural language processing (NLP).To address this, existing works introduce a reasoning step of event hiding (a.k.a.perspective-taking), where events unknown to a character are removed before question answering.However, resorting to event hiding for ToM reasoning presents a performance degradation issue due to the strict output format constraints involved in event hiding.To mitigate this issue, we propose generating perspective-taking outputs as free-form explanations without event hiding, but this poses a notable yet underexplored challenge: LLMs need to inhibit responses to events unknown to characters, because the absence of event hiding exposes LLMs to these events throughout reasoning.To address this challenge, we hypothesize and empirically verify that LLMs can achieve such inhibition if a character's lack of knowledge about events is made explicit during reasoning.Based on this finding, we introduce PICTURE, a new prompting method that enables LLMs to generate a character's lack of knowledge within free-form CoT.Experimental results show that PICTURE outperforms existing prompting methods by an average of 7.3% on false-belief tasks. 1 Eojin Jeon, SangKeun Lee 0001 |
ACL (1) | 2 |
| 2025 | Forward Knows Efficient Backward Path: Saliency-Guided Memory-Efficient Fine-tuning of Large Language ModelsabstractFine-tuning is widely recognized as a crucial process for aligning large language models (LLMs) with human intentions.However, the substantial memory requirements associated with fine-tuning pose a significant barrier to extending the applicability of LLMs.While parameter-efficient fine-tuning can be a promising approach by reducing trainable parameters, intermediate activations still need to be cached to compute gradients during the backward pass, thereby limiting overall memory efficiency.In this work, we propose Saliency-Guided Gradient Flow (SAGE), a memoryefficient fine-tuning method designed to minimize the memory specifically associated with cached intermediate activations.The key strategy is to selectively cache activations based on their saliency during the forward pass and then use these activations for the backward pass.This process transforms the dense backward pass into a sparse one, thereby enhancing memory efficiency.To verify whether SAGE can serve as an efficient alternative for fine-tuning, we conduct comprehensive experiments across diverse fine-tuning scenarios and setups.The experimental results show that SAGE substantially improves memory efficiency without a significant loss in accuracy, highlighting its broad value in real-world applications 1 . Yeachan Kim, SangKeun Lee 0001 |
ACL (1) | 2 |
| 2025 | Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in KoreanabstractWe introduce the Korean Grammar Evaluation BenchMark (KoGEM), designed to assess the linguistic competence of LLMs and humans in Korean.KoGEM consists of 1.5k multiplechoice QA pairs covering five main categories and 16 subcategories.The zero-shot evaluation of 27 LLMs of various sizes and types reveals that while LLMs perform remarkably well on straightforward tasks requiring primarily definitional knowledge, they struggle with tasks that demand the integration of realworld experiential knowledge, such as phonological rules and pronunciation.Furthermore, our in-depth analysis suggests that incorporating such experiential knowledge could enhance the linguistic competence of LLMs.With Ko-GEM, we not only highlight the limitations of current LLMs in linguistic competence but also uncover hidden facets of LLMs in linguistic competence, paving the way for enhancing comprehensive language understanding.Our code and dataset are available at Sung-Ho Kim 0010, Nayeon Kim 0002, Taehee Jeon, SangKeun Lee 0001 |
ACL (1) | 4 |
| 2025 | Curriculum Debiasing: Toward Robust Parameter-Efficient Fine-Tuning Against Dataset BiasesabstractParameter-efficient fine-tuning (PEFT) addresses the memory footprint issue of full fine-tuning by modifying only a subset of model parameters. However, on datasets exhibiting spurious correlations, we observed that PEFT slows down the model’s convergence on unbiased examples, while the convergence on biased examples remains fast. This leads to the model’s overfitting on biased examples, causing significant performance degradation in out-of-distribution (OOD) scenarios. Traditional debiasing methods mitigate this issue by emphasizing unbiased examples during training but often come at the cost of in-distribution (ID) performance drops. To address this trade-off issue, we propose a curriculum debiasing framework that presents examples in a biased-to-unbiased order. Our framework initially limits the model’s exposure to unbiased examples, which are harder to learn, allowing it to first establish a foundation on easier-to-converge biased examples. As training progresses, we gradually increase the proportion of unbiased examples in the training set, guiding the model away from reliance on spurious correlations. Compared to the original PEFT methods, our method accelerates convergence on unbiased examples by approximately twofold and improves ID and OOD performance by 1.2% and 8.0%, respectively. Yeachan Kim, Wing-Lam Mok, SangKeun Lee 0001 |
ACL (1) | 4 |
| 2025 | Incorporating Domain Knowledge into Materials TokenizationabstractWhile language models are increasingly utilized in materials science, typical models rely on frequency-centric tokenization methods originally developed for natural language processing.However, these methods frequently produce excessive fragmentation and semantic loss, failing to maintain the structural and semantic integrity of material concepts.To address this issue, we propose MATTER, a novel tokenization approach that integrates material knowledge into tokenization.Based on MatDetector trained on our materials knowledge base and a re-ranking method prioritizing material concepts in token merging, MATTER maintains the structural integrity of identified material concepts and prevents fragmentation during tokenization, ensuring their semantic meaning remains intact.The experimental results demonstrate that MATTER outperforms existing tokenization methods, achieving an average performance gain of 4% and 2% in the generation and classification tasks, respectively.These results underscore the importance of domain knowledge for tokenization strategies in scientific text processing. 1 Yerim Oh, Jun-Hyung Park, Sung-Ho Kim 0010, SangKeun Lee 0001 |
ACL (1) | 5 |
| 2025 | Connecting the Knowledge Dots: Retrieval-augmented Knowledge Connection for Commonsense ReasoningabstractWhile large language models (LLMs) have achieved remarkable performance across various natural language processing (NLP) tasks, LLMs exhibit a limited understanding of commonsense reasoning due to the necessity of implicit knowledge that is rarely expressed in text.Recently, retrieval-augmented language models (RALMs) have enhanced their commonsense reasoning ability by incorporating background knowledge from external corpora.However, previous RALMs overlook the implicit nature of commonsense knowledge, potentially leading to the retrieved documents not directly contain information needed to answer questions.In this paper, we propose Retrieval-augmented knowledge Connection, RECONNECT, which transforms indirectly relevant documents into a direct explanation to answer the given question.To this end, we extract relevant knowledge from various retrieved document subsets and aggregate them into a direct explanation.Experimental results show that RECONNECT outperforms state-of-the-art (SOTA) baselines, achieving improvements of +2.0% and +4.6% average accuracy on in-domain (ID) and outof-domain (OOD) benchmarks, respectively 1 . Soyeon Bak, Minju Hong, Songha Kim, Tae-Eui Kam, SangKeun Lee 0001 |
EMNLP | 7 |
| 2025 | Bridging the Gap Between Molecule and Textual Descriptions via Substructure-aware AlignmentabstractMolecule and text representation learning has gained increasing interest due to its potential for enhancing the understanding of chemical information.However, existing models often struggle to capture subtle differences between molecules and their descriptions, as they lack the ability to learn fine-grained alignments between molecular substructures and chemical phrases.To address this limitation, we introduce MolBridge, a novel molecule-text learning framework based on substructure-aware alignments.Specifically, we augment the original molecule-description pairs with additional alignment signals derived from molecular substructures and chemical phrases.To effectively learn from these enriched alignments, Mol-Bridge employs substructure-aware contrastive learning, coupled with a self-refinement mechanism that filters out noisy alignment signals.Experimental results show that MolBridge effectively captures fine-grained correspondences and outperforms state-of-the-art baselines on a wide range of molecular benchmarks, underscoring the importance of substructure-aware alignment in molecule-text learning. Hyuntae Park, Yeachan Kim, SangKeun Lee 0001 |
EMNLP | 3 |
| 2025 | Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
Nayeon Kim 0002, Eojin Jeon, Jun-Hyung Park, SangKeun Lee 0001 |
PAKDD (3) | 4 |
| 2025 | Continual debiasing: A bias mitigation framework for natural language understanding systems
Jun-Hyung Park, SangKeun Lee 0001 |
Expert Syst. Appl. | 4 |
| 2024 | SparseFlow: Accelerating Transformers by Sparsifying Information FlowsabstractTransformers have become the de-facto standard for natural language processing.However, dense information flows within transformers pose significant challenges for real-time and resource-constrained devices, as computational complexity grows quadratically with sequence length.To counteract such dense information flows, we propose SPARSEFLOW, a novel efficient method designed to sparsify the dense pathways of token representations across all transformer blocks.To this end, SPARSEFLOW parameterizes the information flows linking token representations to transformer blocks.These parameterized information flows are optimized to be sparse, allowing only the salient information to pass through into the blocks.To validate the efficacy of SPARSEFLOW, we conduct comprehensive experiments across diverse benchmarks (understanding and generation), scales (ranging from millions to billions), architectures (including encoders, decoders, and seq-to-seq models), and modalities (such as language-only and vision-language).The results convincingly demonstrate that sparsifying the dense information flows leads to substantial speedup gains without compromising task accuracy.For instance, SPARSEFLOW reduces computational costs by half on average, without a significant loss in accuracy 1 . Yeachan Kim, SangKeun Lee 0001 |
ACL (1) | 2 |
| 2024 | Towards Robust and Generalized Parameter-Efficient Fine-Tuning for Noisy Label LearningabstractParameter-efficient fine-tuning (PEFT) has enabled the efficient optimization of cumbersome language models in real-world settings.However, as datasets in such environments often contain noisy labels that adversely affect performance, PEFT methods are inevitably exposed to noisy labels.Despite this challenge, the adaptability of PEFT to noisy environments remains underexplored.To bridge this gap, we investigate various PEFT methods under noisy labels.Interestingly, our findings reveal that PEFT has difficulty in memorizing noisy labels due to its inherently limited capacity, resulting in robustness.However, we also find that such limited capacity simultaneously makes PEFT more vulnerable to interference of noisy labels, impeding the learning of clean samples.To address this issue, we propose Clean Routing (CleaR), a novel routing-based PEFT approach that adaptively activates PEFT modules.In CleaR, PEFT modules are preferentially exposed to clean data while bypassing the noisy ones, thereby minimizing the noisy influence.To verify the efficacy of CleaR, we perform extensive experiments on diverse configurations of noisy labels.The results convincingly demonstrate that CleaR leads to substantially improved performance in noisy environments 1 . Yeachan Kim, SangKeun Lee 0001 |
ACL (1) | 3 |
| 2024 | Mentor-KD: Making Small Language Models Better Multi-step ReasonersabstractLarge Language Models (LLMs) have displayed remarkable performances across various complex tasks by leveraging Chain-of-Thought (CoT) prompting.Recently, studies have proposed a Knowledge Distillation (KD) approach, reasoning distillation, which transfers such reasoning ability of LLMs through fine-tuning language models of multi-step rationales generated by LLM teachers.However, they have inadequately considered two challenges regarding insufficient distillation sets from the LLM teacher model, in terms of 1) data quality and 2) soft label provision.In this paper, we propose Mentor-KD, which effectively distills the multi-step reasoning capability of LLMs to smaller LMs while addressing the aforementioned challenges.Specifically, we exploit a mentor, intermediate-sized task-specific fine-tuned model, to augment additional CoT annotations and provide soft labels for the student model during reasoning distillation.We conduct extensive experiments and confirm Mentor-KD's effectiveness across various models and complex reasoning tasks 1 . Hojae Lee, SangKeun Lee 0001 |
EMNLP | 3 |
| 2024 | MolTRES: Improving Chemical Language Representation Learning for Molecular Property PredictionabstractChemical representation learning has gained increasing interest due to the limited availability of supervised data in fields such as drug and materials design.This interest particularly extends to chemical language representation learning, which involves pre-training Transformers on SMILES sequences -textual descriptors of molecules.Despite its success in molecular property prediction, current practices often lead to overfitting and limited scalability due to early convergence.In this paper, we introduce a novel chemical language representation learning framework, called MolTRES, to address these issues.MolTRES incorporates generator-discriminator training, allowing the model to learn from more challenging examples that require structural understanding.In addition, we enrich molecular representations by transferring knowledge from scientific literature by integrating external materials embedding.Experimental results show that our models outperform existing state-of-the-art models on popular molecular property prediction tasks. github.com/irishev/MolTRES Jun-Hyung Park, Yeachan Kim, Hyuntae Park, SangKeun Lee 0001 |
EMNLP | 5 |
| 2023 | Dynamic Structure Pruning for Compressing CNNsabstractStructure pruning is an effective method to compress and accelerate neural networks. While filter and channel pruning are preferable to other structure pruning methods in terms of realistic acceleration and hardware compatibility, pruning methods with a finer granularity, such as intra-channel pruning, are expected to be capable of yielding more compact and computationally efficient networks. Typical intra-channel pruning methods utilize a static and hand-crafted pruning granularity due to a large search space, which leaves room for improvement in their pruning performance. In this work, we introduce a novel structure pruning method, termed as dynamic structure pruning, to identify optimal pruning granularities for intra-channel pruning. In contrast to existing intra-channel pruning methods, the proposed method automatically optimizes dynamic pruning granularities in each layer while training deep neural networks. To achieve this, we propose a differentiable group learning method designed to efficiently learn a pruning granularity based on gradient-based learning of filter groups. The experimental results show that dynamic structure pruning achieves state-of-the-art pruning performance and better realistic acceleration on a GPU compared with channel pruning. In particular, it reduces the FLOPs of ResNet50 by 71.85% without accuracy degradation on the ImageNet dataset. Our code is available at https://github.com/irishev/DSP. Jun-Hyung Park, Yeachan Kim, Joon-Young Choi, SangKeun Lee 0001 |
AAAI | 5 |
| 2023 | SMoP: Towards Efficient and Effective Prompt Tuning with Sparse Mixture-of-PromptsabstractPrompt tuning has emerged as a successful parameter-efficient alternative to the full finetuning of language models.However, prior works on prompt tuning often utilize long soft prompts of up to 100 tokens to improve performance, overlooking the inefficiency associated with extended inputs.In this paper, we propose a novel prompt tuning method SMoP (Sparse Mixture-of-Prompts) that utilizes short soft prompts for efficient training and inference while maintaining performance gains typically induced from longer soft prompts.To achieve this, SMoP employs a gating mechanism to train multiple short soft prompts specialized in handling different subsets of the data, providing an alternative to relying on a single long soft prompt to cover the entire data.Experimental results demonstrate that SMoP outperforms baseline methods while reducing training and inference costs.We release our code at https://github.com/jyjohnchoi/SMoP. Joon-Young Choi, Jun-Hyung Park, Wing-Lam Mok, SangKeun Lee 0001 |
EMNLP | 5 |
| 2023 | Improving Bias Mitigation through Bias Experts in Natural Language UnderstandingabstractBiases in the dataset often enable the model to achieve high performance on in-distribution data, while poorly performing on out-ofdistribution data.To mitigate the detrimental effect of the bias on the networks, previous works have proposed debiasing methods that down-weight the biased examples identified by an auxiliary model, which is trained with explicit bias labels.However, finding a type of bias in datasets is a costly process.Therefore, recent studies have attempted to make the auxiliary model biased without the guidance (or annotation) of bias labels, by constraining the model's training environment or the capability of the model itself.Despite the promising debiasing results of recent works, the multiclass learning objective, which has been naively used to train the auxiliary model, may harm the bias mitigation effect due to its regularization effect and competitive nature across classes.As an alternative, we propose a new debiasing framework that introduces binary classifiers between the auxiliary model and the main model, coined bias experts.Specifically, each bias expert is trained on a binary classification task derived from the multi-class classification task via the One-vs-Rest approach.Experimental results demonstrate that our proposed strategy improves the bias identification ability of the auxiliary model.Consequently, our debiased model consistently outperforms the state-ofthe-art on various challenge datasets.1 Eojin Jeon, Juhyeong Park, Yeachan Kim, Wing-Lam Mok, SangKeun Lee 0001 |
EMNLP | 6 |
| 2023 | Leap-of-Thought: Accelerating Transformers via Dynamic Token RoutingabstractComputational inefficiency in transformers has been a long-standing challenge, hindering the deployment in resource-constrained or realtime applications.One promising approach to mitigate this limitation is to progressively remove less significant tokens, given that the sequence length strongly contributes to the inefficiency.However, this approach entails a potential risk of losing crucial information due to the irrevocable nature of token removal.In this paper, we introduce Leap-of-Thought (LoT), a novel token reduction approach that dynamically routes tokens within layers.Unlike previous work that irrevocably discards tokens, LoT enables tokens to 'leap' across layers.This ensures that all tokens remain accessible in subsequent layers while reducing the number of tokens processed within layers.We achieve this by pairing the transformer with dynamic token routers, which learn to selectively process tokens essential for the task.Evaluation results clearly show that LoT achieves substantial improvement on computational efficiency.Specifically, LoT attains up to 25× faster inference time without a significant loss in accuracy 1 . Yeachan Kim, Jun-Hyung Park, SangKeun Lee 0001 |
EMNLP | 5 |
| 2023 | DIVE: Towards Descriptive and Diverse Visual Commonsense GenerationabstractTowards human-level visual understanding, visual commonsense generation has been introduced to generate commonsense inferences beyond images.However, current research on visual commonsense generation has overlooked an important human cognitive ability: generating descriptive and diverse inferences.In this work, we propose a novel visual commonsense generation framework, called DIVE, which aims to improve the descriptiveness and diversity of generated inferences.DIVE involves two methods, generic inference filtering and contrastive retrieval learning, which address the limitations of existing visual commonsense resources and training objectives.Experimental results verify that DIVE outperforms state-ofthe-art models for visual commonsense generation in terms of both descriptiveness and diversity, while showing a superior quality in generating unique and novel inferences.Notably, DIVE achieves human-level descriptiveness and diversity on Visual Commonsense Graphs.Furthermore, human evaluations confirm that DIVE aligns closely with human judgments on descriptiveness and diversity 1 . Jun-Hyung Park, Hyuntae Park, Youjin Kang, Eojin Jeon, SangKeun Lee 0001 |
EMNLP | 5 |
| 2023 | Examining Consistency of Visual Commonsense Reasoning based on Person GroundingabstractHuiju Kim, Youjin Kang, SangKeun Lee. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Huiju Kim, Youjin Kang, SangKeun Lee 0001 |
IJCNLP (1) | 3 |
| 2022 | Break it Down into BTS: Basic, Tiniest Subword Units for KoreanabstractWe introduce Basic, Tiniest Subword (BTS) units for the Korean language, which are inspired by the invention principle of Hangeul, the Korean writing system.Instead of relying on 51 Korean consonant and vowel letters, we form the letters from BTS units by adding strokes or combining them.To examine the impact of BTS units on Korean language processing, we develop a novel BTSbased word embedding framework that is readily applicable to various models.Our experiments reveal that BTS units significantly improve the performance of Korean word embedding on all intrinsic and extrinsic tasks in our evaluation.In particular, BTS-based word embedding outperforms the state-of-theart Korean word embedding by 11.8% in word analogy.We further investigate the unique advantages provided by BTS units through indepth analysis.Our code is available at https: //github.com/irishev/BTS. Nayeon Kim 0002, Jun-Hyung Park, Joon-Young Choi, Eojin Jeon, Youjin Kang, SangKeun Lee 0001 |
EMNLP | 6 |
| 2022 | Tutoring Helps Students Learn Better: Improving Knowledge Distillation for BERT with Tutor NetworkabstractPre-trained language models have achieved remarkable successes in natural language processing tasks, coming at the cost of increasing model size.To address this issue, knowledge distillation (KD) has been widely applied to compress language models.However, typical KD approaches for language models have overlooked the difficulty of training examples, suffering from incorrect teacher prediction transfer and sub-efficient training.In this paper, we propose a novel KD framework, Tutor-KD, which improves the distillation effectiveness by controlling the difficulty of training examples during pre-training.We introduce a tutor network that generates samples that are easy for the teacher but difficult for the student, with training on a carefully designed policy gradient method.Experimental results show that Tutor-KD significantly and consistently outperforms the state-of-the-art KD methods with variously sized student models on the GLUE benchmark, demonstrating that the tutor can effectively generate training examples for the student 1 . Jun-Hyung Park, Wing-Lam Mok, Joon-Young Choi, SangKeun Lee 0001 |
EMNLP | 6 |
| 2022 | Efficient Pre-training of Masked Language Model via Concept-based Curriculum MaskingabstractMasked language modeling (MLM) has been widely used for pre-training effective bidirectional representations, but incurs substantial training costs.In this paper, we propose a novel concept-based curriculum masking (CCM) method to efficiently pre-train a language model.CCM has two key differences from existing curriculum learning approaches to effectively reflect the nature of MLM.First, we introduce a carefully-designed linguistic difficulty criterion that evaluates the MLM difficulty of each token.Second, we construct a curriculum that gradually masks words related to the previously masked words by retrieving a knowledge graph.Experimental results show that CCM significantly improves pre-training efficiency.Specifically, the model trained with CCM shows comparative performance with the original BERT on the General Language Understanding Evaluation benchmark at half of the training cost.Code is available at https://github.com/KoreaMGLEE/Concept- based-curriculum-masking. Jun-Hyung Park, Kang-Min Kim, SangKeun Lee 0001 |
EMNLP | 5 |
| 2022 | Examining the impact of adaptive convolution on natural language understanding
Jun-Hyung Park, Byung-Ju Choi, SangKeun Lee 0001 |
Expert Syst. Appl. | 3 |
| 2022 | Quantized Sparse Training: A Unified Trainable Framework for Joint Pruning and Quantization in DNNsabstractDeep neural networks typically have extensive parameters and computational operations. Pruning and quantization techniques have been widely used to reduce the complexity of deep models. Both techniques can be jointly used for realizing significantly higher compression ratios. However, separate optimization processes and difficulties in choosing the hyperparameters limit the application of both the techniques simultaneously. In this study, we propose a novel compression framework, termed as quantized sparse training, that prunes and quantizes networks jointly in a unified training process. We integrate pruning and quantization into a gradient-based optimization process based on the straight-through estimator. Quantized sparse training enables us to simultaneously train, prune, and quantize a network from scratch. The empirical results validate the superiority of the proposed methodology over the recent state-of-the-art baselines with respect to both the model size and accuracy. Specifically, quantized sparse training achieves a 135 KB model size in the case of VGG16, without any accuracy degradation, which is 40% of the model size feasible based on the state-of-the-art pruning and quantization approach. Jun-Hyung Park, Kang-Min Kim, SangKeun Lee 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2021 | Handling Out-Of-Vocabulary Problem in Hangeul Word EmbeddingsabstractWord embedding is considered an essential factor in improving the performance of various Natural Language Processing (NLP) models.However, it is hardly applicable in realworld datasets as word embedding is generally studied with a well-refined corpus.Notably, in Hangeul (Korean writing system), which has a unique writing system, various kinds of Out-Of-Vocabulary (OOV) appear from typos.In this paper, we propose a robust Hangeul word embedding model against typos, while maintaining high performance.The proposed model utilizes a Convolutional Neural Network (CNN) architecture with a channel attention mechanism that learns to infer the original word embeddings.The model train with a dataset that consists of a mix of typos and correct words.To demonstrate the effectiveness of the proposed model, we conduct three kinds of intrinsic and extrinsic tasks.While the existing embedding models fail to maintain stable performance as the noise level increases, the proposed model shows stable performance. Ohjoon Kwon, Soo-Ryeon Lee, SangKeun Lee 0001 |
EACL | 5 |
| 2020 | Adaptive Compression of Word EmbeddingsabstractDistributed representations of words have been an indispensable component for natural language processing (NLP) tasks.However, the large memory footprint of word embeddings makes it challenging to deploy NLP models to memory-constrained devices (e.g., selfdriving cars, mobile devices).In this paper, we propose a novel method to adaptively compress word embeddings.We fundamentally follow a code-book approach that represents words as discrete codes such as (8, 5, 2, 4).However, unlike prior works that assign the same length of codes to all words, we adaptively assign different lengths of codes to each word by learning downstream tasks.The proposed method works in two steps.First, each word directly learns to select its code length in an end-to-end manner by applying the Gumbel-softmax tricks.After selecting the code length, each word learns discrete codes through a neural network with a binary constraint.To showcase the general applicability of the proposed method, we evaluate the performance on four different downstream tasks.Comprehensive evaluation results clearly show that our method is effective and makes the highly compressed word embeddings without hurting the task accuracy.Moreover, we show that our model assigns word to each code-book by considering the significance of tasks. Yeachan Kim, Kang-Min Kim, SangKeun Lee 0001 |
ACL | 3 |
| 2020 | Representation Learning for Unseen Words by Bridging Subwords to Semantic NetworksabstractPre-trained word embeddings are widely used in various fields. However, the coverage of pre-trained word embeddings only includes words that appeared in corpora where pre-trained embeddings are learned. It means that the words which do not appear in training corpus are ignored in tasks, and it could lead to the limited performance of neural models. In this paper, we propose a simple yet effective method to represent out-of-vocabulary (OOV) words. Unlike prior works that solely utilize subword information or knowledge, our method makes use of both information to represent OOV words. To this end, we propose two stages of representation learning. In the first stage, we learn subword embeddings from the pre-trained word embeddings by using an additive composition function of subwords. In the second stage, we map the learned subwords into semantic networks (e.g., WordNet). We then re-train the subword embeddings by using lexical entries on semantic lexicons that could include newly observed subwords. This two-stage learning makes the coverage of words broaden to a great extent. The experimental results clearly show that our method provides consistent performance improvements over strong baselines that use subwords or lexical resources separately. Yeachan Kim, Kang-Min Kim, SangKeun Lee 0001 |
LREC | 3 |
| 2019 | From Text Classification to Keyphrase Extraction for Short TextabstractExisting keyphrase extraction approaches often suffer from issues such as the sparsity and brevity of short text (e.g., headlines, queries, and tweets). In this paper, we propose a novel keyphrase extraction method for short text by utilizing recurrent neural networks. The main idea behind our approach is to classify short text into a relevant class or category and extract keyphrases from important words in the class or category. Unlike previous supervised approaches that need the information of annotated keyphrases, our approach requires only a text classification dataset (i.e., DBpedia), which is easier to use and requires less human effort. In our approach, we first feed short text into the attention-based neural network for text classification. We then compute attention weights of each word in input short text. Subsequently, we detect keyphrase candidates by chunking phrases and summing the attention weights of compositional words in the chunked phrase. The experimental results clearly show the efficacy of our approach on real-world datasets, such as headlines, queries, and tweets. The proposed method outperforms the Microsoft Cognitive Services and IBM Watson Natural Language Understanding service for keyphrase extraction in terms of F1-score and acceptable percentage on the NYT and Question datasets. Further, we confirm that the proposed method is comparable to supervised methods for keyphrase extraction from short text in the Tweet dataset. Song-Eun Lee, Kang-Min Kim, Woo-Jong Ryu, Jemin Park, SangKeun Lee 0001 |
IEEE BigData | 5 |
| 2019 | meChat: In-device Conversational Photo Sharing ServiceabstractIn this demo, we demonstrate an in-device conversational photo sharing service, termed meChat, which helps users share in-device photos easily in messaging applications by searching conversation-related photos automatically. In particular, meChat understands the semantics of on-going conversation and in-device photos by projecting both of them into a single semantic space. Subsequently, it retrieves in-device photos related to the conversation context. Through this process, meChat makes it easy for the users to share the photos while communicating in messaging applications. In addition, it is worth noting that meChat works in a stand-alone, privacy-protecting manner without sending out any in-device photos and conversations to the external servers. This is different from existing photo services (e.g., Google Photos) which resort on the cloud server. Woo-Jong Ryu, Yoonjoo Ahn, Song-Eun Lee, Kang-Min Kim, Jun-Hyung Park, SangKeun Lee 0001 |
MobiSys | 7 |
| 2019 | From Small-scale to Large-scale Text ClassificationabstractNeural network models have achieved impressive results in the field of text classification. However, existing approaches often suffer from insufficient training data in a large-scale text classification involving a large number of categories (e.g., several thousands of categories). Several neural network models have utilized multi-task learning to overcome the limited amount of training data. However, these approaches are also limited to small-scale text classification. In this paper, we propose a novel neural network-based multi-task learning framework for large-scale text classification. To this end, we first treat the different scales of text classification (i.e., large and small numbers of categories) as multiple, related tasks. Then, we train the proposed neural network, which learns small- and large-scale text classification tasks simultaneously. In particular, we further enhance this multi-task learning architecture by using a gate mechanism, which controls the flow of features between the small- and large-scale text classification tasks. Experimental results clearly show that our proposed model improves the performance of the large-scale text classification task with the help of the small-scale text classification task. The proposed scheme exhibits significant improvements of as much as 14% and 5% in terms of micro-averaging and macro-averaging F1-score, respectively, over state-of-the-art techniques. Kang-Min Kim, Yeachan Kim, Ji-Min Lee, SangKeun Lee 0001 |
WWW | 5 |
| 2018 | Learning to Generate Word Representations using Subword InformationabstractDistributed representations of words play a major role in the field of natural language processing by encoding semantic and syntactic information of words. However, most existing works on learning word representations typically regard words as individual atomic units and thus are blind to subword information in words. This further gives rise to a difficulty in representing out-of-vocabulary (OOV) words. In this paper, we present a character-based word representation approach to deal with this limitation. The proposed model learns to generate word representations from characters. In our model, we employ a convolutional neural network and a highway network over characters to extract salient features effectively. Unlike previous models that learn word representations from a large corpus, we take a set of pre-trained word embeddings and generalize it to word entries, including OOV words. To demonstrate the efficacy of the proposed model, we perform both an intrinsic and an extrinsic task which are word similarity and language modeling, respectively. Experimental results show clearly that the proposed model significantly outperforms strong baseline models that regard words or their subwords as atomic units. For example, we achieve as much as 18.5% improvement on average in perplexity for morphologically rich languages compared to strong baselines in the language modeling task. Yeachan Kim, Kang-Min Kim, Ji-Min Lee, SangKeun Lee 0001 |
COLING | 4 |
| 2018 | Utilizing Probase in Open Directory Project-based Text ClassificationabstractOpen Directory Project (ODP) has been successfully utilized in text classification due to its representation ability of various categories. However, ODP includes a limited number of entities, which play an important role in classification tasks. In this paper, we enrich the semantics of ODP categories with Probase entities. To effectively incorporate Probase entities in ODP categories, we first represent each ODP category and Probase entity in terms of concepts. Next, we measure the semantic relevance between an ODP category and a Probase entity based on the concept vector. Finally, we use Probase entity to enrich the semantics of the ODP categories. Our experimental results show that the proposed methodology exhibits a significant improvement over state-of-the-art techniques in the ODP-based text classification. So-Young Jun, Dinara Aliyeva, Ji-Min Lee, SangKeun Lee 0001 |
FUZZ-IEEE | 4 |
| 2018 | Incorporating Word Embeddings into Open Directory Project Based Large-Scale Classification
Kang-Min Kim, Dinara Aliyeva, Byung-Ju Choi, SangKeun Lee 0001 |
PAKDD (2) | 4 |
| 2018 | Deriving human activity from geo-located data by ontological and statistical reasoning
Zolzaya Dashdorj, Stanislav Sobolevsky, SangKeun Lee 0001, Carlo Ratti |
Knowl. Based Syst. | 3 |
| 2017 | Demo: sigSocial: A Novel Social Media Aggregation Service using a Tiny Text IntelligenceabstractWe present an entirely novel concept of retrieving social media data, called sigSocial. It integrates social media data of various sources, using a semantic classifier. Nowadays, people use multiple social media simultaneously, acquiring information with ease. However, accessing numerous services to reach different channels is bothersome. Also, the volume of information one can process is limited. Our aim is to reduce this burden, providing easiness and efficiency. In other words, we attempt to build a single service that integrates information from various platforms. The application has three main features. First, it enables users to explore multiple social media without accessing them separately. Second, it organizes information retrieved from social medias into well-defined classes. Finally, it works as a stand-alone application, the mechanism of which is internal to the device, not relying on any external servers or networks. This method respects user privacy, which has recently gained much attention. Hyunwoong Bang, Hyunsub Kim, SangKeun Lee 0001 |
MobiSys | 3 |
| 2017 | Demo: Mobile Contextual Advertising Platform based on Tiny Text IntelligenceabstractIn-app advertising has become a significant source of revenue for mobile app. In order to improve the effectiveness of in-app ads, most ad networks focus on targeting a user based on the user's personal information collected from their ad library inside mobile apps and the global knowledge built from big data on ad servers. However, sharing user's sensitive information with the ad servers may raise privacy concerns. As opposed to targeting users, mobile contextual advertising seeks to target the app page a user is viewing. In this demo, we present a novel mobile contextual advertising platform, called MoCA, which is designed to improve the semantic relevance of in-app ads in a stand-alone, privacy-protecting manner on mobile devices. MoCA understands the semantics of app page and ads, and then matches semantically relevant ads to the page inside mobile devices. To the best of our knowledge, this is the first work to implement the mobile contextual advertising platform based on the semantic approach without resort to ad servers. So-Young Jun, So-Jung Park, Kang-Min Kim, SangKeun Lee 0001 |
MobiSys | 5 |
| 2017 | Hashtag-based topic evolution in social media
Md. Hijbul Alam, Woo-Jong Ryu, SangKeun Lee 0001 |
World Wide Web | 3 |
| 2016 | Joint multi-grain topic sentiment: modeling semantic aspects for online reviews
Md. Hijbul Alam, Woo-Jong Ryu, SangKeun Lee 0001 |
Inf. Sci. | 3 |
| 2015 | XQStream++: Fast tuple extraction algorithm for streaming XML data
Byung-Gul Ryu, JongWoo Ha, SangKeun Lee 0001 |
Inf. Sci. | 3 |
| 2014 | Toward robust classification using the Open Directory ProjectabstractThe Open Directory Project (ODP) is a large scale, high quality and publicly available web directory utilized in many studies and real-world applications. In this paper, we explore training data expansion techniques for text classification as one of the possible directions to deal with the sparse characteristic of the ODP dataset. We propose a dozen classification methods, which can be differentiated by (1) from which categories training data is expanded, and (2) how the expanded training data is merged to generate centroid vectors. Evaluation results show that training data expansion significantly improves the classification performance more than representative classifiers. We also find that (1) child and descendant categories are more valuable sources to expand training data than parent and ancestor categories, and (2) distance-based weighting is superior to simple averaging to merge the expanded training data. JongWoo Ha, Won-Jun Jang, Yong-Ku Lee, SangKeun Lee 0001 |
DSAA | 5 |
| 2013 | Impact of node distance on selfish replica allocation in a mobile ad-hoc network
Byung-Gul Ryu, Jae-Ho Choi 0001, SangKeun Lee 0001 |
Ad Hoc Networks | 3 |
| 2013 | Semantic contextual advertising based on the open directory projectabstractContextual advertising seeks to place relevant textual ads within the content of generic webpages. In this article, we explore a novel semantic approach to contextual advertising. This consists of three tasks: (1) building a well-organized hierarchical taxonomy of topics, (2) developing a robust classifier for effectively finding the topics of pages and ads, and (3) ranking ads based on the topical relevance to pages. First, we heuristically build our own taxonomy of topics from the Open Directory Project (ODP). Second, we investigate how to increase classification accuracy by taking the unique characteristics of the ODP into account. Last, we measure the topical relevance of ads by applying a link analysis technique to the similarity graph carefully derived from our taxonomy. Experiments show that our classification method improves the performance of Ma- F 1 by as much as 25.7% over the baseline classifier. In addition, our ranking method enhances the relevance of ads substantially, up to 10% in terms of precision at k , compared to a representative strategy. JongWoo Ha, Jin-Yong Jung, SangKeun Lee 0001 |
ACM Trans. Web | 4 |
| 2012 | Semantic Aspect Discovery for Online ReviewsabstractThe number of opinions and reviews about different products and services is growing online. Users frequently look for important aspects of a product or service in the reviews. Usually, they are interested in semantic (i.e., sentiment-oriented) aspects. However, extracting semantic aspects with supervised methods is very expensive. We propose a domain independent unsupervised model to extract semantic aspects, and conduct qualitative and quantitative experiments to evaluate the extracted aspects. The experiments show that our model effectively extracts semantic aspects with correlated top words. In addition, the conducted evaluation on aspect sentiment classification shows that our model outperforms other models by 5-7% in terms of macro-average F1. Md. Hijbul Alam, SangKeun Lee 0001 |
ICDM | 2 |
| 2012 | Examining the impact of data-access cost on XML twig pattern matching
SangKeun Lee 0001, Byung-Gul Ryu, Kun-Lung Wu |
Inf. Sci. | 1 |
| 2012 | Novel approaches to crawling important pages early
Md. Hijbul Alam, JongWoo Ha, SangKeun Lee 0001 |
Knowl. Inf. Syst. | 3 |
| 2012 | Handling Selfishness in Replica Allocation over a Mobile Ad Hoc NetworkabstractIn a mobile ad hoc network, the mobility and resource constraints of mobile nodes may lead to network partitioning or performance degradation. Several data replication techniques have been proposed to minimize performance degradation. Most of them assume that all mobile nodes collaborate fully in terms of sharing their memory space. In reality, however, some nodes may selfishly decide only to cooperate partially, or not at all, with other nodes. These selfish nodes could then reduce the overall data accessibility in the network. In this paper, we examine the impact of selfish nodes in a mobile ad hoc network from the perspective of replica allocation. We term this selfish replica allocation. In particular, we develop a selfish node detection algorithm that considers partial selfishness and novel replica allocation techniques to properly cope with selfish replica allocation. The conducted simulations demonstrate the proposed approach outperforms traditional cooperative replica allocation techniques in terms of data accessibility, communication cost, and average query delay. Jae-Ho Choi 0001, Kyu-Sun Shim, SangKeun Lee 0001, Kun-Lung Wu |
IEEE Trans. Mob. Comput. | 3 |
| 2010 | A Cluster-Based Group Key Management Scheme for Wireless Sensor NetworksabstractIn this paper, we propose a cluster-based group key management scheme for wireless sensor networks(WSNs) that targets at reduce the communication overhead and storage cost of sensor nodes. In the proposed scheme, a group key is generated by the collaboration of cluster head and nodes within the cluster. Only cluster heads take responsible for reconstruct and delivery the group key. Performance evaluations demonstrate that the proposed scheme maintains a good level of security while significantly reduced the communication overhead compared with the existing schemes, especially in a large scale WSN. Yongluo Shen, SangKeun Lee 0001 |
APWeb | 3 |
| 2010 | EUI: an embedded engine for understanding user intents from mobile devicesabstractWe design and implement a novel embedded software engine, called EUI, to understand user intents from usage data within mobile devices. By developing the EUI engine in mobile devices, we expect to move towards proactive devices for mobile personalized services. To this end, we seek to embed the Open Directory Project (ODP) into mobile devices, and build a robust classifier with the embedded ODP. Thus, the EUI engine classifies the usage data within mobile devices into some ODP categories. Our implementation handles some challenging issues in embedding the ODP and building a robust classifier. The demonstration shows that our implementation understands the semantics of the usage data effectively. JongWoo Ha, Kyu-Sun Shim, SangKeun Lee 0001 |
CIKM | 4 |
| 2010 | Tree-Based Index Overlay in Hybrid Peer-to-Peer Systems
InSung Kang, SungJin Choi, Soon Young Jung, SangKeun Lee 0001 |
J. Comput. Sci. Technol. | 4 |
| 2010 | A Binary String Approach for Updates in Dynamic Ordered XML DataabstractTo facilitate XML query processing, several labeling schemes have been proposed, in which the ancestor-descendant and parent-child relationships in XML queries can be quickly determined without accessing the original XML file. However, all of these existing schemes have to relabel the existing nodes or recalculate certain values when order-sensitive updates cause insertions, thus causing the label update cost to be high. In this paper, we propose a novel labeling scheme, called IBSL (Improved Binary String Labeling), which supports order-sensitive updates without relabeling or recalculation. In addition, we reuse the deleted labels at the same position in the XML tree. The conducted experimental results show that IBSL efficiently processes order-sensitive queries and leaf node/subtree updates. Hye-Kyeong Ko, SangKeun Lee 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2009 | Fractional PageRank Crawler: Prioritizing URLs Efficiently for Crawling Important Pages Early
Md. Hijbul Alam, JongWoo Ha, SangKeun Lee 0001 |
DASFAA | 3 |
| 2009 | Energy Efficient and Progressive Strategy for Processing Skyline Queries on Air
JongWoo Ha, Yoon Kwon, Jae-Ho Choi 0001, SangKeun Lee 0001 |
DEXA | 4 |
| 2007 | Bottom-up nearest neighbor search for R-trees
MoonBae Song, KwangJin Park, Ki-Sik Kong, SangKeun Lee 0001 |
Inf. Process. Lett. | 4 |
| 2007 | On the efficiency of secure XML broadcasting
Hye-Kyeong Ko, Min-Jeong Kim, SangKeun Lee 0001 |
Inf. Sci. | 3 |
| 2006 | An Effective, Efficient XML Data Broadcasting Method in a Mobile Wireless Network
Sang-Hyun Park 0002, Jae-Ho Choi 0001, SangKeun Lee 0001 |
DEXA | 3 |
| 2006 | Fast and Memory-Efficient NN Search in Wireless Data Broadcast
Myong-Soo Lee, SangKeun Lee 0001 |
HPCC | 2 |
| 2006 | An Efficient Scheme to Completely Avoid Re-labeling in XML Updates
Hye-Kyeong Ko, SangKeun Lee 0001 |
WISE | 2 |
| 2006 | Efficient, Energy Conserving Transaction Processing in Wireless Data BroadcastabstractBroadcasting in wireless mobile computing environments is an effective technique to disseminate information to a massive number of clients equipped with powerful, battery operated devices. To conserve the usage of energy, which is a scarce resource, the information to be broadcast must be organized so that the client can selectively tune in at the desired portion of the broadcast. In this paper, the efficient, energy conserving transaction processing in mobile broadcast environments is examined with widely accepted approaches to indexed data organizations suited for a single item retrieval. The basic idea is to share the index information on multiple data items based on the predeclaration technique. The analytical and simulation studies have been performed to evaluate the effectiveness of our methodology, showing that predeclaration-based transaction processing with selective tuning ability can provide a significant performance improvement of battery life, while retaining a low access time. Tolerance to access failures during transaction processing is also described. SangKeun Lee 0001, Chong-Sun Hwang, Masaru Kitsuregawa |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2005 | Efficient Dissemination of Filtered Data in XML-Based SDI
Jae-Ho Choi 0001, Young-Jin Yoon, SangKeun Lee 0001 |
DEXA | 3 |
| 2005 | A Stochastic Viewpoint on the Generation of Spatiotemporal Datasets
MoonBae Song, KwangJin Park, Ki-Sik Kong, SangKeun Lee 0001 |
ICCSA (2) | 4 |
| 2005 | An Efficient Cache Access Protocol in a Mobile Computing Environment
Jae-Ho Choi 0001, SangKeun Lee 0001 |
ISPA | 2 |
| 2004 | Energy Efficient Transaction Processing in Mobile Broadcast Environments
SangKeun Lee 0001 |
ADBIS | 1 |
| 2004 | Efficient Transaction Processing in Mobile Data Broadcast Environments
SangKeun Lee 0001, SungSuk Kim |
DASFAA | 1 |
| 2004 | Energy-Efficient Message Management Algorithms in HMIPv6
Sun Ok Yang, SungSuk Kim, Chong-Sun Hwang, SangKeun Lee 0001 |
ICCSA (1) | 4 |
| 2004 | Performance Evaluation of a Predeclaration-Based Transaction Processing in a Hybrid Data DeliveryabstractPush-based broadcasting in wireless information services is a very effective technique to disseminate information to a massive number of clients when the number of data items is small. When the database is large, however, it may be beneficial to integrate a pull-based (client-to-server) backchannel with the push-based broadcast approach, resulting in a hybrid data delivery. In this paper, we analyze the performance behavior of a predeclaration-based transaction processing, which was originally devised for a push-based data broadcast, in the hybrid data delivery through an extensive simulation. Simulation results show that the use of predeclaration-based transaction processing can provide significant performance improvement not only in a pure push data delivery, but also in a hybrid data delivery. SangKeun Lee 0001, SungSuk Kim |
Mobile Data Management | 1 |
| 2004 | Maintaining mobile transactional consistency in hybrid broadcast environments
SungSuk Kim, Sun Ok Yang, SangKeun Lee 0001 |
Acta Informatica | 3 |
| 2003 | Considering Mobility Patterns in Moving Objects DatabaseabstractWhat is important in location-aware services is how to track moving objects efficiently. To this end, an efficient protocol which updates location information in a location server is highly needed. In fact, the performance of a location update strategy highly depends on the assumed mobility pattern. In most existing works, however, the mobility issue has been disregarded and too simplified as linear function of time. We propose a new mobility model, namely state-based mobility model (SMM) to provide more generalized framework for both describing the mobility and updating location information of moving objects. We also introduce the state-based location update protocol (SLUP) based on this mobility model. MoonBae Song, JeHyok Ryu, SangKeun Lee 0001, Chong-Sun Hwang |
ICPP | 3 |
| 2003 | Using reordering technique for mobile transaction management in broadcast environments
SungSuk Kim, SangKeun Lee 0001, Chong-Sun Hwang |
Data Knowl. Eng. | 2 |
| 2003 | Using Predeclaration for Efficient Read-Only Transaction Processing in Wireless Data BroadcastabstractWireless data broadcast allows a large number of users to retrieve data simultaneously in mobile databases, resulting in an efficient way of using the scarce wireless bandwidth. However, the efficiency of data access methods is limited by an inherent property that data can only be accessed strictly sequentially by users. To properly cope with the inherent property, this paper presents three predeclaration-based transaction processing methods that yield a significant performance improvement in wireless data broadcast. SangKeun Lee 0001, Chong-Sun Hwang, Masaru Kitsuregawa |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2002 | Using Predeclaration for Efficient Read-only Transaction Processing in Wireless Data BroadcastabstractWireless data broadcast allows a large number of users to retrieve data simultaneously in mobile databases, resulting in an efficient way of using the scarce wireless bandwidth. The efficiency of data access methods, however, is limited by an inherent property that data can only be accessed strictly sequentially by users. The paper addresses the issue of ensuring consistency and currency of data items requested in a certain order by wireless read-only transactions. To properly cope with the inherent property of data broadcast, we explore a predeclaration-based query optimization and devise three predeclaration-based transaction processing methods. SangKeun Lee 0001, Masaru Kitsuregawa, Chong-Sun Hwang |
ICDCS | 1 |
| 2001 | O-PreH: Optimistic Transaction Processing Algorithm based on Pre-Reordering in Hybrid Broadcast EnvironmentsabstractIn recent years, there has been a lot of research effort in the periodic push model where the server repetitively disseminates information without explicit request. We call the broadcast model supporting backchannel as hybrid broadcast. In this paper, we devise a new transaction processing algorithm called O-PreH, which is based on the notion of pre-reordering. If one or more conflicts for mobile transactions are found from server's periodic invalidation report, conflict orders are determined not to violate the consistency( pre-reordering) and then the remaining operations have to be executed pessimistically. SungSuk Kim, SangKeun Lee 0001, Soon Young Jung, Chong-Sun Hwang |
CIKM | 2 |
| 2001 | Optimistic Scheduling Algorithm for Mobile Transactions Based on Reordering
SungSuk Kim, Chong-Sun Hwang, Heon-Chang Yu, SangKeun Lee 0001 |
Mobile Data Management | 4 |
| 2001 | Unified Protocols of Concurrency Control and Recovery in Distributed Object-based DatabasesabstractThis paper provides unified protocols of concurrency control and recovery in distributed object-based databases by using two unified conflict notions: preservation and weak preservation. The two conflict relations provide the solutions to (i) the low-level heterogeneity of different recovery mechanisms and/or object models, and (ii) the correct schedules from both concurrency control and recovery points of view. In particular, preservation can be used for accepting serializable and strict (SR-ST) and/or serializable and avoiding cascading aborts (SR-ACA) schedules, whereas weak preservation can be used for accepting serializable and recoverable (SR-RC) schedules. It is also shown that the unified protocols are general enough for object-based databases in addition to the classical read/write databases. SangKeun Lee 0001, Chong-Sun Hwang |
Comput. J. | 1 |
| 2001 | Revisiting Transaction Management in Multidatabase Systems
SangKeun Lee 0001, Chong-Sun Hwang, Heon-Chang Yu |
Distributed Parallel Databases | 1 |
| 1997 | A Uniform Approach to Global Concurrency Control and Recovery in Multidatabase EnvironmentabstractIn this paper, we provide a uniform approach to global con-' currency control and recovery in multidatabase environment.Instead of considering global serializability and global atomicity as two orthogonal concepts, we simply adopt global serializability as the only correctness criterion and require global serializability to be maintained even in a failure-prone multidatabase environment.We first propose rigid conflict serializability (R-CSR) as a sufficient condition for the global transaction manager to ensure global serializability in an autonomous: heterogeneous, and failure-free multidatabase environment.Following this, we show that the combination of cascadeless R-CSR of global transactions and a wntezt-setlJitiue and lute redo recovery leads to the achievement of global serializability in a failure-prone multidatabase environment. SangKeun Lee 0001, Chong-Sun Hwang, Won-Gyu Lee |
CIKM | 1 |
| 1997 | A unified approach to global concurrency control and global deadlocks in a multidatabase environmentabstractOur objective is to provide a theoretical foundation for multidatabase transaction management that deals with global concurrency control and global deadlocks in a uniform manner. We first propose rigid conflict serializability as a sufficient condition for the global transaction management to ensure global serializability in multidatabase environment. Subsequently, it is shown that the enforcement of rigid conflict serializability through a rigid method at the time each global subtransaction begins its execution avoids global deadlocks. The deadlock-free policy in the paper seems to be attractive due to the simple and uniform approach it takes. The basic advantage of the approach is that the global transaction manager can allow any interleavings among normal database operations belonging to global transactions without any mechanism at global level. SangKeun Lee 0001, Chong-Sun Hwang, Won-Gyu Lee |
ICPADS | 1 |
| 1996 | A New Conflict Relation for Concurrency Control and Recovery in object-based DatabasesabstractThis paper proposes preservation as a new conflict relation in an object-based database.By explicitly including reverse-operations which bridge the gap between concurrency control and recovery, preservation can be used independently of execution contexts to which different recovery algorithms and/or object models give rise, and further it forms a basis for formulating semantics-based recovery.This paper also makes a t wo-dimensional( i.e., execution cent exts and operations' specifications) comparison bet ween preservation and other conflict relations.Irr each execution cent ext, our formal comparison reveala that preservationbased concurrency control achieves more concurrency than commutativity-based one. SangKeun Lee 0001, Soon Young Jung, Chong-Sun Hwang |
CIKM | 1 |