EDBT 2026 Demo / reviewers in the wild / expert
Gary Geunbae Lee
dblp:l/GGLee · also Gary Lee 0001, Geunbae Lee
· DBLP profile ↗
162ranked-venue papers
7as first author
50since 2021 · last 2026
0000-0002-3692-6732ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 118 · 4 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 15 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 1 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?abstractEstimating the cognitive complexity of reading comprehension (RC) items is crucial for assessing item difficulty before it is administered to learners.Unlike syntactic and semantic features, such as passage length or semantic similarity between options, cognitive features that arise during answer reasoning are not readily extractable using existing NLP tools and have traditionally relied on human annotation.In this study, we examine whether large language models (LLMs) can estimate the cognitive complexity of RC items by focusing on two dimensions-Evidence Scope and Transformation Level-that indicate the degree of cognitive burden involved in reasoning about the answer.Our experimental results demonstrate that LLMs can approximate the cognitive complexity of items, indicating their potential as tools for prior difficulty analysis.Further analysis reveals a gap between LLMs' reasoning ability and their metacognitive awareness: even when they produce correct answers, they sometimes fail to correctly identify the features underlying their own reasoning process. Seonjeong Hwang, Hyounghun Kim, Gary Geunbae Lee |
ACL (1) | 3 |
| 2026 | A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item GenerationabstractRecent studies in difficulty-controlled reading comprehension item generation have leveraged large language models (LLMs) to produce items by adjusting difficulty-related features.However, existing methods typically rely on a single-agent prompting approach, which often fails to consistently satisfy specified feature constraints, resulting in items that deviate from the target difficulty level.To address this limitation, we introduce MAFIG, a Multiagent Framework for Feature-constrained Item Generation, where multiple LLM agents and feature-specific evaluators collaborate to generate and iteratively revise items based on intended constraints.Furthermore, to verify the efficacy of MAFIG in difficulty control, we propose a method for constructing a sequence of feature constraint sets that yield items with monotonically increasing difficulty.Experimental results demonstrate that MAFIG generates items that adhere to target constraints at a significantly higher rate than baselines, achieving robust difficulty control through the difficulty-calibrated constraint sequence. Seonjeong Hwang, Jun Seo, Hyounghun Kim, Gary Geunbae Lee |
ACL (1) | 4 |
| 2026 | Difficulty-Controllable Cloze Question Distractor GenerationabstractMultiple-choice cloze questions are commonly used to assess linguistic proficiency and comprehension.However, generating high-quality distractors remains challenging, as existing methods often lack adaptability and control over difficulty levels, and the absence of difficulty-annotated datasets further hinders progress.To address these issues, we propose a novel framework for generating distractors with controllable difficulty by leveraging both data augmentation and a multitask learning strategy.First, to create a high-quality, difficultyannotated dataset, we introduce a two-way distractor generation process to produce diverse and plausible distractors.These candidates are filtered and then categorized by difficulty using an ensemble QA system.Second, this newly created dataset is used to train a difficultycontrollable generation model via multitask learning.Experimental results demonstrate that our method generates high-quality distractors across difficulty levels and substantially outperforms GPT-4o in aligning distractor difficulty with human perception. Seokhoon Kang, Yejin Jeon, Seonjeong Hwang, Gary Geunbae Lee |
ACL (1) | 4 |
| 2026 | Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language ModelsabstractWarning: This paper contains examples that may be offensive or upsetting.Large Language Models (LLMs) have greatly advanced Natural Language Processing (NLP), particularly through instruction tuning, which enables broad task generalization without additional fine-tuning.However, their reliance on large-scale datasets-often collected from human or web sources-makes them vulnerable to backdoor attacks, where adversaries poison a small subset of data to implant hidden behaviors.Despite this growing risk, defenses for instruction-tuned models remain underexplored.We propose MB-Defense (Merging & Breaking Defense Framework), a novel training pipeline that immunizes instruction-tuned LLMs against diverse backdoor threats.MB-Defense comprises two stages: (i) Defensive Poisoning, which merges attacker and defensive triggers into a unified backdoor representation, and (ii) Backdoor Neutralization, which breaks this representation through additional training to restore clean behavior.Extensive experiments across multiple LLMs show that MB-Defense substantially lowers attack success rates while preserving instruction-following ability.Our method offers a generalizable and data-efficient defense strategy, improving the robustness of instruction-tuned LLMs against unseen backdoor attacks. Gary Geunbae Lee |
ACL (1) | 2 |
| 2026 | Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech RecognitionabstractAccented speech remains a persistent challenge for automatic speech recognition (ASR), as most models are trained on data dominated by a few high-resource English varieties, leading to substantial performance degradation for other accents.Accent-agnostic approaches improve robustness yet struggle with heavily accented or unseen varieties, while accent-specific methods rely on limited and often noisy labels.We introduce MOE-CTC, a Mixture-of-Experts architecture with intermediate CTC supervision that jointly promotes expert specialization and generalization.During training, accent-aware routing encourages experts to capture accentspecific patterns, which gradually transitions to label-free routing for inference.Each expert is equipped with its own CTC head to align routing with transcription quality, and a routing-augmented loss further stabilizes optimization.Experiments on the MCV-ACCENT benchmark demonstrate consistent gains across both seen and unseen accents in low-and highresource conditions, achieving up to 29.3% relative WER reduction over strong FastConformer baselines. Hyounghun Kim, Gary Geunbae Lee |
ACL (1) | 3 |
| 2026 | Adaptive Planning for Multi-Attribute Controllable Summarization with Monte Carlo Tree SearchabstractControllable summarization moves beyond generic outputs toward human-aligned summaries guided by specified attributes.In practice, the interdependence among attributes makes it challenging for language models to satisfy correlated constraints consistently.Moreover, previous approaches often require perattribute fine-tuning, limiting flexibility across diverse summary attributes.In this paper, we propose adaptive planning for multi-attribute controllable summarization (PACO), a trainingfree framework that reframes the task as planning the order of sequential attribute control with a customized Monte Carlo Tree Search (MCTS).In PACO, nodes represent summaries, and actions correspond to single-attribute adjustments, enabling progressive refinement of only the attributes requiring further control.This strategy adaptively discovers optimal control orders, ultimately producing summaries that effectively meet all constraints.Extensive experiments across diverse domains and models demonstrate that PACO achieves robust multi-attribute controllability, surpassing both LLM-based self-planning models and finetuned baselines.Remarkably, PACO with Llama-3.2-1Brivals the controllability of the much larger Llama-3.3-70Bbaselines.With larger models, PACO achieves superior control performance, outperforming all competitors. Sangwon Ryu, Heejin Do, Yunsu Kim 0001, Gary Geunbae Lee, Jungseul Ok |
ACL (1) | 4 |
| 2026 | Teach-to-reason with scoring: Self-explainable rationale-driven multi-trait essay scoringabstract• We propose RaDME, a self-explainable multi-trait AES model, achieving four key gains. • Explainability : generate explicit trait-wise rationales for model-generated scores. • Accuracy : rationale generation even improves scoring performance. • Efficiency : distilling LLM reasoning capacity yields a lightweight scorer. • Consistency : reveal that a scoring-first design stabilizes scores and explanations. Multi-trait automated essay scoring (AES) systems provide a fine-grained evaluation of an essay’s diverse aspects. While they excel in scoring, prior systems fail to explain why specific trait scores are assigned. This lack of transparency leaves instructors and learners unconvinced of the AES outputs, hindering their practical use. To address this, we propose a self-explainable Rationale-Driven Multi-trait automated Essay scoring (RaDME) 1 1 Codes and all generated results will be publicly available. framework. RaDME leverages the reasoning capabilities of large language models (LLMs) by distilling them into a smaller yet effective scorer. This more manageable student model is optimized to sequentially generate a trait score followed by the corresponding rationale, thereby inherently learning to select a more justifiable score by considering the subsequent rationale during training. Our findings indicate that while LLMs underperform in direct AES tasks, they excel in rationale generation when provided with precise numerical scores. Thus, RaDME integrates the superior reasoning capacities of LLMs into the robust scoring accuracy of an optimized, smaller model. Extensive experiments demonstrate that RaDME achieves both accurate and adequate reasoning while supporting high-quality multi-trait scoring, significantly enhancing the transparency of AES. Heejin Do, Sangwon Ryu, Gary Geunbae Lee |
Expert Syst. Appl. | 3 |
| 2026 | Semantic parsing with candidate expressions for knowledge base question answering
Daehwan Nam, Gary Geunbae Lee |
Expert Syst. Appl. | 2 |
| 2025 | Multi-Facet Blending for Faceted Query-by-Example RetrievalabstractWith the growing demand to fit fine-grained user intents, faceted query-by-example (QBE), which retrieves similar documents conditioned on specific facets, has gained recent attention.However, prior approaches mainly depend on document-level comparisons using basic indicators like citations due to the lack of facet-level relevance datasets; yet, this limits their use to citation-based domains and fails to capture the intricacies of facet constraints.In this paper, we propose a multi-facet blending (FaBle) augmentation method, which exploits modularity by decomposing and recomposing to explicitly synthesize facet-specific training sets.We automatically decompose documents into facet units and generate (ir)relevant pairs by leveraging LLMs' intrinsic distinguishing capabilities; then, dynamically recomposing the units leads to facet-wise relevance-informed document pairs.Our modularization eliminates the need for pre-defined facet knowledge or labels.Further, to prove the FaBle's efficacy in a new domain beyond citation-based scientific paper retrieval, we release a benchmark dataset for educational exam item QBE.FaBle augmentation on 1K documents remarkably assists training in obtaining facet conditional embeddings. Heejin Do, Sangwon Ryu, Jonghwi Kim, Gary Geunbae Lee |
ACL (1) | 4 |
| 2025 | Retrieval-Augmented Fine-Tuning With Preference Optimization For Visual Program GenerationabstractDeokhyung Kang, Jeonghun Cho, Yejin Jeon, Sunbin Jang, Minsub Lee, Jawoon Cho, Gary Lee. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Deokhyung Kang, Jeonghun Cho 0002, Yejin Jeon, Sunbin Jang, Minsub Lee, Jawoon Cho, Gary Geunbae Lee |
ACL (1) | 7 |
| 2025 | MiLQ: Benchmarking IR Models for Bilingual Web Search with Mixed Language QueriesabstractDespite bilingual speakers frequently using mixed-language queries in web searches, Information Retrieval (IR) research on them remains scarce.To address this, we introduce MiLQ, Mixed-Language Query test set, the first public benchmark of mixed-language queries, qualified as realistic and relatively preferred.Experiments show that multilingual IR models perform moderately on MiLQ and inconsistently across native, English, and mixed-language queries, also suggesting code-switched training data's potential for robust IR models handling such queries.Meanwhile, intentional English mixing in queries proves an effective strategy for bilinguals searching English documents, which our analysis attributes to enhanced token matching compared to native queries. 1 * This work was done when the author was at aiXplain 1 The code and data for this work are available at : https://github.com/jonghwi-kim/milq.2 In this study, code-switching, mixed-language, and codemixing are used synonymously.Was sind die Vorteile und Nachteile einer einheitlichen europäischen Währung?Was sind die Advantages und Disadvantages einer single European Currency?What are the advantages and disadvantages of a single European currency? Jonghwi Kim, Deokhyung Kang, Seonjeong Hwang, Yunsu Kim 0001, Jungseul Ok, Gary Geunbae Lee |
EMNLP | 6 |
| 2025 | MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with ResistanceabstractRecent studies have explored the use of large language models (LLMs) in psychotherapy; however, text-based cognitive behavioral therapy (CBT) models often struggle with client resistance, which can weaken therapeutic alliance. To address this, we propose a multimodal approach that incorporates nonverbal cues, which allows the AI therapist to better align its responses with the client’s negative emotional state.Specifically, we introduce a new synthetic dataset, Mirror (Multimodal Interactive Rolling with Resistance), which is a novel synthetic dataset that pairs each client’s statements with corresponding facial images. Using this dataset, we train baseline vision language models (VLMs) so that they can analyze facial cues, infer emotions, and generate empathetic responses to effectively manage client resistance.These models are then evaluated in terms of both their counseling skills as a therapist, and the strength of therapeutic alliance in the presence of client resistance. Our results demonstrate that Mirror significantly enhances the AI therapist’s ability to handle resistance, which outperforms existing text-based CBT approaches.Human expert evaluations further confirm the effectiveness of our approach in managing client resistance and fostering therapeutic alliance. Hoonrae Kim, Yejin Jeon, Gary Geunbae Lee |
EMNLP | 5 |
| 2025 | KoBLEX: Open Legal Question Answering with Multi-hop ReasoningabstractLarge Language Models (LLM) have achieved remarkable performances in general domains and are now extending into the expert domain of law.Several benchmarks have been proposed to evaluate LLMs' legal capabilities.However, these benchmarks fail to evaluate open-ended and provisiongrounded Question Answering (QA).To address this, we introduce a Korean Benchmark for Legal EXplainable QA (KOBLEX), designed to evaluate provision-grounded, multihop legal reasoning.KOBLEX includes 226 scenario-based QA instances and their supporting provisions, created using a hybrid LLM-human expert pipeline.We also propose a method called Parametric provisionguided Selection Retrieval (PARSER), which uses LLM-generated parametric provisions to guide legally grounded and reliable answers.PARSER facilitates multi-hop reasoning on complex legal questions by generating parametric provisions and employing a three-stage sequential retrieval process.Furthermore, to better evaluate the legal fidelity of the generated answers, we propose Legal Fidelity Evaluation (LF-EVAL).LF-EVAL is an automatic metric that jointly considers the question, answer, and supporting provisions and shows a high correlation with human judgments.Experimental results show that PARSER consistently outperforms strong baselines, achieving the best results across multiple LLMs.Notably, compared to standard retrieval with GPT-4o, PARSER achieves 37.91 higher F-1 and 30.81 higher LF-EVAL.Further analyses reveal that PARSER efficiently delivers consistent performance across reasoning depths, with ablations confirming the effectiveness of PARSER. 1 * Equal Contribution. 1 The code and dataset are available at https://github. com/daehuikim/ Jihyung Lee, Daehui Kim, Seonjeong Hwang, Hyounghun Kim, Gary Geunbae Lee |
EMNLP | 5 |
| 2025 | PanicToCalm: A Proactive Counseling Agent for Panic AttacksabstractPanic attacks are acute episodes of fear and distress, in which timely, appropriate intervention can significantly help individuals regain stability.However, suitable datasets for training such models remain scarce due to ethical and logistical issues.To address this, we introduce PACE, which is a dataset that includes high-distress episodes constructed from firstperson narratives, and structured around the principles of Psychological First Aid (PFA).Using this data, we train PACER, a counseling model designed to provide both empathetic and directive support, which is optimized through supervised learning and simulated preference alignment.To assess its effectiveness, we propose PANICEVAL, a multi-dimensional framework covering general counseling quality and crisis-specific strategies.Experimental results show that PACER outperforms strong baselines in both counselor-side metrics and client affect improvement.Human evaluations further confirm its practical value, with PACER consistently preferred over general, CBT-based, and GPT-4-powered models in panic scenarios 1 . Yejin Min, Yejin Jeon, SungJun Yang, Hyounghun Kim, Gary Geunbae Lee |
EMNLP | 7 |
| 2025 | Leveraging What's Overfixed: Post-Correction via LLM Grammatical Error OvercorrectionabstractRobust supervised fine-tuned small Language Models (sLMs) often show high reliability but tend to undercorrect.They achieve high precision at the cost of low recall.Conversely, Large Language Models (LLMs) often show the opposite tendency, making excessive overcorrection, leading to low precision.To effectively harness the strengths of LLMs to address the recall challenges in sLMs, we propose Post-Correction via Overcorrection (PoCO), a novel approach that strategically balances recall and precision.PoCO first intentionally triggers overcorrection via LLM to maximize recall by allowing comprehensive revisions, then applies a targeted post-correction step via fine-tuning smaller models to identify and refine erroneous outputs.We aim to harmonize both aspects by leveraging the generative power of LLMs while preserving the reliability of smaller supervised models.Our extensive experiments demonstrate that PoCO effectively balances GEC performance by increasing recall with competitive precision, ultimately improving the overall quality of grammatical error correction. Taehee Park, Heejin Do, Gary Geunbae Lee |
EMNLP | 3 |
| 2025 | A Domain-Specific Multilingual Speech Translation Corpus via Simultaneous InterpretationabstractThis paper presents a novel multilingual speech translation corpus for complex, domain-specific content in Korean, English, Spanish, and Japanese. The corpus contains 4,000 hours of parallel speech, including 1,000 hours of Korean audio with simultaneous sight interpretations in the other three languages by 294 professionals (242 interpreters and 52 Korean voice actors). It also includes transcriptions, translations, and annotations for all languages. The Dewey Decimal Classification was adapted to balance knowledge representation, and speech tasks were conducted in a controlled studio environment to ensure data consistency. Translation, transcription, and annotation workflows were managed through a custom-built platform. The corpus captures nuanced contexts, cultural sensitivities, and domain-specific terminology, addressing linguistic challenges like structural differences between SOV (Korean, Japanese) and SVO languages (English, Spanish). Preliminary evaluations indicate its potential to enhance end-to-end speech translation models, support cross-lingual transfer learning, and tackle real-time translation issues. Gary Geunbae Lee, Hung Soon Kim, Sunhee Kim, Minhwa Chung |
ICASSP | 2 |
| 2025 | Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum LearningabstractDysarthric speakers experience substantial communication challenges due to impaired motor control of the speech apparatus, which leads to reduced speech intelligibility. This creates significant obstacles in dataset curation since actual recording of long, articulate sentences for the objective of training personalized TTS models becomes infeasible. Thus, the limited availability of audio data, in addition to the articulation errors that are present within the audio, complicates personalized speech synthesis for target dysarthric speaker adaptation. To address this, we frame the issue as a domain transfer task and introduce a knowledge anchoring framework that leverages a teacher-student model, enhanced by curriculum learning through audio augmentation. Experimental results show that the proposed zero-shot multi-speaker TTS model effectively generates synthetic speech with markedly reduced articulation errors and high speaker fidelity, while maintaining prosodic naturalness. Yejin Jeon, Solee Im, Gary Geunbae Lee |
INTERSPEECH | 4 |
| 2025 | Revisiting Early Detection of Sexual Predators via Turn-level OptimizationabstractJinMyeong An, Sangwon Ryu, Heejin Do, Yunsu Kim, Jungseul Ok, Gary Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jinmyeong An, Sangwon Ryu, Heejin Do, Yunsu Kim 0001, Jungseul Ok, Gary Geunbae Lee |
NAACL (Long Papers) | 6 |
| 2025 | K-COMP: Retrieval-Augmented Medical Domain Question Answering With Knowledge-Injected CompressorabstractJeonghun Cho, Gary Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jeonghun Cho 0002, Gary Geunbae Lee |
NAACL (Long Papers) | 2 |
| 2025 | Multimodal Cognitive Reframing Therapy via Multi-hop Psychotherapeutic ReasoningabstractSubin Kim, Hoonrae Kim, Heejin Do, Gary Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hoonrae Kim, Heejin Do, Gary Geunbae Lee |
NAACL (Long Papers) | 4 |
| 2025 | DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech RecognitionabstractWonjun Lee, Solee Im, Heejin Do, Yunsu Kim, Jungseul Ok, Gary Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Solee Im, Heejin Do, Yunsu Kim 0001, Jungseul Ok, Gary Geunbae Lee |
NAACL (Long Papers) | 6 |
| 2025 | PicPersona-TOD : A Dataset for Personalizing Utterance Style in Task-Oriented Dialogue with Image PersonaabstractJihyun Lee, Yejin Jeon, Seungyeon Seo, Gary Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yejin Jeon, Seungyeon Seo, Gary Geunbae Lee |
NAACL (Long Papers) | 4 |
| 2025 | Structured reasoning and answer verification: Enhancing question answering system accuracy and explainability
Jihyung Lee, Gary Geunbae Lee |
Knowl. Based Syst. | 2 |
| 2024 | Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker RepresentationsabstractZero-shot multi-speaker TTS aims to synthesize speech with the voice of a chosen target speaker without any fine-tuning. Prevailing methods, however, encounter limitations at adapting to new speakers of out-of-domain settings, primarily due to inadequate speaker disentanglement and content leakage. To overcome these constraints, we propose an innovative negation feature learning paradigm that models decoupled speaker attributes as deviations from the complete audio representation by utilizing the subtraction operation. By eliminating superfluous content information from the speaker representation, our negation scheme not only mitigates content leakage, thereby enhancing synthesis robustness, but also improves speaker fidelity. In addition, to facilitate the learning of diverse speaker attributes, we leverage multi-stream Transformers, which retain multiple hypotheses and instigate a training paradigm akin to ensemble learning. To unify these hypotheses and realize the final speaker representation, we employ attention pooling. Finally, in light of the imperative to generate target text utterances in the desired voice, we adopt adaptive layer normalizations to effectively fuse the previously generated speaker representation with the target text representations, as opposed to mere concatenation of the text and audio modalities. Extensive experiments and validations substantiate the efficacy of our proposed approach in preserving and harnessing speaker-specific attributes vis-à-vis alternative baseline models. Yejin Jeon, Yunsu Kim 0001, Gary Geunbae Lee |
AAAI | 3 |
| 2024 | Multi-Dimensional Optimization for Text Summarization via Reinforcement LearningabstractThe evaluation of summary quality encompasses diverse dimensions such as consistency, coherence, relevance, and fluency.However, existing summarization methods often target a specific dimension, facing challenges in generating well-balanced summaries across multiple dimensions.In this paper, we propose multiobjective reinforcement learning tailored to generate balanced summaries across all four dimensions.We introduce two multi-dimensional optimization (MDO) strategies for adaptive learning: 1) MDO min , rewarding the current lowest dimension score, and 2) MDO pro , optimizing multiple dimensions similar to multi-task learning, resolves conflicting gradients across dimensions through gradient projection.Unlike prior ROUGE-based rewards relying on reference summaries, we use a QA-based reward model that aligns with human preferences.Further, we discover the capability to regulate the length of summaries by adjusting the discount factor, seeking the generation of concise yet informative summaries that encapsulate crucial points.Our approach achieved substantial performance gains compared to baseline models on representative summarization datasets, particularly in the overlooked dimensions. Sangwon Ryu, Heejin Do, Yunsu Kim 0001, Gary Geunbae Lee, Jungseul Ok |
ACL (1) | 4 |
| 2024 | Aspect-Based Semantic Textual Similarity for Educational Test Items
Heejin Do, Gary Geunbae Lee |
AIED (2) | 2 |
| 2024 | Explainable Multi-hop Question Generation: An End-to-End Approach without Intermediate Question LabelingabstractIn response to the increasing use of interactive artificial intelligence, the demand for the capacity to handle complex questions has increased. Multi-hop question generation aims to generate complex questions that requires multi-step reasoning over several documents. Previous studies have predominantly utilized end-to-end models, wherein questions are decoded based on the representation of context documents. However, these approaches lack the ability to explain the reasoning process behind the generated multi-hop questions. Additionally, the question rewriting approach, which incrementally increases the question complexity, also has limitations due to the requirement of labeling data for intermediate-stage questions. In this paper, we introduce an end-to-end question rewriting model that increases question complexity through sequential rewriting. The proposed model has the advantage of training with only the final multi-hop questions, without intermediate questions. Experimental results demonstrate the effectiveness of our model in generating complex questions, particularly 3- and 4-hop questions, which are appropriately paired with input answers. We also prove that our model logically and incrementally increases the complexity of questions, and the generated multi-hop questions are also beneficial for training question answering models. Seonjeong Hwang, Yunsu Kim 0001, Gary Geunbae Lee |
LREC/COLING | 3 |
| 2024 | Leveraging the Interplay between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause FormationabstractContemporary neural speech synthesis models have indeed demonstrated remarkable proficiency in synthetic speech generation as they have attained a level of quality comparable to that of human-produced speech. Nevertheless, it is important to note that these achievements have predominantly been verified within the context of high-resource languages such as English. Furthermore, the Tacotron and FastSpeech variants show substantial pausing errors when applied to the Korean language, which affects speech perception and naturalness. In order to address the aforementioned issues, we propose a novel framework that incorporates comprehensive modeling of both syntactic and acoustic cues that are associated with pausing patterns. Remarkably, our framework possesses the capability to consistently generate natural speech even for considerably more extended and intricate out-of-domain (OOD) sentences, despite its training on short audio clips. Architectural design choices are validated through comparisons with baseline models and ablation studies using subjective and objective metrics, thus confirming model performance. Yejin Jeon, Yunsu Kim 0001, Gary Geunbae Lee |
LREC/COLING | 3 |
| 2024 | Denoising Table-Text Retrieval for Open-Domain Question AnsweringabstractIn table-text open-domain question answering, a retriever system retrieves relevant evidence from tables and text to answer questions. Previous studies in table-text open-domain question answering have two common challenges: firstly, their retrievers can be affected by false-positive labels in training datasets; secondly, they may struggle to provide appropriate evidence for questions that require reasoning across the table. To address these issues, we propose Denoised Table-Text Retriever (DoTTeR). Our approach involves utilizing a denoised training dataset with fewer false positive labels by discarding instances with lower question-relevance scores measured through a false positive detection model. Subsequently, we integrate table-level ranking information into the retriever to assist in finding evidence for questions that demand reasoning across the table. To encode this ranking information, we fine-tune a rank-aware column encoder to identify minimum and maximum values within a column. Experimental results demonstrate that DoTTeR significantly outperforms strong baselines on both retrieval recall and downstream QA tasks. Our code is available at https://github.com/deokhk/DoTTeR. Deokhyung Kang, Baikjin Jung, Yunsu Kim 0001, Gary Geunbae Lee |
LREC/COLING | 4 |
| 2024 | Autoregressive Multi-trait Essay Scoring via Reinforcement Learning with Scoring-aware Multiple RewardsabstractRecent advances in automated essay scoring (AES) have shifted towards evaluating multiple traits to provide enriched feedback.Like typical AES systems, multi-trait AES employs the quadratic weighted kappa (QWK) to measure agreement with human raters, aligning closely with the rating schema; however, its non-differentiable nature prevents its direct use in neural network training.In this paper, we propose Scoring-aware Multi-reward Reinforcement Learning (SaMRL), which integrates actual evaluation schemes into the training process by designing QWK-based rewards with a mean-squared error penalty for multi-trait AES.Existing reinforcement learning (RL) applications in AES are limited to classification models despite associated performance degradation, as RL requires probability distributions; instead, we adopt an autoregressive score generation framework to leverage token generation probabilities for robust multi-trait score predictions.Empirical analyses demonstrate that SaMRL facilitates model training, notably enhancing scoring of previously inferior prompts. Heejin Do, Sangwon Ryu, Gary Geunbae Lee |
EMNLP | 3 |
| 2024 | Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target LanguagesabstractAutomatic question generation (QG) serves a wide range of purposes, such as augmenting question-answering (QA) corpora, enhancing chatbot systems, and developing educational materials.Despite its importance, most existing datasets predominantly focus on English, resulting in a considerable gap in data availability for other languages.Cross-lingual transfer for QG (XLT-QG) addresses this limitation by allowing models trained on high-resource language datasets to generate questions in lowresource languages.In this paper, we propose a simple and efficient XLT-QG method that operates without the need for monolingual, parallel, or labeled data in the target language, utilizing a small language model.Our model, trained solely on English QA datasets, learns interrogative structures from a limited set of question exemplars, which are then applied to generate questions in the target language.Experimental results show that our method outperforms several XLT-QG baselines and achieves performance comparable to GPT-3.5-turbo across different languages.Additionally, the synthetic data generated by our model proves beneficial for training multilingual QA models.With significantly fewer parameters than large language models and without requiring additional training for target languages, our approach offers an effective solution for QG and QA tasks across various languages 1 . Seonjeong Hwang, Yunsu Kim 0001, Gary Geunbae Lee |
EMNLP | 3 |
| 2024 | Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic ParsingabstractRecent efforts have aimed to utilize multilingual pretrained language models (mPLMs) to extend semantic parsing (SP) across multiple languages without requiring extensive annotations.However, achieving zero-shot cross-lingual transfer for SP remains challenging, leading to a performance gap between source and target languages.In this study, we propose Cross-lingual Back-Parsing (CBP), a novel data augmentation methodology designed to enhance cross-lingual transfer for SP.Leveraging the representation geometry of the mPLMs, CBP synthesizes target language utterances from source meaning representations.Our methodology effectively performs cross-lingual data augmentation in challenging zero-resource settings, by utilizing only labeled data in the source language and monolingual corpora.Extensive experiments on two cross-lingual SP benchmarks (Mschema2QA and Xspider) demonstrate that CBP brings substantial gains in the target language.Further analysis of the synthesized utterances shows that our method successfully generates target language utterances with high slot value alignment rates while preserving semantic integrity. Deokhyung Kang, Seonjeong Hwang, Yunsu Kim 0001, Gary Geunbae Lee |
EMNLP | 4 |
| 2024 | Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
Heejin Do, Gary Geunbae Lee |
INTERSPEECH | 3 |
| 2024 | Key-Element-Informed sLLM Tuning for Document SummarizationabstractRemarkable advances in large language models (LLMs) have enabled high-quality text summarization.However, this capability is currently accessible only through LLMs of substantial size or proprietary LLMs with usage fees.In response, smallerscale LLMs (sLLMs) of easy accessibility and low costs have been extensively studied, yet they often suffer from missing key information and entities, i.e., low relevance, in particular, when input documents are long.We hence propose a key-elementinformed instruction tuning for summarization, so-called KEIT-Sum, which identifies key elements in documents and instructs sLLM to generate summaries capturing these key elements.Experimental results on dialogue and news datasets demonstrate that sLLM with KEITSum indeed provides high-quality summarization with higher relevance and less hallucinations, competitive to proprietary LLM. Sangwon Ryu, Heejin Do, Yunsu Kim 0001, Gary Geunbae Lee, Jungseul Ok |
INTERSPEECH | 4 |
| 2024 | An Investigation into Explainable Audio Hate Speech DetectionabstractResearch on hate speech has predominantly revolved around detection and interpretation from textual inputs, leaving verbal content largely unexplored.While there has been limited exploration into hate speech detection within verbal acoustic speech inputs, the aspect of interpretability has been overlooked.Therefore, we introduce a new task of explainable audio hate speech detection.Specifically, we aim to identify the precise time intervals, referred to as audio frame-level rationales, which serve as evidence for hate speech classification.Towards this end, we propose two different approaches: cascading and End-to-End (E2E).The cascading approach initially converts audio to transcripts, identifies hate speech within these transcripts, and subsequently locates the corresponding audio time frames.Conversely, the E2E approach processes audio utterances directly, which allows it to pinpoint hate speech within specific time frames.Additionally, due to the lack of explainable audio hate speech datasets that include audio frame-level rationales, we curated a synthetic audio dataset to train our models.We further validated these models on actual human speech utterances and found that the E2E approach outperforms the cascading method in terms of the audio frame Intersection over Union (IoU) metric.Furthermore, we observed that including frame-level rationales significantly enhances hate speech detection accuracy for the E2E approach. DisclaimerThe reader may encounter content of an offensive or hateful nature.However, given the nature of the work, this cannot be avoided. Jinmyeong An, Yejin Jeon, Jungseul Ok, Yunsu Kim 0001, Gary Geunbae Lee |
SIGDIAL | 6 |
| 2024 | Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation LearningabstractRecent dialogue systems rely on turn-based spoken interactions, requiring accurate Automatic Speech Recognition (ASR).Errors in ASR can significantly impact downstream dialogue tasks.To address this, using dialogue context from user and agent interactions for transcribing subsequent utterances has been proposed.This method incorporates the transcription of the user's speech and the agent's response as model input, using the accumulated context generated by each turn.However, this context is susceptible to ASR errors because it is generated by the ASR model in an auto-regressive fashion.Such noisy context can further degrade the benefits of context input, resulting in suboptimal ASR performance.In this paper, we introduce Context Noise Representation Learning (CNRL) to enhance robustness against noisy context, ultimately improving dialogue speech recognition accuracy.To maximize the advantage of context awareness, our approach includes decoder pre-training using text-based dialogue data and noise representation learning for a context encoder.Based on the evaluation of speech dialogues, our method shows superior results compared to baselines.Furthermore, the strength of our approach is highlighted in noisy environments where user speech is barely audible due to real-world noise, relying on contextual information to transcribe the input accurately. Gary Geunbae Lee |
SIGDIAL | 3 |
| 2024 | DiagESC: Dialogue Synthesis for Integrating Depression Diagnosis into Emotional Support ConversationabstractDialogue systems for mental health care aim to provide appropriate support to individuals experiencing mental distress.While extensive research has been conducted to deliver adequate emotional support, existing studies cannot identify individuals who require professional medical intervention and cannot offer suitable guidance.We introduce the Diagnostic Emotional Support Conversation task for an advanced mental health management system.We develop the DESC dataset 1 to assess depression symptoms while maintaining user experience by utilizing task-specific utterance generation prompts and a strict filtering algorithm.Evaluations by professional psychological counselors indicate that DESC has a superior ability to diagnose depression than existing data.Additionally, conversational quality evaluation reveals that DESC maintains fluent, consistent, and coherent dialogues. Seungyeon Seo, Gary Geunbae Lee |
SIGDIAL | 2 |
| 2023 | Exploring the Viability of Synthetic Audio Data for Audio-Based Dialogue State TrackingabstractDialogue state tracking plays a crucial role in extracting information in task-oriented dialogue systems. However, preceding research are limited to textual modalities, primarily due to the shortage of authentic human audio datasets. We address this by investigating synthetic audio data for audio-based DST. To this end, we develop cascading and end-to-end models, train them with our synthetic audio dataset, and test them on actual human speech data. To facilitate evaluation tailored to audio modalities, we introduce a novel PhonemeF1 to capture pronunciation similarity. Experimental results showed that models trained solely on synthetic datasets can generalize their performance to human voice data. By eliminating the dependency on human speech data collection, these insights pave the way for significant practical advancements in audio-based DST. Data and code are available at https://github.com/JihyunLee1/E2E-DST.1 Yejin Jeon, Yunsu Kim 0001, Gary Geunbae Lee |
ASRU | 5 |
| 2023 | Optimizing Two-Pass Cross-Lingual Transfer Learning: Phoneme Recognition And Phoneme To Grapheme TranslationabstractThis research optimizes two-pass cross-lingual transfer learning in low-resource languages by enhancing phoneme recognition and phoneme-to-grapheme translation models. Our approach optimizes these two stages to improve speech recognition across languages. We optimize phoneme vocabulary coverage by merging phonemes based on shared articulatory characteristics, thus improving recognition accuracy. Additionally, we introduce a global phoneme noise generator for realistic ASR noise during phoneme-to-grapheme training to reduce error propagation. Experiments on the CommonVoice 12.0 dataset show significant reductions in Word Error Rate (WER) for low-resource languages, highlighting the effectiveness of our approach. This research contributes to the advancements of two-pass ASR systems in low-resource languages, offering the potential for improved cross-lingual transfer learning. Gary Geunbae Lee, Yunsu Kim 0001 |
ASRU | 2 |
| 2023 | AutoCAT: Reinforcement Learning for Automated Exploration of Cache-Timing AttacksabstractThe aggressive performance optimizations in modern microprocessors can result in security vulnerabilities. For example, timing-based attacks in processor caches can steal secret keys or break randomization. So far, finding cache-timing vulnerabilities is mostly performed by human experts, which is inefficient and laborious. There is a need for automatic tools that can explore vulnerabilities given that unreported vulnerabilities leave the systems at risk.In this paper, we propose AutoCAT, an automated exploration framework that finds cache timing-channel attack sequences using reinforcement learning (RL). Specifically, AutoCAT formulates the cache timing-channel attack as a guessing game between an attack program and a victim program holding a secret. This guessing game can thus be solved via modern deep RL techniques. AutoCAT can explore attacks in various cache configurations without knowing design details and under different attack and victim program configurations. AutoCAT can also find attacks to bypass certain detection and defense mechanisms. In particular, AutoCAT discovered StealthyStreamline, a new attack that is able to bypass performance counter-based detection and has up to a 71% higher information leakage rate than the state-of-the-art LRU-based attacks on real processors. AutoCAT is the first of its kind in using RL for crafting microarchitectural timing-channel attack sequences and can accelerate cache timing-channel exploration for secure microprocessor designs. Mulong Luo, Wenjie Xiong 0001, Gary Geunbae Lee, Amy Zhang 0001, Yuandong Tian, Hsien-Hsin S. Lee, G. Edward Suh |
HPCA | 3 |
| 2023 | Hierarchical Pronunciation Assessment with Multi-Aspect AttentionabstractAutomatic pronunciation assessment is a major component of a computer-assisted pronunciation training system. To provide in-depth feedback, scoring pronunciation at various levels of granularity such as phoneme, word, and utterance, with diverse aspects such as accuracy, fluency, and completeness, is essential. However, existing multi-aspect multi-granularity methods simultaneously predict all aspects at all granularity levels; therefore, they have difficulty in capturing the linguistic hierarchy of phoneme, word, and utterance. This limitation further leads to neglecting intimate cross-aspect relations at the same linguistic unit. In this paper, we propose a Hierarchical Pronunciation Assessment with Multi-aspect Attention (HiPAMA) model, which hierarchically represents the granularity levels to directly capture their linguistic structures and introduces multi-aspect attention that reflects associations across aspects at the same level to create more connotative representations. By obtaining relational information from both the granularity- and aspect-side, HiPAMA can take full advantage of multi-task learning. Remarkable improvements in the experimental results on the speachocean762 datasets demonstrate the robustness of HiPAMA, particularly in the difficult-to-assess aspects. Heejin Do, Yunsu Kim 0001, Gary Geunbae Lee |
ICASSP | 3 |
| 2023 | MACTA: A Multi-agent Reinforcement Learning Approach for Cache Timing Attacks and Detection
Jiaxun Cui, Mulong Luo, Gary Geunbae Lee, Peter Stone 0001, Hsien-Hsin S. Lee, G. Edward Suh, Wenjie Xiong 0001, Yuandong Tian |
ICLR | 4 |
| 2023 | Score-balanced Loss for Multi-aspect Pronunciation AssessmentabstractWith rapid technological growth, automatic pronunciation assessment has transitioned toward systems that evaluate pronunciation in various aspects, such as fluency and stress.However, despite the highly imbalanced score labels within each aspect, existing studies have rarely tackled the data imbalance problem.In this paper, we suggest a novel loss function, score-balanced loss, to address the problem caused by uneven data, such as bias toward the majority scores.As a re-weighting approach, we assign higher costs when the predicted score is of the minority class, thus, guiding the model to gain positive feedback for sparse score prediction.Specifically, we design two weighting factors by leveraging the concept of an effective number of samples and using the ranks of scores.We evaluate our method on the speechocean762 dataset, which has noticeably imbalanced scores for several aspects.Improved results particularly on such uneven aspects prove the effectiveness of our method. Heejin Do, Yunsu Kim 0001, Gary Geunbae Lee |
INTERSPEECH | 3 |
| 2023 | Tracking Must Go On : Dialogue State Tracking with Verified Self-TrainingabstractIn task-oriented dialogues, dialogue state tracking (DST) is a critical component as it identifies specific information for the user's purpose.However, as annotating DST data requires a significant amount of human effort, leveraging raw dialogue is crucial.To address this, we propose a new self-training (ST) framework with a verification model.Unlike previous ST methods that rely on extensive hyper-parameter searching to filter out inaccurate data, our verification methodology ensures the accuracy and validity of the dataset without using a fixed threshold.Furthermore, to mitigate overfitting, we augment the dataset by generating diverse user utterances.Even when using only 10% of the labeled data, our approach achieves comparable results to a fully labeled MultiWOZ2.0dataset.The evaluation of scalability also demonstrates enhanced robustness in predicting unseen values. Chaebin Lee, Yunsu Kim 0001, Gary Geunbae Lee |
INTERSPEECH | 4 |
| 2023 | Self-feeding training method for semi-supervised grammatical error correction
Soonchoul Kown, Gary Geunbae Lee |
Comput. Speech Lang. | 2 |
| 2023 | Target-Oriented Knowledge Distillation with Language-Family-Based Grouping for Multilingual NMTabstractMultilingual NMT has developed rapidly, but still has performance degradation caused by language diversity and model capacity constraints. To achieve the competitive accuracy of multilingual translation despite such limitations, knowledge distillation, which improves the student network by matching the teacher network’s output, has been applied and shown enhancement by focusing on the important parts of the teacher distribution. However, existing knowledge distillation methods for multilingual NMT rarely consider the knowledge, which has an important function as the student model’s target, in the process. In this article, we propose two distillation strategies that effectively use the knowledge to improve the accuracy of multilingual NMT. First, we introduce a language-family-based approach, guiding to select appropriate knowledge for each language pair. By distilling the knowledge of multilingual teachers that each processes a group of languages classified by language families, the multilingual model overcomes accuracy degradation caused by linguistic diversity. Second, we propose target-oriented knowledge distillation, which intensively focuses on the ground-truth target of knowledge with a penalty strategy. Our method provides a sensible distillation by penalizing samples without actual targets, while additionally targeting the ground-truth targets. Experiments using TED Talk datasets demonstrate the effectiveness of our method with BLEU scores increment. Discussions of distilled knowledge and further observations of the methods also validate our results. Heejin Do, Gary Geunbae Lee |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Conversational QA Dataset Generation with Answer RevisionabstractConversational question-answer generation is a task that automatically generates a large-scale conversational question answering dataset based on input passages. In this paper, we introduce a novel framework that extracts question-worthy phrases from a passage and then generates corresponding questions considering previous conversations. In particular, our framework revises the extracted answers after generating questions so that answers exactly match paired questions. Experimental results show that our simple answer revision approach leads to significant improvement in the quality of synthetic data. Moreover, we prove that our framework can be effectively utilized for domain adaptation of conversational question answering. Seonjeong Hwang, Gary Geunbae Lee |
COLING | 2 |
| 2022 | Schema Encoding for Transferable Dialogue State TrackingabstractDialogue state tracking (DST) is an essential sub-task for task-oriented dialogue systems. Recent work has focused on deep neural models for DST. However, the neural models require a large dataset for training. Furthermore, applying them to another domain needs a new dataset because the neural models are generally trained to imitate the given dataset. In this paper, we propose Schema Encoding for Transferable Dialogue State Tracking (SET-DST), which is a neural DST method for effective transfer to new domains. Transferable DST could assist developments of dialogue systems even with few dataset on target domains. We use a schema encoder not just to imitate the dataset but to comprehend the schema of the dataset. We aim to transfer the model to new domains by encoding new schemas and using them for DST on multi-domain settings. As a result, SET-DST improved the joint accuracy by 1.46 points on MultiWOZ 2.1. Hyunmin Jeon, Gary Geunbae Lee |
COLING | 2 |
| 2022 | SF-DST: Few-Shot Self-Feeding Reading Comprehension Dialogue State Tracking with Auxiliary TaskabstractFew-shot dialogue state tracking (DST) model tracks user requests in dialogue with reliable accuracy even with a small amount of data.In this paper, we introduce an ontology-free few-shot DST with self-feeding belief state input.The selffeeding belief state input increases the accuracy in multi-turn dialogue by summarizing previous dialogue.Also, we newly developed a slot-gate auxiliary task.This new auxiliary task helps classify whether a slot is mentioned in the dialogue.Our model achieved the best score in a few-shot setting for four domains on multiWOZ 2.0. Gary Geunbae Lee |
INTERSPEECH | 2 |
| 2022 | DORA: Towards policy optimization for task-oriented dialogue system with efficient context
Hyunmin Jeon, Gary Geunbae Lee |
Comput. Speech Lang. | 2 |
| 2019 | Adversarial approach to domain adaptation for reinforcement learning on dialog systems
Sangjun Koo, Hwanjo Yu, Gary Geunbae Lee |
Pattern Recognit. Lett. | 3 |
| 2018 | Out-of-domain Detection based on Generative Adversarial NetworkabstractThe main goal of this paper is to develop out-of-domain (OOD) detection for dialog systems.We propose to use only indomain (IND) sentences to build a generative adversarial network (GAN) of which the discriminator generates low scores for OOD sentences.To improve basic GANs, we apply feature matching loss in the discriminator, use domain-category analysis as an additional task in the discriminator, and remove the biases in the generator.Thereby, we reduce the huge effort of collecting OOD sentences for training OOD detection.For evaluation, we experimented OOD detection on a multi-domain dialog system.The experimental results showed the proposed method was most accurate compared to the existing methods. Seonghan Ryu, Sangjun Koo, Hwanjo Yu, Gary Geunbae Lee |
EMNLP | 4 |
| 2018 | Living with AI in Connected Devices for valuable ExperienceabstractThis talk describes how AI technology can be used in multiple connected devices around us such as smartphone, TV, refrigerator and several consumer electronic devices, thus giving new and exciting customer experiences and values. This talk starts with Samsung's AI vision as a device company, and introduces 5 strategic principles along with industrial usages of AI technology including speech and natural language, visual understanding, data intelligence and autonomous driving all with deep learning techniques heavily involved. These applications naturally form a platform for both on device/edge, cloud and machine learning services for various current and future Samsung devices. Gary Geunbae Lee |
ACM Multimedia | 1 |
| 2017 | Automatic sentence stress feedback for non-native English learnersabstractThis paper proposes a sentence stress feedback system in which sentence stress prediction, detection, and feedback provision models are combined. This system provides non-native learners with feedback on sentence stress errors so that they can improve their English rhythm and fluency in a self-study setting. The sentence stress feedback system was devised to predict and detect the sentence stress of any practice sentence. The accuracy of the prediction and detection models was 96.6% and 84.1%, respectively. The stress feedback provision model offers positive or negative stress feedback for each spoken word by comparing the probability of the predicted stress pattern with that of the detected stress pattern. In an experiment that evaluated the educational effect of the proposed system incorporated in our CALL system, significant improvements in accentedness and rhythm were seen with the students who trained with our system but not with those in the control group. Gary Geunbae Lee, Jieun Song, Byeongchang Kim 0001, Sechun Kang, Jinsik Lee, Hyosung Hwang |
Comput. Speech Lang. | 1 |
| 2017 | Two-stage multi-intent detection for spoken language understanding
Byeongchang Kim 0001, Seonghan Ryu, Gary Geunbae Lee |
Multim. Tools Appl. | 3 |
| 2017 | Neural sentence embedding using only in-domain sentences for out-of-domain sentence detection in dialog systems
Seonghan Ryu, Seokhwan Kim, Junhwi Choi, Hwanjo Yu, Gary Geunbae Lee |
Pattern Recognit. Lett. | 5 |
| 2015 | Open-domain personalized dialog system using user-interested topics in system responsesabstractWe built a personalized example-based dialog system that constructs its responses by considering entities that the user has uttered, and topics in which the user has expressed interest. The system analyzes user input utterances, then uses DBpedia and Freebase to extract relevant entities and topics. The extracted entities and topics are stored in personal knowledge memory and are used when the system selects responses from the example database and generates responses. We conducted a human experiment in which evaluators rated dialog systems based on subjective metrics. The proposed dialog system that uses topics that are of interest to the user achieved higher evaluation scores for both personalization and satisfaction than the baseline systems. These results demonstrate that the use of topics in the system response provides a sense that the system pays attention to the user's utterances; as a consequence the user has a satisfactory dialog experience. Jeesoo Bang, Sangdo Han, Kyusong Lee, Gary Geunbae Lee |
ASRU | 4 |
| 2015 | Implementation of generic positive-negative tracker in extensible dialog systemabstractDialog state tracking is one of the most challenging tasks in the implementation of statistical Dialog Management (DM) systems. In development of a Korean dialog system, we implemented a generic tracking approach that can be used to extend a given initial set of system-output types. Our approach uses two methods: confidence estimation for error modeling, and dialog abstraction for dialog state tracking. We adopted a phoneme-sequence matching algorithm to estimate confidence for erroneous Korean user input. We also adopted a positive-negative model to abstract and generalize the effect of given user input and corresponding system output on dialog-state updating. Experiment result implies that our model can be used for dialog tracking without significant loss of performance. We implemented dialog system to verify that our approach is feasible in a practical Korean dialog system that can be adopted for other languages. Sangjun Koo, Seonghan Ryu, Gary Geunbae Lee |
ASRU | 3 |
| 2015 | Question Answering System using Multiple Information Source and Open Type Answer MergeabstractSeonyeong Park, Soonchoul Kwon, Byungsoo Kim, Sangdo Han, Hyosup Shim, Gary Geunbae Lee. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015. Seonyeong Park, Soonchoul Kown, Byungsoo Kim 0002, Sangdo Han, Hyosup Shim, Gary Geunbae Lee |
HLT-NAACL | 6 |
| 2015 | Exploiting knowledge base to generate responses for natural language dialog listening agentsabstractWe developed a natural language dialog listening agent that uses a knowledge base (KB) to generate rich and relevant responses. Our system extracts an important named entity from a user utterance, then scans the KB to extract contents related to this entity. The system can generate diverse and relevant responses by assembling the related KB contents into appropriate sentences. Fifteen students tested our system; they gave it higher approval scores than they gave other systems. These results demonstrate that our system generated various responses and encouraged users to continue talking. Sangdo Han, Jeesoo Bang, Seonghan Ryu, Gary Geunbae Lee |
SIGDIAL Conference | 4 |
| 2015 | Conversational Knowledge Teaching Agent that uses a Knowledge BaseabstractWhen implementing a conversational educational teaching agent, user-intent understanding and dialog management in a dialog system are not sufficient to give users educational information.In this paper, we propose a conversational educational teaching agent that gives users some educational information or triggers interests on educational contents.The proposed system not only converses with a user but also answer questions that the user asked or asks some educational questions by integrating a dialog system with a knowledge base.We used the Wikipedia corpus to learn the weights between two entities and embedding of properties to calculate similarities for the selection of system questions and answers. Kyusong Lee, Hongsuck Seo, Junhwi Choi, Sangjun Koo, Gary Geunbae Lee |
SIGDIAL Conference | 5 |
| 2014 | Vowel-reduction feedback system for non-native learners of EnglishabstractIn spoken English, vowels in non-stressed syllables are often reduced to a brief neutral vowel (e.g, e or ι). Non-native speakers of English may not use this `vowel reduction' correctly, so their utterances may sound unnatural. We propose an automatic system to provide feedback about vowel-reduction to non-native speakers of English. The system has three parts: it predicts vowel reduction, detects vowel reduction in speech, compares the prediction to the detected sound to generate a score then uses this score to provide corrective feedback to the speaker. The system had good accuracy and provided positive learning results for the user. The proposed system can be used as a part of a computer-assisted language learning system. Jeesoo Bang, Kyusong Lee, Seonghan Ryu, Gary Geunbae Lee |
ICASSP | 4 |
| 2014 | Grammatical error correction based on learner comprehension model in oral conversationabstractWe aim to provide grammar error feedback to learners. It is known that grammar error detection and feedback are challenging problems in written language, however, they become much more difficult tasks in oral conversation because it is difficult for a system to judge whether an error is due to grammar or automatic speech recognition (ASR). False alarms occur when a learner correctly utters a remark, but the system gives feedback implying an error. Minimizing the false alarm rate is especially critical in education applications because it is imperative that the tutor give correct instruction to learners. Thus, to reduce the false alarm rate in grammar error detection and feedback, we apply a partially observable Markov decision process (POMDP) when the system provides feedback about a learner's mistake. The POMDP models uncertainty between grammar errors and ASR errors. An additional advantage of our method is that “belief states” in POMDP can be used for learner models which indicate each individual learner's grammar comprehension level. Kyusong Lee, Seonghan Ryu, Hongsuck Seo, Seokhwan Kim, Gary Geunbae Lee |
SLT | 5 |
| 2014 | Pronunciation Variants Prediction Method to Detect Mispronunciations by Korean Learners of EnglishabstractThis article presents an approach to nonnative pronunciation variants modeling and prediction. The pronunciation variants prediction method was developed by generalized transformation-based error-driven learning (GTBL). The modified goodness of pronunciation (GOP) score was applied to effective mispronunciation detection using logistic regression machine learning under the pronunciation variants prediction. English-read speech data uttered by Korean-speaking learners of English were collected, then pronunciation variation knowledge was extracted from the differences between the canonical phonemes and the actual phonemes of the speech data. With this knowledge, an error-driven learning approach was designed that automatically learns phoneme variation rules from phoneme-level transcriptions. The learned rules generate an extended recognition network to detect mispronunciations. Three different mispronunciation detection methods were tested including our logistic regression machine learning method with modified GOP scores and mispronunciation preference features; all three methods yielded significant improvement in predictions of pronunciation variants, and our logistic regression method showed the best performance. Jeesoo Bang, Gary Geunbae Lee, Minhwa Chung |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2014 | Cross-Lingual Annotation Projection for Weakly-Supervised Relation ExtractionabstractAlthough researchers have conducted extensive studies on relation extraction in the last decade, statistical systems based on supervised learning are still limited, because they require large amounts of training data to achieve high performance level. In this article, we propose cross-lingual annotation projection methods that leverage parallel corpora to build a relation extraction system for a resource-poor language without significant annotation efforts. To make our method more reliable, we introduce two types of projection approaches with noise reduction strategies. We demonstrate the merit of our method using a Korean relation extraction system trained on projected examples from an English-Korean parallel corpus. Experiments show the feasibility of our approaches through comparison to other systems based on monolingual resources. Seokhwan Kim, Minwoo Jeong, Gary Geunbae Lee |
ACM Trans. Asian Lang. Inf. Process. | 4 |
| 2013 | Counseling Dialog System with 5W1H Extraction
Sangdo Han, Kyusong Lee, Gary Geunbae Lee |
SIGDIAL Conference | 4 |
| 2013 | Two scalable algorithms for associative text classification
Yongwook Yoon, Gary Geunbae Lee |
Inf. Process. Manag. | 2 |
| 2013 | Unsupervised Spoken Language Understanding for a Multi-Domain Dialog SystemabstractThis paper proposes an unsupervised spoken language understanding (SLU) framework for a multi-domain dialog system. Our unsupervised SLU framework applies a non-parametric Bayesian approach to dialog acts, intents and slot entities, which are the components of a semantic frame. The proposed approach reduces the human effort necessary to obtain a semantically annotated corpus for dialog system development. In this study, we analyze clustering results using various evaluation metrics for four dialog corpora. We also introduce a multi-domain dialog system that uses the unsupervised SLU framework. We argue that our unsupervised approach can help overcome the annotation acquisition bottleneck in developing dialog systems. To verify this claim, we report a dialog system evaluation, in which our method achieves competitive results in comparison with a system that uses a manually annotated corpus. In addition, we conducted several experiments to explore the effect of our approach on reducing development costs. The results show that our approach be helpful for the rapid development of a prototype system and reducing the overall development costs. Minwoo Jeong, Kyungduk Kim, Seonghan Ryu, Gary Geunbae Lee |
IEEE Trans. Speech Audio Process. | 5 |
| 2012 | Seamless error correction interface for voice word processorabstractIn this paper, we propose an error correction interface for a voice word processor. This correction interface includes user intention understanding and automatic error region detection. For accurate correction, we include a confirmation process that includes an error region control command and a re-uttering command. We evaluate the performance of the user intention understanding first, and we evaluate the effectiveness of our interface compare to a general two-step error correction interface. Junhwi Choi, Kyungduk Kim, Seokhwan Kim, Injae Lee, Gary Geunbae Lee |
ICASSP | 7 |
| 2012 | Unsupervised modeling of user actions in a dialog corpusabstractIn data-driven spoken dialog system development, developers should prepare a dialog corpus with semantic annotation. However, the labeling process is a laborious and time consuming task. To reduce human efforts, we propose an unsupervised approach based on non-parametric Bayesian Hidden Markov Model to the problem of modeling user actions. With the non-parametric model, system designers do not need to determine the number and type of user actions. In the experiments, we evaluated the clustering results by comparing them to the human annotation. We also tested a dialog system that used models trained from the automatically annotated corpus with a user simulation. Minwoo Jeong, Kyungduk Kim, Gary Geunbae Lee |
ICASSP | 4 |
| 2012 | Grammatical Error Annotation for Korean Learners of Spoken English
Hongsuck Seo, Kyusong Lee, Gary Geunbae Lee, Soo-Ok Kweon, Hae-Ri Kim |
LREC | 3 |
| 2012 | An automatic pitch accent feedback system for english learners with adaptation of an english corpus spoken by KoreansabstractTo improve the English proficiency of Korean learners, we design a system for pitch accents, which consists of prediction, detection and feedback parts. The prediction and detection parts adopt Conditional Random Field models to achieve a prediction accuracy of 87.25%, which is based on the Boston University radio news corpus, and a detection accuracy of 81.21%, which is based on the Korean Learner's English Accentuation corpus. In the learner experiment with our system, learners' pitch accent proficiency, as assessed by English experts, was improved from 2.67 to 3.25 on a scale of 1-to-5, and the accuracy of not-wrong feedback was measured at 82.77%. The learners assessed the learning effectiveness of our system at 4.3 on a scale of 1-to-5. Sechun Kang, Gary Geunbae Lee, Byeongchang Kim 0001 |
SLT | 2 |
| 2012 | Generating grammar questions using corpus data in L2 learningabstractThis paper examines how grammar questions are automatically generated for L2 learning by applying a sequential labeling technique to learner corpora. We developed a model that helps detect possible error positions and select the most appropriate form among choices. Discriminant models such as conditional random field and maximum entropy are used to generate the error identification question. Questions generated by the proposed method corresponded highly to questions that experts made. Our data-driven approach lends itself to any language without costing expensive expertise. Kyusong Lee, Soo-Ok Kweon, Hongsuck Seo, Gary Geunbae Lee |
SLT | 4 |
| 2012 | Stacking Model-Based Korean Prosodic Phrasing Using Speaker Variability Reduction and Linguistic Feature EngineeringabstractThis article presents a prosodic phrasing model for a general purpose Korean speech synthesis system. To reflect the factors affecting prosodic phrasing in the model, linguistically motivated machine-learning features were investigated. These features were effectively incorporated using a stacking model. The phrasing performance was also improved through feature engineering. The corpus used in the experiment is a 4,392-sentence corpus (55,015 words with an average of 13 words per sentence). Because the corpus contains speaker-dependent variability and such variability is not appropriately reflected in a general purpose speech synthesis system, a method to reduce such variability is proposed. In addition, the entire set of data used in the experiment is provided to the public for future use in comparative research. Jinsik Lee, Byeongchang Kim 0001, Gary Geunbae Lee |
ACM Trans. Asian Lang. Inf. Process. | 5 |
| 2012 | Subcellular Localization Prediction through Boosting Association RulesabstractComputational methods for predicting protein subcellular localization have used various types of features, including N-terminal sorting signals, amino acid compositions, and text annotations from protein databases. Our approach does not use biological knowledge such as the sorting signals or homologues, but use just protein sequence information. The method divides a protein sequence into short $k$-mer sequence fragments which can be mapped to word features in document classification. A large number of class association rules are mined from the protein sequence examples that range from the N-terminus to the C-terminus. Then, a boosting algorithm is applied to those rules to build up a final classifier. Experimental results using benchmark datasets show our method is excellent in terms of both the classification performance and the test coverage. The result also implies that the $k$-mer sequence features which determine subcellular locations do not necessarily exist in specific positions of a protein sequence. Online prediction service implementing our method is available at http://isoft.postech.ac.kr/research/BCAR/subcell. Yongwook Yoon, Gary Geunbae Lee |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2011 | Grammatical Error Detection for Corrective Feedback Provision in Oral ConversationsabstractThe demand for computer-assisted language learning systems that can provide corrective feedback on language learners’ speaking has increased. However, it is not a trivial task to detect grammatical errors in oral conversations because of the unavoidable errors of automatic speech recognition systems. To provide corrective feedback, a novel method to detect grammatical errors in speaking performance is proposed. The proposed method consists of two sub-models: the grammaticality-checking model and the error-type classification model. We automatically generate grammatical errors that learners are likely to commit and construct error patterns based on the articulated errors. When a particular speech pattern is recognized, the grammaticality-checking model performs a binary classification based on the similarity between the error patterns and the recognition result using the confidence score. The error-type classification model chooses the error type based on the most similar error pattern and the error frequency extracted from a learner corpus. The grammaticality checking method largely outperformed the two comparative models by 56.36% and 42.61% in F-score while keeping the false positive rate very low. The error-type classification model exhibited very high performance with a 99.6% accuracy rate. Because high precision and a low false positive rate are important criteria for the language-tutoring setting, the proposed method will be helpful for intelligent computer-assisted language learning systems. Hyungjong Noh, Kyusong Lee, Gary Geunbae Lee |
AAAI | 4 |
| 2011 | A Cross-lingual Annotation Projection-based Self-supervision Approach for Open Information Extraction
Seokhwan Kim, Minwoo Jeong, Gary Geunbae Lee |
IJCNLP | 4 |
| 2011 | Web-Enhanced Content Retrieval for Information Access Dialogue SystemabstractWe consider the problem of content retrieval with complex queries for an information access dialogue system. Traditional information access dialogue systems rely on exact query matching and heuristic rules to find relevant content in a relational database. To deal with complex queries, a dialogue system is used to attain deep semantic processing such as full semantic parsing and ontology-based reasoning. However, these systems require a large amount of semantic annotation and domain expert knowledge that are often very expensive to obtain and thus have been limited in practice. In this paper, we present a simple alternative method where web-searched documents can contribute to enhanced vector space model-based content retrieval. Our model captures underlying co-occurrence patterns between the query and the contents. An efficient ranking algorithm is applied to retrieve the relevant contents. One merit of the proposed approach is that it does not require heavy semantic processing, and therefore, it results in efficient content retrieval. We demonstrate that our method is beneficial in an electronic program-guided dialogue system. Index Terms: web-enhanced content retrieval, information access dialogue system Cheongjae Lee, Minwoo Jeong, Kyungduk Kim, Seokhwan Kim, Junhwi Choi, Gary Geunbae Lee |
INTERSPEECH | 7 |
| 2011 | POMY: A Conversational Virtual Environment for Language Learning in POSTECH
Hyungjong Noh, Kyusong Lee, Gary Geunbae Lee |
SIGDIAL Conference | 4 |
| 2011 | Hybrid user intention modeling to diversify dialog simulations
Sangkeun Jung, Cheongjae Lee, Kyungduk Kim, Gary Geunbae Lee |
Comput. Speech Lang. | 5 |
| 2011 | A local tree alignment approach to relation extraction of multiple arguments
Seokhwan Kim, Minwoo Jeong, Gary Geunbae Lee |
Inf. Process. Manag. | 3 |
| 2011 | Grammatical error simulation for computer-assisted language learning
Hyungjong Noh, Kyusong Lee, Gary Geunbae Lee |
Knowl. Based Syst. | 5 |
| 2011 | Iteratively constrained selection of word alignment links using knowledge and statistics
Hyeongjong Noh, Kyusong Lee, Gary Geunbae Lee |
Knowl. Based Syst. | 5 |
| 2010 | A Cross-lingual Annotation Projection Approach for Relation Detection
Seokhwan Kim, Minwoo Jeong, Gary Geunbae Lee |
COLING | 4 |
| 2010 | Intention-based Corrective Feedback Generation using Context-aware Model
Cheongjae Lee, Hyungjong Noh, Gary Geunbae Lee |
CSEDU (1) | 5 |
| 2010 | Script-description Pair Extraction from Text Documents of English as Second Language Podcast
Hyungjong Noh, Minwoo Jeong, Gary Geunbae Lee |
CSEDU (1) | 5 |
| 2010 | Modeling confirmations for example-based dialog managementabstractThis paper proposes a method to model confirmations for example-based dialog management. To enable the system to provide a confirmation to the user in an appropriate time, we employed a multiple dialog state representation approach for keeping track of user input uncertainty and implemented a confirmation agent which decides when the information gathered from the user contains an error. We developed a car navigation dialog system to evaluate our proposed method. Evaluations with simulated dialogs show our approach is useful for handling misunderstanding errors in example-based dialog management. Kyungduk Kim, Cheongjae Lee, Junhwi Choi, Sangkeun Jung, Gary Geunbae Lee |
SLT | 6 |
| 2010 | Affective effects of speech-enabled robots for language learningabstractThis study introduces the speech and language technologies used in the educational assistant robots that we developed for language learning and exploring the affective effects of robot-assisted language learning (RALL). To achieve this purpose, a course was designed in which students have meaningful interaction with intelligent robots in an immersive environment. A total of 24 elementary students, ranging in age over 9-13, were enrolled in English lessons. Descriptive statistics and pre-test/post-test design were used to investigate the affective effects of RALL approach. The result showed that RALL is promoting and improving students' satisfaction, interest, confidence, and motivation at the significance level of 0.01. Changgu Kim, Hyungjong Noh, Kyusong Lee, Gary Geunbae Lee |
SLT | 6 |
| 2010 | Let's Buy Books: Finding eBooks using voice searchabstractWe describe Let's Buy Books, a dialog system that helps users search for eBook titles. In this paper we compare different vector space approaches to voice search and find that a hybrid approach using a weighted sub-space model smoothed with a general model provides the best performance over different conditions and evaluated using both synthetic queries and queries collected from users through questionnaires. Cheongjae Lee, Alexander I. Rudnicky, Gary Geunbae Lee |
SLT | 3 |
| 2010 | Hybrid approach to robust dialog management using agenda and dialog examples
Cheongjae Lee, Sangkeun Jung, Kyungduk Kim, Gary Geunbae Lee |
Comput. Speech Lang. | 4 |
| 2009 | Correlation-based query relaxation for example-based dialog modelingabstractQuery relaxation refers to the process of reducing the number of constraints on a query if it returns no result when searching a database. This is an important process to enable extraction of an appropriate number of query results because queries that are too strictly constrained may return no result, whereas queries that are too loosely constrained may return too many results. This paper proposes an automated method of correlation-based query relaxation (CBQR) to select an appropriate constraint subset. The example-based dialog modeling framework was used to validate our algorithm. Preliminary results show that the proposed method facilitates the automation of query relaxation. We believe that the CBQR algorithm effectively relaxes constraints on failed queries to return more dialog examples. Cheongjae Lee, Sangkeun Jung, Kyungduk Kim, Gary Geunbae Lee |
ASRU | 6 |
| 2009 | Semi-supervised Speech Act Recognition in Emails and Forums
Minwoo Jeong, Chin-Yew Lin, Gary Geunbae Lee |
EMNLP | 3 |
| 2009 | Hybrid approach to grapheme to phoneme conversion for KoreanabstractIn the grapheme to phoneme conversion problem for Korean, two main approaches have been discussed: knowledge-based and data-driven methods. However, both camps have limita-tions: the knowledge-based hand-written rules cannot handle some of the pronunciation changes due to the lack of capabili-ty of linguistic analyzers and many exceptions; data-driven methods always suffer from data sparseness. To overcome the shortages of both camps, this paper presents a novel combin-ing method which effectively integrates two components: (1) a rule-based converting system based on linguistically motivated hand-written rules and (2) a statistical converting system using a Maximum Entropy model. The experimental results clearly show the effectiveness of our proposed method. Index Terms: grapheme to phoneme conversion, letter to sound rules, hybrid approach 1. Jinsik Lee, Byeongchang Kim 0001, Gary Geunbae Lee |
INTERSPEECH | 3 |
| 2009 | Key node selection for containing infectious disease spread using particle swarm optimizationabstractIn recent years, some emerging and reemerging infectious diseases have grown into global health threats due to high human mobility. It is important to have intervention plans for containing the spread of such infectious diseases. Among various intervention strategies, screening infected people is an efficient way for evaluating the infection scale and controlling the spread of infectious diseases. Considering the cost in manpower and limited screening machines available, we face to challenges for selecting the optimal nodes (sites) in order to obtain better screening and control effects. In this paper, particle swarm optimization technique is used to determine key nodes for controlling infectious disease spread, through evaluating the number of people captured at each key node. The research example is shown on evaluating the screening control over train stations in Singapore. The optimization algorithm and control concept can be easily extended to large-scale infectious disease control in other kinds of key nodes and in other geographical regions. The selection for optimal control set of the multi objective optimization problem is done using particle swarm optimization. Numerical simulation shows the effectiveness of the proposed algorithm. Xiuju Fu, Sonja Lim, Lipo Wang 0001, Gary Geunbae Lee, Stefan Ma, Limsoon Wong, Gaoxi Xiao |
SIS | 4 |
| 2009 | Data-driven user simulation for automated evaluation of spoken dialog systems
Sangkeun Jung, Cheongjae Lee, Kyungduk Kim, Minwoo Jeong, Gary Geunbae Lee |
Comput. Speech Lang. | 5 |
| 2009 | A data-driven grapheme-to-phoneme conversion method using dynamic contextual converting rules for Korean TTS systems
Jinsik Lee, Gary Geunbae Lee |
Comput. Speech Lang. | 2 |
| 2009 | Multi-domain spoken language understanding with transfer learning
Minwoo Jeong, Gary Geunbae Lee |
Speech Commun. | 2 |
| 2009 | Example-based dialog modeling for practical multi-domain dialog system
Cheongjae Lee, Sangkeun Jung, Seokhwan Kim, Gary Geunbae Lee |
Speech Commun. | 4 |
| 2008 | Robust Dialog Management with N-Best Hypotheses Using Dialog Examples and Agenda
Cheongjae Lee, Sangkeun Jung, Gary Geunbae Lee |
ACL | 3 |
| 2008 | Transformation-based Sentence Splitting method for Statistical Machine Translation
Gary Geunbae Lee |
IJCNLP | 3 |
| 2008 | An alignment-based pattern representation model for information extractionabstractNo abstract available. Seokhwan Kim, Minwoo Jeong, Gary Geunbae Lee |
SIGIR | 3 |
| 2008 | Practical use of non-local features for statistical spoken language understanding
Minwoo Jeong, Gary Geunbae Lee |
Comput. Speech Lang. | 2 |
| 2008 | DialogStudio: A workbench for data-driven spoken dialog system development and management
Sangkeun Jung, Cheongjae Lee, Seokhwan Kim, Gary Geunbae Lee |
Speech Commun. | 4 |
| 2008 | Improving Speech Recognition and Understanding using Error-Corrective RerankingabstractThe main issues of practical spoken-language applications for human-computer interface are how to overcome speech recognition errors and guarantee the reasonable end-performance of spoken-language applications. Therefore, handling the erroneously recognized outputs is a key in developing robust spoken-language systems. To address this problem, we present a method to improve the accuracy of speech recognition and performance of spoken-language applications. The proposed error corrective reranking approach exploits recognition environment characteristics and domain-specific semantic information to provide robustness and adaptability for a spoken-language system. We demonstrate some experiments of spoken dialogue tasks and empirical results that show an improvement in accuracy for both speech recognition and spoken-language understanding. In our experiment, we show an error reduction of up to 9.7% and 16.8%; of word error rate, and 5.5% and 7.9% of understanding error for the air travel and telebanking service domains. Minwoo Jeong, Gary Geunbae Lee |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2008 | Triangular-Chain Conditional Random FieldsabstractSequential modeling is a fundamental task in scientific fields, especially in speech and natural language processing, where many problems of sequential data can be cast as a sequential labeling or a sequence classification. In many applications, the two problems are often correlated, for example named entity recognition and dialog act classification for spoken language understanding. This paper presents triangular-chain conditional random fields (CRFs), a unified probabilistic model combining two related problems. Triangular-chain CRFs jointly represent the sequence and meta-sequence labels in a single graphical structure that both explicitly encodes their dependencies and preserves uncertainty between them. An efficient inference and parameter estimation method is described for triangular-chain CRFs by extending linear-chain CRFs. This method outperforms baseline models on synthetic data and real-world dialog data for spoken language understanding. Minwoo Jeong, Gary Geunbae Lee |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | A Joint Statistical Model for Simultaneous Word Spacing and Spelling Error Correction for Korean
Hyungjong Noh, Jeongwon Cha, Gary Geunbae Lee |
ACL | 3 |
| 2007 | Example-based error recovery strategy for spoken dialog systemabstractError handling has become an important issue in spoken dialog systems. We describe an example-based approach to detect and repair errors in an example-based dialog modeling framework. Our approach to error recovery is focused on the re-phrase strategy with a system and a task guidance to help the novice users to re-phrase well-recognizable and well-understandable input. The dialog system gives possible utterance templates and contents related to the current situation when errors are detected. An empirical evaluation of the car navigation system shows that our approach is effective to the novice users for operating the spoken dialog system. Cheongjae Lee, Sangkeun Jung, Gary Geunbae Lee |
ASRU | 4 |
| 2007 | Structures for Spoken Language Understanding: A Two-Step ApproachabstractSpoken language understanding (SLU) aims to map a user's speech into a semantic frame. Since most of the previous works use the semantic structures for SLU, we verify that the structure is valuable even for noisy input. We apply a structured prediction method to SLU problem with comparison to unstructured one. In addition, we present a combined method to embed long-distance dependency between entities in a cascaded manner. On air travel data, we show that our approach improves performance over baseline models. Minwoo Jeong, Gary Geunbae Lee |
ICASSP (4) | 2 |
| 2007 | A semi-supervised method for efficient construction of statistical spoken language understanding resourcesabstractWe present a semi-supervised framework to construct spoken language understanding resources with very low cost. We generate context patterns with a few seed entities and a large amount of unlabeled utterances. Using these context patterns, we extract new entities from the unlabeled utterances. The extracted entities are appended to the seed entities, and we can obtain the extended entity list by repeating these steps. Our method is based on an utterance alignment algorithm which is a variant of the biological sequence alignment algorithm. Using this method, we can obtain precise entity lists with high coverage, which is of help to reduce the cost of building resources for statistical spoken language understanding systems. Index Terms: semi-supervised method, spoken language understanding 1. Seokhwan Kim, Minwoo Jeong, Gary Geunbae Lee |
INTERSPEECH | 3 |
| 2007 | Building ubiquitous and robust speech and natural language interfaces
Gary Geunbae Lee, Shimei Pan |
IUI | 1 |
| 2007 | Improving Speech Recognition Using Semantic and Reference Features in a Multimodal Dialog SystemabstractCurrent Speech-based dialog system undergo a practical problem; a speech recognizer is defective due to inevitable errors. Even in multimodal dialog systems, which have multiple input channels, errors in the speech recognition are a major problem because speech contains a large portion of user's intention. In this paper, we propose a re-ranking method to improve the performance of speech recognition in a multimodal dialog system. To re-rank the n-best speech recognition hypotheses, we use the multimodal understanding features that are orthogonal to the speech as well as the speech recognizer features. We demonstrate our method to smart home domain, and the results show that the multimodal understanding features are promising in overcoming many speech errors. Kyungduk Kim, Minwoo Jeong, Gary Geunbae Lee |
RO-MAN | 3 |
| 2007 | A Spoken Dialogue System for Electronic Program Guide Information AccessabstractIn this paper, we present POSTECH Spoken Dialogue System for Electronic Program Guide Information Access (POSSDS-EPG). POSSDS-EPG consists of automatic speech recognizer, spoken language understanding, dialogue manager, system utterance generator, text-to-speech synthesizer, and EPG database manager. Each module is designed and implemented to make an effective and practical spoken dialogue system. In particular, in order to reflect the up-to-date EPG information which is updated frequently and periodically, we applied a web-mining technology to the EPG database manager, which builds the content database based on automatically extracted information from popular EPG websites. The automatically generated content database is used by other modules in the system for building their own resources. Evaluations show that our system performs EPG access task in high performance and can be managed with low cost. Seokhwan Kim, Cheongjae Lee, Sangkeun Jung, Gary Geunbae Lee |
RO-MAN | 4 |
| 2007 | Emotion Recognition for Affective User Interfaces using Natural Language DialogsabstractIn a real world, emotion plays a significant role in rational actions in human communication. Given the potential and importance of emotions, in recent years, there has been growing interest in the study of emotions to improve the capabilities of current human-robot interaction. The emotion recognition from text modality is a necessary step to develop affective conversational interfaces. In this paper, we present an effective hybrid approach to improve the performance of emotion recognition from text by combining linguistic, pragmatic, and keyword spotting features. Cheongjae Lee, Gary Geunbae Lee |
RO-MAN | 2 |
| 2007 | Dependency structure language model for topic detection and tracking
Changki Lee, Gary Geunbae Lee, Myung-Gil Jang |
Inf. Process. Manag. | 2 |
| 2007 | Special issue on AIRS2005: Information retrieval research in Asia
Gary Geunbae Lee, Sung-Hyon Myaeng |
Inf. Process. Manag. | 1 |
| 2007 | Efficient implementation of associative classifiers for document classification
Yongwook Yoon, Gary Geunbae Lee |
Inf. Process. Manag. | 2 |
| 2007 | Exploring phrasal context and error correction heuristics in bootstrapping for geographic named entity annotation
Gary Geunbae Lee |
Inf. Syst. | 2 |
| 2006 | Exploiting Non-Local Features for Spoken Language Understanding
Minwoo Jeong, Gary Geunbae Lee |
ACL | 2 |
| 2006 | A Situation-Based Dialogue Management using Dialogue ExamplesabstractIn this paper, we present POSTECH Situation-Based Dialogue Manager (POSSDM) for a spoken dialogue system using both example- and rule-based dialogue management techniques for effective generation of appropriate system responses. A spoken dialogue system should generate cooperative responses to smoothly control dialogue flow with the users. We introduce a new dialogue management technique incorporating dialogue examples and situation-based rules for the electronic program guide (EPG) domain. For the system response generation, we automatically construct and index a dialogue example database from the dialogue corpus, and the proper system response is determined by retrieving the best dialogue example for the current dialogue situation, which includes a current user utterance, dialogue act, semantic frame and discourse history. When the dialogue corpus is not enough to cover the domain, we also apply manually constructed situation-based rules mainly for meta-level dialogue management. Experiments show that our example-based dialogue modeling is very useful and effective in domain-oriented dialogue processing Cheongjae Lee, Sangkeun Jung, Jihyun Eun, Minwoo Jeong, Gary Geunbae Lee |
ICASSP (1) | 5 |
| 2006 | C-TOBI-Based Pitch Accent Prediction Using Maximum-Entropy Model
Byeongchang Kim 0001, Gary Geunbae Lee |
ICCSA (3) | 2 |
| 2006 | Incorporating second-order information into two-step major phrase break prediction for KoreanabstractIn this paper, we present a new phrase break prediction method that integrates second-order information into general maximum entropy model. The phrase break prediction problem was mapped into a classification problem in our research. The features we used for the prediction of phrase breaks are of several layers such as local features (part-of-speech (POS) tags, a lexicon, lengths of eojeols 1 and location of juncture in the sentence), global features (chunk label derived from a eojeol parse tree) and second-order features (distance probability of previous and next phrase break). These three features were combined and used in the experiments, and we were able to generate good performance especially in the major phrase break prediction. Index Terms: phrase break, prosodic phrasing, speech synthesis, ToBI Seungwon Kim, Jinsik Lee, Byeongchang Kim 0001, Gary Geunbae Lee |
INTERSPEECH | 4 |
| 2006 | Grapheme-to-phoneme conversion using automatically extracted associative rules for Korean TTS systemabstractIn this paper, we describe a method for automatically extracting grapheme-to-phoneme conversion rules directly from the transcription of speech synthesis database and introduce a weighted score and jamo * similarity to overcome the rule application difficulties. We make a structured rule tree by rule pruning and rule association, and can eliminate most of the rules with almost no decrease of the performance. Our system achieves over 99.5 percent of phoneme-level accuracy and this performance is easily achievable even with the small amount of training data. Index Terms: grapheme-to-phoneme conversion, letter-tosound rule, text-to-speech system Jinsik Lee, Seungwon Kim, Gary Geunbae Lee |
INTERSPEECH | 3 |
| 2006 | Improving phrase-based Korean-English statistical machine translationabstractIn this paper, we describe several techniques to improve Korean-English statistical machine translation. We have built a phrase-based statistical machine translation system in a travel domain. On the baseline phrase-based system, several techniques are applied to improve the translation quality. Each technique can be applied or removed easily since the techniques are part of the preprocessing method or corpus processing method. Our experiments show that most of the techniques were successful except reordering the word sequence. The combination of the successful techniques has significantly improved the translation quality. Index Terms: statistical machine translation Gary Geunbae Lee |
INTERSPEECH | 3 |
| 2006 | Three phase verification for spoken dialog clarificationabstractSpoken dialog tasks incur many errors including speech recognition errors, understanding errors, and even dialog management errors. These errors create a big gap between user's will and the system's understanding, and eventually result in a misinterpretation. To fill in the gap, people in human-to-human dialog try to clarify the major causes of the misunderstanding and selectively correct them. This paper presents a method for applying the human's clarification techniques to human-machine spoken dialog systems. To increase the error detection precision and error recovery efficiency for the clarification dialogs, error detection phase is organized into three systematic phases and a clarification expert is devised for recovering the errors using the three phase verification. The experiment results demonstrate that the three phase verification could effectively catch the word and utterance-level errors in order to increase the SLU (spoken language understanding) performance and the clarification experts can actually increase the dialog success rate and the dialog efficiency. Sangkeun Jung, Cheongjae Lee, Gary Geunbae Lee |
IUI | 3 |
| 2006 | MMR-based Active Machine Learning for Bio Named Entity Recognition
Seokhwan Kim, Kyungduk Kim, Jeongwon Cha, Gary Geunbae Lee |
HLT-NAACL | 5 |
| 2006 | Jointly Predicting Dialog Act and Named Entity for spoken Language UnderstandingabstractSpoken language understanding (SLU) addresses the problem of mapping natural language speech into semantic frame for structure encoding of its meaning. Most of the SLU systems separate out the dialog act (DA) identification from the named entity (NE) recognition to generate the semantic frames. In previous works, these two subtasks are treated by independent or cascaded approaches. In the cascaded systems, however, DA and NE influence only to one side, rather than to both sides. In this paper, we develop a new joint SLU model with a triangular-chain conditional random field (CRF) to encode inter-dependence between DA and NE. On four real dialog data, we show that our joint approach outperforms both independent and cascaded approaches. Minwoo Jeong, Gary Geunbae Lee |
SLT | 2 |
| 2006 | Chat and Goal-Oriented Dialog Together: a Unified Example-Based Architecture for Multi-Domain Dialog ManagementabstractThis paper discusses development of a multi-domain conversational dialog system for simultaneously managing chats and goal-oriented dialogs. In this paper, we present a UMDM (unified multi-domain dialog manager) using a novel example-based dialog management technique. We have developed an effective utterance classifier with linguistic, semantic, and keyword features for domain switching and an example-based dialog modeling technique for domain-portable dialog models. Our experiments show that our approach is very useful and effective in multi-domain dialog system. Cheongjae Lee, Sangkeun Jung, Minwoo Jeong, Gary Geunbae Lee |
SLT | 4 |
| 2006 | Information gain and divergence-based feature selection for machine learning-based text categorization
Changki Lee, Gary Geunbae Lee |
Inf. Process. Manag. | 2 |
| 2006 | An effective procedure for constructing a hierarchical text classification systemabstractAbstract In text categorization tasks, classification on some class hierarchies has better results than in cases without the hierarchy. Currently, because a large number of documents are divided into several subgroups in a hierarchy, we can appropriately use a hierarchical classification method. However, we have no systematic method to build a hierarchical classification system that performs well with large collections of practical data. In this article, we introduce a new evaluation scheme for internal node classifiers, which can be used effectively to develop a hierarchical classification system. We also show that our method for constructing the hierarchical classification system is very effective, especially for the task of constructing classifiers applied to hierarchy tree with a lot of levels. Yongwook Yoon, Changki Lee, Gary Geunbae Lee |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2006 | Two-phase learning for biological event extraction and verificationabstractMany previous biological event-extraction systems were based on hand-crafted rules which were specifically tuned to a specific biological application domain. But manually constructing and tuning the rules are time-consuming processes and make the systems less portable. So supervised machine-learning methods were developed to generate the extraction rules automatically, but accepting the trade-off between precision and recall (high recall with low precision, and vice versa) is a barrier to improving performance. To make matters worse, a text in the biological domain is more complex because it often contains more than two biological events in a sentence, and one event in a noun chunk can be an entity for the other event. As a result, there are as yet no systems that give a good performance in extracting events in biological domains by using supervised machine learning.To overcome the limitations of previous systems and the complexity of biological texts, we present the following new ideas. First, we adopted a supervised machine-learning method to reduce the human effort in making extraction rules in order to obtain a highly domain-portable system. Second, we overcame the classical trade-off between precision and recall by using an event component verification method. Thus, machine learning occurs in two phases in our architecture. In the first phase, the system focuses on improving recall in extracting events between biological entities during a supervised machine-learning period. After extracting the biological events with automatically learned rules, in the second phase the system removes incorrect biological events by verifying the extracted event components with a maximum entropy (ME) classification method. In other words, the system targets for high recall in the first phase and tries to achieve high precision with a classifier in the second phase. Finally, we improved a supervised machine-learning algorithm so that it could learn a rule in a noun chunk and a rule extending throughout a sentence at two different levels, separately, for nested biological events. Eunju Kim, Cheongjae Lee, Kyungduk Kim, Gary Geunbae Lee, Byoung-Kee Yi, Jeongwon Cha |
ACM Trans. Asian Lang. Inf. Process. | 5 |
| 2006 | AUTHOR: Text mining and management in biomedicineabstract10.1145/1131348.1131349 Jong-Chan Park, Gary Geunbae Lee, Limsoon Wong |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2005 | Heuristic Methods for Reducing Errors of Geographic Named Entities Learned by Bootstrapping
Gary Geunbae Lee |
IJCNLP | 2 |
| 2005 | A multiple classifier-based concept-spotting approach for robust spoken language understandingabstractIn this paper, we present a concept spotting approach using manifold machine learning techniques for robust spoken language understanding. The goal of this approach is to find proper values for pre-defined slots of given meaning representation. Especially we propose a voting-based selection using multiple classifiers for robust spoken language understanding. This approach proposes no full level of language understanding but partial understanding because the method is only interested in the pre-defined meaning representation slots. In spite of this partial understanding, we can acquire necessary information to make interesting applications from the slot values because the slots are properly designed for specific domain-oriented understanding tasks. In several experimental results, the SLU (Spoken Language Understanding) performance degradation of spoken inputs compared with textual inputs are only F-measure 10.72, 11.43 and 11.51 for speech act, main goal and component slot extraction task respectively although the WER of spoken inputs is as high as 18.71%. That is, the evaluation results show that our concept spotting approach for SLU system is especially robust for spoken language input which has large recognition errors. 1. Jihyun Eun, Minwoo Jeong, Gary Geunbae Lee |
INTERSPEECH | 3 |
| 2005 | An error-corrective language-model adaptation for automatic speech recognitionabstractWe present a new language model adaptation framework integrated with error handling method to improve accuracy of speech recognition and performance of spoken language applications. The proposed error corrective language model adaptation approach exploits domain-specific language variations and recognition environment characteristics to provide robustness and adaptability for a spoken language system. We demonstrate some experiments of spoken dialogue tasks and empirical results which show an improvement of the accuracy for both speech recognition and spoken language understanding. 1. Minwoo Jeong, Jihyun Eun, Sangkeun Jung, Gary Geunbae Lee |
INTERSPEECH | 4 |
| 2005 | POSBIOTM-NER: a trainable biomedical named-entity recognition systemabstractSUMMARY: POSBIOTM-NER is a trainable biomedical named-entity recognition system. POSBIOTM-NER can be automatically trained and adapted to new datasets without performance degradation, using CRF (conditional random field) machine learning techniques and automatic linguistic feature analysis. Currently, we have trained our system on three different datasets. GENIA-NER was trained based on GENIA Corpus, GENE-NER based on BioCreative data and GPCR-NER based on our own POSBIOTM/NE corpus, respectively, which would be used in GPCR-related pathway extraction. Eunju Kim, Gary Geunbae Lee, Byoung-Kee Yi |
Bioinform. | 3 |
| 2005 | nformation extraction with automatic knowledge expansion
Hanmin Jung, Eunji Yi, Dongseok Kim, Gary Geunbae Lee |
Inf. Process. Manag. | 4 |
| 2005 | Probabilistic information retrieval model for a dependency structured indexing system
Changki Lee, Gary Geunbae Lee |
Inf. Process. Manag. | 2 |
| 2004 | High Speed Unknown Word Prediction Using Support Vector Machine for Chinese Text-to-Speech Systems
Juhong Ha, Byeongchang Kim 0001, Gary Geunbae Lee, Yoon-Suk Seong |
IJCNLP | 4 |
| 2004 | SVM-Based Biological Named Entity Recognition Using Minimum Edit-Distance Feature Boosted by Virtual Examples
Eunji Yi, Gary Geunbae Lee, Soo-Jun Park |
IJCNLP | 2 |
| 2004 | Systematic Construction of Hierarchical Classifier in SVM-Based Text Categorization
Yongwook Yoon, Changki Lee, Gary Geunbae Lee |
IJCNLP | 3 |
| 2004 | An information extraction approach for spoken language understandingabstractThis paper presents an Information Extraction (IE) approach for spoken language understanding. The goal in IE is to find proper values for pre-defined slots of given templates. IE for spoken language understanding proposes a concept spotting approach for spoken language because IE approach is interested in only pre-defined concept slots. In spite of this partial understanding, we can acquire necessary information for an application from the values of pre-defined slots because the slots are properly designed for speech understanding in a specific domain. Spoken language has so many recognition errors especially in a poor environment so it is more difficult to understand than textual language. Considering this fact, we attempt to understand the languages by concentrating on the specified information. In experiments on the car navigation domain, F-measure for concept spotting for textual input (WER 0%) and spoken input (WER 39%) are 96.33 % and 78.30 % respectively. 1. Jihyun Eun, Changki Lee, Gary Geunbae Lee |
INTERSPEECH | 3 |
| 2004 | High quality text-to-pinyin conversion using two-phase unknown word predictionabstractOne of the enduring problems in developing high-quality Chinese TTS (text-to-speech) systems is accurate text-to-Pinyin conversion. To solve the problem, identification of words and assignment of correct POS (Part-of-Speech) tags for an input sentence are very important tasks. Also, determining the correct Pinyin of polyphonic characters in unknown words is an important problem in a Chinese TTS system. The unknown word problem has significant effects on the accuracy of the synthesized sound, so accurately predicting the category of unknown words can help a Chinese TTS system to pronounce more naturally. In this paper, we present an SVM(support vector machine)-based method that predicts the unknown words for the results of Chinese word segmentation and POS tagging. For high speed SVM processing to be used in TTS, we pre-detect the candidate boundary of the unknown words before starting the actual prediction. Results of the experiments are very promising by showing high precision and high recall with also very high speed. Juhong Ha, Gary Geunbae Lee, Yoon-Suk Seong, Byeongchang Kim 0001 |
INTERSPEECH | 3 |
| 2004 | Speech recognition error correction using maximum entropy language modelabstractA speech interface is often required in many application environments, such as telephone-based information retrieval, car navigation systems, and user-friendly interfaces, but the low speech recognition rate makes it difficult to extend its application to new fields. We propose a domain adaptation technique via error correction with a maximum entropy language model, which is a general and elegant framework to combine higher level linguistic knowledge. Our approach has the ability to correct both semantic and lexical errors in 1-best output from the black-box style speech recognizer, and can improve the performance of speech recognition and application system. Through extensive experiments using a speechdriven in-vehicle telematics information retrieval and spoken language understanding, we demonstrate the superior performance of our approach and some advantages over previous lexical-oriented error correction approaches. Sangkeun Jung, Minwoo Jeong, Gary Geunbae Lee |
INTERSPEECH | 3 |
| 2004 | Using multiple linguistic features for Mandarin phrase break prediction in maximum-entropy classification frameworkabstractAbstract We model Mandarin phrase break prediction as a classification problem with three level prosodic structures and apply conditional maximum entropy classification to this problem. We acquire multiple levels of linguistic knowledge from an annotated corpus to become well-integrated features for maximum entropy framework. Five kinds of features were used to represent various linguistic constraints including POS tag features, lexical features, phonetic features, length features, and distance features. Experiment results show that our method performs better than the previous methods and the conditional maximum entropy (ME) model is very effective for data sparseness problem in Mandarin phrase break prediction. 1. Introduction Assigning the appropriate phrase breaks in text-to-speech systems is important for naturalness and intelligibility. Linguistic researchers have shown that the spoken language is structured as a hierarchy of prosodic units, including phonological phrase, intonation phrase, and utterance [1]. However, the text language is often structured by syntactic units, such as words and phrases, which are not equivalent to prosodic ones. But we suppose the syntactic information would provide important cues for prosodic phrase prediction. Many techniques have been introduced to predict phrase break, such as using Recurrent Neural Network (RNN)[2], Hidden Markov Model (HMM)[3], POS bi-gram and CART [4] and rule learning with C4.5 or TBL [5][6]. Min Chu and Yao Qian [4] proposed CART-based approach in four-class prosodic structure and their method shows high accuracy of 83%. Zhao and Tao with others [5][6] proposed automatic rule-learning approach with two typical rule-learning algorithms (C4.5 and TBL). They report a higher accuracy of 87.9% where they used POS features, lexical features and length features. They added chunking features and achieved even better accuracy of 90.0%, but they only used two-class prosodic structure to evaluate the accuracy. We treat entire phrase break prediction as a classification problem and apply conditional maximum entropy (ME) model. Various linguistic information is represented in the form of features, and five kinds of features were used in our system including POS tag features, lexical features, phonetic features, length features, and distance features. One serious problem of Mandarin phrase break prediction is that we usually do not have a large sized phrase break annotated corpus. So, we have to acquire multiple levels of linguistic information from only a small sized annotated corpus. Here, the data sparseness is a critical problem in Mandarin phrase break prediction, and we show that ME framework is very effective to capture useful features in the sparse training data environment. The remainder of this paper is organized as follows: Section 2 briefly introduces the ME framework with some justification. In section 3, we present our prosodic phrase framework and the five kinds of features used in the system. The effectiveness of our proposed method is verified by the experimental results given in section 4. Finally, conclusions are provided in section 5. Gary Geunbae Lee, Byeongchang Kim 0001 |
INTERSPEECH | 2 |
| 2003 | Exploring term dependences in probabilistic information retrieval model
Bong-Hyun Cho, Changki Lee, Gary Geunbae Lee |
Inf. Process. Manag. | 3 |
| 2003 | Unsupervised learning of mDTD extraction patterns for Web text mining
Dongseok Kim, Hanmin Jung, Gary Geunbae Lee |
Inf. Process. Manag. | 3 |
| 2002 | Syllable-Pattern-Based Unknown-Morpheme Segmentation and Estimation for Hybrid Part-of-Speech Tagging of KoreanabstractMost errors in Korean morphological analysis and part-of-speech (POS) tagging are caused by unknown morphemes. This paper presents a syllable-pattern-based generalized unknown-morpheme-estimation method with POSTAG (POStech TAGger), 1 which is a statistical and rule-based hybrid POS tagging system. This method of guessing unknown morphemes is based on a combination of a morpheme pattern dictionary that encodes general lexical patterns of Korean morphemes with a posteriori syllable trigram estimation. The syllable trigrams help to calculate lexical probabilities of the unknown morphemes and are utilized to search for the best tagging result. This method can guess the POS tags of unknown morphemes regardless of their numbers and/or positions in an eojeol (a Korean spacing unit similar to an English word), which is not possible with other systems for tagging Korean. In a series of experiments using three different domain corpora, the system achieved a 97% tagging accuracy even though 10% of the morphemes in the test corpora were unknown. It also achieved very high coverage and accuracy of estimation for all classes of unknown morphemes. 1 The binary code of POSTAG is open to the public for research and evaluation purposes at http://nlp.postech.ac.kr/. Follow the link OpenResources→DownLoad. Gary Geunbae Lee, Jeongwon Cha, Jong-Hyeok Lee |
Comput. Linguistics | 1 |
| 2002 | Integrated multi-strategic Web document pre-processing for sentence and word boundary detection
Junhyeok Shim, Dongseok Kim, Jeongwon Cha, Gary Geunbae Lee, Jungyun Seo |
Inf. Process. Manag. | 4 |
| 2002 | Morpheme-based grapheme to phoneme conversion using phonetic patterns and morphophonemic connectivity informationabstractBoth dictionary-based and rule-based methods on grapheme-to-phoneme conversion have their own advantages and limitations. For example, a large sized phonetic dictionary and complex morphophonemic rules are required for the dictionary-based method and the LTS (letter to sound) rule-based method itself cannot model the complete morphophonemic constraints.This paper describes a grapheme-to-phoneme conversion method for Korean using a dictionary-based and rule-based hybrid method with a phonetic pattern dictionary and CCV (consonant consonant vowel) LTS (letter to sound) rules. The phonetic pattern dictionary, standing for the dictionary-based method, contains entries in the form of a morpheme pattern and its phonetic pattern. The patterns represent candidate phonological changes in left and right boundaries of morphemes. Obviously, the CCV LTS rules stand for the rule-based method. The rules are in charge of grapheme-to-phoneme conversion within morphemes.The conversion method consists of mainly two steps including morpheme to phoneme conversion and morphophonemic connectivity check, and two preprocessing steps including phrase break prediction and morpheme normalization. Phrase break prediction presumes phrase breaks using the stochastic method on part-of-speech (POS) information. Morpheme normalization is to replace non-Korean symbols with their corresponding standard Korean graphemes. In the morpheme-phoneticizing module, each morpheme in the phrase is converted into phonetic patterns by looking it up in the phonetic pattern dictionary. Graphemes within a morpheme are grouped into CCV units and converted into phonemes by the CCV LTS rules. The morphophonemic connectivity table supports grammaticality checking of the two adjacent phonetic morphemes.In experiments with a non-Korean symbol free corpus of 4,973 sentences, we achieved a 99.98% grapheme-to-phoneme conversion performance rate and a 99.0% sentence conversion performance rate. With a broadcast news corpus of 621 sentences, 99.7% of the graphemes and 86.6% of the sentences are correctly converted. The full Korean TTS (Text-to-Speech) system is now being implemented using this conversion method. Byeongchang Kim 0001, Gary Geunbae Lee, Jong-Hyeok Lee |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2002 | Automatic corpus-based tone and break-index prediction using K-ToBI representationabstractIn this article we present a prosody generation architecture based on K-ToBI (Korean Tone and Break Index) representation. ToBI is a multitier representation system based on linguistic knowledge that transcribes events in an utterance. The TTS (Text-To-Speech) system, which adopts ToBI as an intermediate representation, is known to exhibit higher flexibility, modularity, and domain/task portability compared to the direct prosody generation TTS systems. However, for practical-level performance, the cost of corpus preparation is very expensive because the ToBI labeled corpus is constructed manually by many prosody experts, and normally requires large amounts of data for statistical prosody modeling. Unlike previous ToBI-based systems, this article proposes a new method, which transcribes the K-ToBI labels in Korean speech completely automatically. We develop automatic corpus-based K-ToBI labeling tools and prediction methods based on several lexico-syntactic linguistic features for decision-tree induction. We demonstrate the performance of F0 generation from automatically predicted K-ToBI labels, and confirm that the performance is reasonably comparable to state-of-the-art direct prosody generation methods and previous ToBI-based methods. Byeongchang Kim 0001, Gary Geunbae Lee |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2001 | Automatic Corpus-based Tone Prediction using K-ToBI Representation
Byeongchang Kim 0001, Gary Geunbae Lee |
EMNLP | 3 |
| 2001 | A Corpus-Based Learning Method of Compound Noun Indexing Rules for Korean
Jee-Hyub Kim, Byung-Kwan Kwak, Gary Geunbae Lee, Jong-Hyeok Lee |
Inf. Retr. | 4 |
| 2000 | Structural disambiguation of morpho-syntactic categorial parsing for Korean
Jeongwon Cha, Gary Geunbae Lee |
COLING | 2 |
| 2000 | Decision-Tree based Error Correction for Statistical Phrase Break Prediction in Korean
Byeongchang Kim 0001, Gary Geunbae Lee |
COLING | 2 |
| 2000 | POSCAT: A Morpheme-based Speech Corpus Annotation Tool
Byeongchang Kim 0001, Jeongwon Cha, Gary Geunbae Lee |
LREC | 4 |
| 2000 | Cross-Language Text Retrieval by Query Translation Using Term ReweightingabstractIn a dictionary-based query translation for cross-language text retrieval, transfer ambiguity is one of the main causes of performance deterioration, but the problem has not received significant attention in this field. To resolve transfer ambiguity, this paper proposes a two-phase query translation based on term reweighting, which uses a bilingual transfer dictionary, originally designed for machine translation. In general, source language query terms each show some word association with others, so that their correct translations should be more likely to co-occur in target documents. Based on this simple intuition, the first phase discriminates more relevant target documents from the others. Using statistical and ranking information from the highly relevant documents, the second phase then converts a translated query vector into reweighted form to add an extra weight on probably correct target terms. In experiments, the results were remarkable: the proposed method achieved almost the same performance as the monolingual IR system, actually contributing to an improvement of precision by about 9% over a baseline system. In-Su Kang, Oh-Woog Kwon, Jong-Hyeok Lee, Gary Geunbae Lee |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 1997 | Multi-level post-processing for Korean character recognition using morphological analysis and linguistic evaluation
Gary Geunbae Lee, Jong-Hyeok Lee, JinHee Yoo |
Pattern Recognit. | 1 |
| 1996 | A Connectionist/Symbolic Dependency Parser for Free Word-Order Languages
Jong-Hyeok Lee, Taeseung Lee, Gary Geunbae Lee |
IEA/AIE | 3 |
| 1996 | Integrating connectionist, statistical and symbolic approaches for continuous spoken Korean processing
Gary Geunbae Lee, Jong-Hyeok Lee, Kyubong Park, Byung-Chang Kim |
ICSLP | 1 |
| 1995 | Goal/Plan Analysis via Distributed Semantic Representations in a Connectionist System
Michael G. Dyer, Gary Geunbae Lee |
Appl. Intell. | 2 |
| 1994 | Table-driven Neural Syntactic Analysis of Spoken Korean
Wonil Lee, Gary Geunbae Lee, Jong-Hyeok Lee |
COLING | 2 |
| 1994 | Integrating TDNN-based diphone recognition with table-driven morphology parsing for understanding of spoken Korean
Kyunghee Kim, Gary Geunbae Lee, Jong-Hyeok Lee, Hong Jeong |
ICSLP | 2 |