VLDB 2026 Research / reviewers in the wild / expert
Mi-Young Kim
dblp:95/2636
· DBLP profile ↗
23ranked-venue papers
8as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Feature-Level Interaction Explanations in Multimodal Transformers
Yeji Kim, Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel |
ICPR (14) | 3 |
| 2026 | Reason2Decide: Rationale-Driven Multi-Task LearningabstractDespite the wide adoption of Large Language Models (LLM)s, clinical decision support systems face a critical challenge: achieving high predictive accuracy while generating explanations aligned with the predictions. Current approaches suffer from exposure bias leading to misaligned explanations. We propose Reason2Decide, a two-stage training framework that addresses key challenges in self-rationalization, including exposure bias and task separation. In Stage-1, our model is trained on rationale generation, while in Stage-2, we jointly train on label prediction and rationale generation, applying scheduled sampling to gradually transition from conditioning on gold labels to model predictions. We evaluate Reason2Decide on three medical datasets, including a proprietary triage dataset and public biomedical QA datasets. Across model sizes, Reason2Decide outperforms other fine-tuning baselines and some zero-shot LLMs in prediction (F1) and rationale fidelity (BERTScore, BLEU, LLM-as-a-Judge). In triage, Reason2Decide is rationale source-robust across LLM-generated, nurse-authored, and nurse-post-processed rationales. In our experiments, while using only LLM-generated rationales in Stage-1, Reason2Decide outperforms other fine-tuning variants. This indicates that LLM-generated rationales are suitable for pretraining models, reducing reliance on human annotations. Remarkably, Reason2Decide achieves these gains with models 40x smaller than contemporary foundation models, making clinical reasoning more accessible for resource-constrained deployments while still providing explainable decision support. H. M. Quamran Hasan, Housam Khalifa Bashier Babiker, Jiayi Dai, Mi-Young Kim, Randy Goebel |
LREC | 4 |
| 2025 | An Overview of the COLIEE 2025 Competition: Legal Case Law and Statute Law Information Retrieval and EntailmentabstractWe summarize the 12th Competition on Legal Information Extraction and Entailment. In this edition, the competition included four tasks on case law and statute law, plus a new pilot task on Tort law. The case law component includes an information retrieval task (Task 1), and the confirmation of an entailment relation between an existing case and an unseen case (Task 2). The statute law component includes an information retrieval task (Task 3), and an entailment/question-answering task based on retrieved civil code statutes (Task 4). The new pilot task is tort prediction (TP) and its rationale extraction (RE). Randy Goebel, Yoshinobu Kano, Mi-Young Kim, Calum Kwan, Ken Satoh, Hiroaki Yamada 0002, Masaharu Yoshioka |
ICAIL | 3 |
| 2024 | Juris-Informatics: Law for AI and Law of AIabstractThis paper presents an outline of our research project developed at our research center for "Juris-Informatics". "Juris-Informatics" is a research field based on two main topics; "Law by AI" and "Law of AI". "Law by Ai" is a research field where we investigate a support tool by AI for legal activities such as legal reasoning and legal document processing. "Law of AI" is a research field where we conduct research on legal control of AI such as considering the legal responsibility of AI and legal compliance of AI. Ken Satoh, Hideaki Takeda 0001, Randy Goebel, Yoshinobu Kano, Mi-Young Kim, Juliano Rabelo 0001, Masaharu Yoshioka |
IEEE Big Data | 5 |
| 2023 | From Intermediate Representations to Explanations: Exploring Hierarchical Structures in NLPabstractInterpretation methods for learned models used in natural language processing (NLP) applications usually provide support for local (specific) explanations, such as quantifying the contribution of each word to the predicted class. But they typically ignore the potential interaction amongst those word tokens. Unlike currently popular methods, we propose a deep model which uses feature attribution and identification of dependencies to support the learning of interpretable representations that will support creation of hierarchical explanations. In addition, hierarchical explanations provide a basis for visualizing how words and phrases are combined at different levels of abstraction, which enables end-users to better understand the prediction process of a deep network. Our study uses multiple well-known datasets to demonstrate the effectiveness of our approach, and provides both automatic and human evaluation. Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel |
ECAI | 2 |
| 2023 | Summary of the Competition on Legal Information, Extraction/Entailment (COLIEE) 2023abstractWe summarize the 10th Competition on Legal Information Extraction and Entailment. In this edition, the competition included four tasks on case law and statute law. The case law component includes an information retrieval task (Task 1), and the confirmation of an entailment relation between an existing case and an unseen case (Task 2). The statute law component includes an information retrieval task (Task 3), and an entailment/question answering task based on retrieved civil code statutes (Task 4). Participation was open to any group based on any approach. Ten different teams participated in the case law competition tasks, most of them in more than one task. We received results from 8 teams for Task 1 (22 runs) and seven teams for Task 2 (18 runs). On the statute law task, there were 9 different teams participating, most in more than one task. 6 teams submitted a total of 16 runs for Task 3, and 9 teams submitted a total of 26 runs for Task 4. We describe the variety of approaches, our official evaluation, and analysis of our data and submission results. Randy Goebel, Yoshinobu Kano, Mi-Young Kim, Juliano Rabelo 0001, Ken Satoh, Masaharu Yoshioka |
ICAIL | 3 |
| 2022 | Locally Distributed Activation Vectors for Guided Feature AttributionabstractExplaining the predictions of a deep neural network (DNN) is a challenging problem. Many attempts at interpreting those predictions have focused on attribution-based methods, which assess the contributions of individual features to each model prediction. However, attribution-based explanations do not always provide faithful explanations to the target model, e.g., noisy gradients can result in unfaithful feature attribution for back-propagation methods. We present a method to learn explanations-specific representations while constructing deep network models for text classification. These representations can be used to faithfully interpret black-box predictions, i.e., highlighting the most important input features and their role in any particular prediction. We show that learning specific representations improves model interpretability across various tasks, for both qualitative and quantitative evaluations, while preserving predictive performance. Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel |
COLING | 2 |
| 2022 | Neural Networks with Feature Attribution and Contrastive Explanations
Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel |
ECML/PKDD (1) | 2 |
| 2022 | Examining the structural relationships among e-learning interactivity, uncertainty avoidance, and perceived risks of COVID-19: Applying extended technology acceptance modelabstractAlthough e-learning has been studied in different contexts, the intention to use e-learning during the Corona virus-19 (COVID-19) pandemic has yet to be explored. This study extends the technology acceptance model (TAM) by treating e-learning interactivity as the antecedent construct of TAM. The constructs of uncertainty avoidance and perceived risks of COVID-19 were also added to understand the attitude and behavioral intentions of students to use e-learning. Two hundred and eighty-eight students from India participated in this study and partial least squares structural equation modeling was used to analyze the data. The results revealed that e-learning offers a high level of interactivity, which influences perceived ease of use and perceived usefulness and enhances a positive attitude of the students. Perceived ease of use was directly found to influence the perceived usefulness of e-learning. Uncertainty avoidance had a positive effect on attitude and intention to use e-learning. Perceived risks of COVID-19 showed a positive and significant effect on attitude, whereas the former did not affect intention to use e-learning. Finally, attitude also influenced intention to use e-learning. Our findings provide academia with theoretical implications and practitioners with practical implications. V. G. Girish, Mi-Young Kim, Indira Sharma, Choong-Ki Lee |
Int. J. Hum. Comput. Interact. | 2 |
| 2021 | DISK-CSV: Distilling Interpretable Semantic Knowledge with a Class Semantic VectorabstractNeural networks (NN) applied to natural language processing (NLP) are becoming deeper and more complex, making them increasingly difficult to understand and interpret.Even in applications of limited scope on fixed data, the creation of these complex "black-boxes" creates substantial challenges for debugging, understanding, and generalization.But rapid development in this field has now lead to building more straightforward and interpretable models.We propose a new technique (DISK-CSV) to distill knowledge concurrently from any neural network architecture for text classification, captured as a lightweight interpretable/explainable classifier.Across multiple datasets, our approach achieves better performance than the target black-box.In addition, our approach provides better explanations than existing techniques. Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel |
EACL | 2 |
| 2020 | RANCC: Rationalizing Neural Networks via Concept ClusteringabstractWe propose a new self-explainable model for Natural Language Processing (NLP) text classification tasks.Our approach constructs explanations concurrently with the formulation of classification predictions.To do so, we extract a rationale from the text, then use it to predict a concept of interest as the final prediction.We provide three types of explanations: 1) rationale extraction, 2) a measure of feature importance, and 3) clustering of concepts.In addition, we show how our model can be compressed without applying complicated compression techniques.We experimentally demonstrate our explainability approach on a number of well-known text classification datasets. Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel |
COLING | 2 |
| 2019 | Statute Law Information Retrieval and EntailmentabstractOur Yes/No statute law question answering system combines components for both statute law information retrieval and confirmation of textual entailment between statues and legal questions. We describe a statute law question answering system that exploits TF-IDF and a language model for information retrieval, and inter-paragraph entailment. We have evaluated our system using the data from the competition on legal information extraction/entailment (COLIEE-2019). The competition consists of four tasks: Tasks 1 and 2 are for the case law information extraction/entailment, and Tasks 3 and 4 are for the statute law information extraction/entailment. Here we explain our methods and evaluation results for Tasks 3 and 4. Task 3 requires the identification of civil law articles relevant to Japan legal bar exam query. For this task, we used TF-IDF and language model-based information retrieval approaches. Task 4 requires a decision on yes/no answer for previously unseen queries given relevant civil law articles. Our approach compares the approximate meanings of queries with relevant articles. Because many statute law and queries consist of more than one paragraph, we need an inter-paragraph entailment method. Our inter-paragraph entailment process exploits an analysis of statute law structure, and negation patterns to predict entailments. Using our heuristic selection of attributes, we perform two experiments which provide the basis for making a decision on the yes/no questions. One experiment uses an SVM model, and the other uses a general heuristic rule. Our experimental evaluation demonstrates the value of our method, and the results show that our method was ranked No. 1 in both of the Tasks 3 and 4 in COLIEE 2019. Mi-Young Kim, Juliano Rabelo 0001, Randy Goebel |
ICAIL | 1 |
| 2019 | Combining Similarity and Transformer Methods for Case Law EntailmentabstractWe tackle the complex problem of determining entailment relationships between case law documents, one of the tasks in the Competition on Legal Information Extraction and Entailment (COLIEE). With input of an entailed fragment from a case coupled with a candidate entailing paragraph from a noticed case, our approach relies on four main components: (1) extraction of similarity measures between the two pieces of text; (2) application of a transformer-based technique on the input text; (3) applying a threshold-based classifier; and (4) post-processing the results considering the a priori probability determined by the data distribution on the training samples and combining the results of (1) and (2). Our experiments achieved an F-score of 0.70 on the official COLIEE test dataset, ranking first among all competitors for that task in the 2019 competition. Juliano Rabelo 0001, Mi-Young Kim, Randy Goebel |
ICAIL | 2 |
| 2017 | Two-step cascaded textual entailment for legal bar exam question answeringabstractOur legal question answering system combines legal information retrieval and textual entailment, and exploits semantic information using a logic-based representation. We have evaluated our system using the data from the competition on legal information extraction/entailment (COLIEE)-2017. The competition focuses on the legal information processing required to answer yes/no questions from Japanese legal bar exams, and it consists of two phases: ad hoc legal information retrieval (Phase 1), and textual entailment (Phase 2). Phase 1 requires the identification of Japan civil law articles relevant to a legal bar exam query. For this phase, we have used an information retrieval approach using TF-IDF combined with a simple language model. Phase 2 requires a yes/no decision for previously unseen queries, which we approach by comparing the approximate meanings of queries with relevant statutes. Our meaning extraction process uses a selection of features based on a kind of paraphrase, coupled with a condition/conclusion/exception analysis of articles and queries. We also extract and exploit negation patterns from the articles. We construct a logic-based representation as a semantic analysis result, and then classify questions into easy and difficult types by analyzing the logic representation. If a question is in our easy category, we simply obtain the entailment answer from the logic representation; otherwise we use an unsupervised learning method to obtain the entailment answer. Experimental evaluation shows that our result ranked highest in the Phase 2 amongst all COLIEE-2017 competitors. Mi-Young Kim, Randy Goebel |
ICAIL | 1 |
| 2015 | Recognition of Patient-Related Named Entities in Noisy Tele-Health TextsabstractWe explore methods for effectively extracting information from clinical narratives that are captured in a public health consulting phone service called HealthLink. Our research investigates the application of state-of-the-art natural language processing and machine learning to clinical narratives to extract information of interest. The currently available data consist of dialogues constructed by nurses while consulting patients by phone. Since the data are interviews transcribed by nurses during phone conversations, they include a significant volume and variety of noise. When we extract the patient-related information from the noisy data, we have to remove or correct at least two kinds of noise: explicit noise , which includes spelling errors, unfinished sentences, omission of sentence delimiters, and variants of terms, and implicit noise , which includes non-patient information and patient's untrustworthy information. To filter explicit noise, we propose our own biomedical term detection/normalization method: it resolves misspelling, term variations, and arbitrary abbreviation of terms by nurses. In detecting temporal terms, temperature, and other types of named entities (which show patients’ personal information such as age and sex), we propose a bootstrapping-based pattern learning process to detect a variety of arbitrary variations of named entities. To address implicit noise, we propose a dependency path-based filtering method. The result of our denoising is the extraction of normalized patient information, and we visualize the named entities by constructing a graph that shows the relations between named entities. The objective of this knowledge discovery task is to identify associations between biomedical terms and to clearly expose the trends of patients’ symptoms and concern; the experimental results show that we achieve reasonable performance with our noise reduction methods. Mi-Young Kim, Ying Xu 0003, Osmar R. Zaïane, Randy Goebel |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2013 | Patient information extraction in noisy tele-health textsabstractWe explore methods for effectively extracting information from clinical narratives, which are captured in a public health consulting phone service called HealthLink. The currently available data consists of dialogues constructed by nurses while consulting patients on the phone. Since the data are interviews transcribed by nurses during phone conversations, they include a significant volume and variety of noise: First is explicit noise, which includes spelling errors, unfinished sentences, omission of sentence delimiters, variants of terms, etc. Second is implicit noise, which includes non-patient's information and negation of patient's information. To filter explicit noise, we propose our biomedical term detection/normalization method: it resolves misspelling, term variations, and arbitrary abbreviation of terms by nurses. In detecting temporal terms and other types of named entities (which show patients' personal information such as age, and sex), we propose a bootstrapping-based pattern learning to detect all kinds of arbitrary variations of the named entities. To address implicit noise, we propose a dependency path-based filtering method. The result of our denoising is the extraction of normalized patient information. The experimental results show that we achieve reasonable performance with our noise reduction methods. Mi-Young Kim, Ying Xu 0003, Osmar R. Zaïane, Randy Goebel |
BIBM | 1 |
| 2013 | Open Information Extraction with Tree Kernels
Ying Xu 0003, Mi-Young Kim, Kevin Quinn 0002, Randy Goebel, Denilson Barbosa 0001 |
HLT-NAACL | 2 |
| 2012 | Adaptive-capacity and robust natural language watermarking for agglutinative languagesabstractABSTRACT We present a robust and adaptive‐capacity watermarking algorithm for agglutinative languages. All processes, including the selection of sentences to be watermarked, watermark embedding, and watermark extraction, are based on syntactic dependency trees. We show that it is more robust to use syntactic dependency trees than the surface forms of sentences in text watermarking. For the agglutinative languages, we embed watermark using the two main characteristics of the languages. First, because a word consists of several morphemes, we can watermark sentences using morphological division/combination without deep linguistic analysis. Second, they permit relatively free word order, so we can move a syntactic constituent within its clause. Finally, to increase the information‐hiding capacity, we adaptively compute the number of watermark bits to be embedded for each sentence. We perform three kinds of evaluation: perceptibility, robustness, and capacity of our method. High capacity is achieved by dynamically determining possibly embedded watermark bits for each sentence. The secret rank based on a syntactic dependency tree strengthens robustness of our method. Finally, we show that the displacement of syntactic constituents and morphological division/combination does not affect the style and naturalness of the text. Copyright © 2011 John Wiley & Sons, Ltd. Mi-Young Kim, Randy Goebel |
Secur. Commun. Networks | 1 |
| 2009 | Natural Language Watermarking by Morpheme SegmentationabstractThis paper explores the method for Korean text watermarking and develops a morpheme-based scheme that a predicate nominal is segmented into a nominal and a predicate. Korean, as an agglutinative language, provides a good ground for the morpheme-based natural language watermarking. Korean word usually consists of a content morpheme and function morphemes. However, predicate nominal has two content morphemes--nominal and predicate. So, we propose a method to separate a predicate nominal into two words and assign a content morpheme into each of the new words. The division of a predicate nominal does not change the meaning of the sentence, and it also ensures the naturalness of the sentence.Our proposed natural language watermarking method consists of five procedures. First, we perform morphological analysis of unmarked text. Next, we choose target predicate nominals for division, and determine the division type. And then, we employ an insertion bit according to the division type. Third, we embed a watermark bit for each predicate nominal. Fourth, if the watermark bit does not correspond to the insertion bit, we divide the predicate nominal into two words. Finally, we obtain marked text. From the experimental results, we show that the rate of unnatural sentences in marked text is significantly lower than that of previous systems. Experimental results also show that the marked text keeps the same style, and it has the same information without semantic distortion. Mi-Young Kim |
ACIIDS | 1 |
| 2005 | Chunking Using Conditional Random Fields in Korean Texts
Yong-Hun Lee, Mi-Young Kim, Jong-Hyeok Lee |
IJCNLP | 2 |
| 2004 | Syntactic Analysis of Long Sentences Based on S-Clauses
Mi-Young Kim, Jong-Hyeok Lee |
IJCNLP | 1 |
| 2004 | T-shape Diamond Search Pattern for New Fast Block Matching Motion Estimation
Mi Gyoung Jung, Mi-Young Kim |
KES | 2 |
| 2004 | Motion Estimation Using Cross Center-Biased Distribution and Spatio-Temporal Correlation of Motion Vector
Mi-Young Kim, Mi Gyoung Jung |
KES | 1 |