VLDB 2026 Research / reviewers in the wild / expert
Seonjeong Hwang
dblp:329/6519
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-1196-2040ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Question answering and dialogue systems · 43% Language models and text generation · 21% Information extraction and text analysis · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computing education · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
question generation |
1.8 | 2 | 2026 | Difficulty-Controllable Cloze Question Distractor Generation · ACL (1) 2026 Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target Languages · EMNLP 2024 |
Natural language and speech › Question answering and dialogue systems › question generation › multiple-choice question generation
distractor generation |
1.0 | 1 | 2026 | Difficulty-Controllable Cloze Question Distractor Generation · ACL (1) 2026 |
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
legal question answering |
0.9 | 1 | 2025 | KoBLEX: Open Legal Question Answering with Multi-hop Reasoning · EMNLP 2025 |
Information retrieval
cross-language information retrieval |
0.9 | 1 | 2025 | MiLQ: Benchmarking IR Models for Bilingual Web Search with Mixed Language Queries · EMNLP 2025 |
Information retrieval
retrieval models |
0.9 | 1 | 2025 | KoBLEX: Open Legal Question Answering with Multi-hop Reasoning · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › semantic parsing
cross-lingual semantic parsing |
0.8 | 1 | 2024 | Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic Parsing · EMNLP 2024 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.8 | 1 | 2024 | Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target Languages · EMNLP 2024 |
Machine learning › Deep learning architectures and training
data augmentation |
0.8 | 1 | 2024 | Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic Parsing · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.8 | 1 | 2024 | Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic Parsing · EMNLP 2024 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.3 | 1 | 2026 | Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items? · ACL (1) 2026 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.3 | 1 | 2025 | KoBLEX: Open Legal Question Answering with Multi-hop Reasoning · EMNLP 2025 |
Natural language and speech › Language models and text generation › multilingual language models
multilingual pretrained language model |
0.2 | 1 | 2024 | Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic Parsing · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
multi-task learning · 2.0multi-agent prompting · 2.0large language model prompting · 2.0iterative revision · 2.0data augmentation · 2.0parametric provision-guided selection retrieval · 1.7legal fidelity evaluation · 1.7synthetic data generation · 0.8small language model · 0.8back-parsing · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?abstractEstimating the cognitive complexity of reading comprehension (RC) items is crucial for assessing item difficulty before it is administered to learners.Unlike syntactic and semantic features, such as passage length or semantic similarity between options, cognitive features that arise during answer reasoning are not readily extractable using existing NLP tools and have traditionally relied on human annotation.In this study, we examine whether large language models (LLMs) can estimate the cognitive complexity of RC items by focusing on two dimensions-Evidence Scope and Transformation Level-that indicate the degree of cognitive burden involved in reasoning about the answer.Our experimental results demonstrate that LLMs can approximate the cognitive complexity of items, indicating their potential as tools for prior difficulty analysis.Further analysis reveals a gap between LLMs' reasoning ability and their metacognitive awareness: even when they produce correct answers, they sometimes fail to correctly identify the features underlying their own reasoning process. Seonjeong Hwang, Hyounghun Kim, Gary Geunbae Lee |
ACL (1) | 1 |
| 2026 | A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item GenerationabstractRecent studies in difficulty-controlled reading comprehension item generation have leveraged large language models (LLMs) to produce items by adjusting difficulty-related features.However, existing methods typically rely on a single-agent prompting approach, which often fails to consistently satisfy specified feature constraints, resulting in items that deviate from the target difficulty level.To address this limitation, we introduce MAFIG, a Multiagent Framework for Feature-constrained Item Generation, where multiple LLM agents and feature-specific evaluators collaborate to generate and iteratively revise items based on intended constraints.Furthermore, to verify the efficacy of MAFIG in difficulty control, we propose a method for constructing a sequence of feature constraint sets that yield items with monotonically increasing difficulty.Experimental results demonstrate that MAFIG generates items that adhere to target constraints at a significantly higher rate than baselines, achieving robust difficulty control through the difficulty-calibrated constraint sequence. Seonjeong Hwang, Jun Seo, Hyounghun Kim, Gary Geunbae Lee |
ACL (1) | 1 |
| 2026 | Difficulty-Controllable Cloze Question Distractor GenerationabstractMultiple-choice cloze questions are commonly used to assess linguistic proficiency and comprehension.However, generating high-quality distractors remains challenging, as existing methods often lack adaptability and control over difficulty levels, and the absence of difficulty-annotated datasets further hinders progress.To address these issues, we propose a novel framework for generating distractors with controllable difficulty by leveraging both data augmentation and a multitask learning strategy.First, to create a high-quality, difficultyannotated dataset, we introduce a two-way distractor generation process to produce diverse and plausible distractors.These candidates are filtered and then categorized by difficulty using an ensemble QA system.Second, this newly created dataset is used to train a difficultycontrollable generation model via multitask learning.Experimental results demonstrate that our method generates high-quality distractors across difficulty levels and substantially outperforms GPT-4o in aligning distractor difficulty with human perception. Seokhoon Kang, Yejin Jeon, Seonjeong Hwang, Gary Geunbae Lee |
ACL (1) | 3 |
| 2025 | MiLQ: Benchmarking IR Models for Bilingual Web Search with Mixed Language QueriesabstractDespite bilingual speakers frequently using mixed-language queries in web searches, Information Retrieval (IR) research on them remains scarce.To address this, we introduce MiLQ, Mixed-Language Query test set, the first public benchmark of mixed-language queries, qualified as realistic and relatively preferred.Experiments show that multilingual IR models perform moderately on MiLQ and inconsistently across native, English, and mixed-language queries, also suggesting code-switched training data's potential for robust IR models handling such queries.Meanwhile, intentional English mixing in queries proves an effective strategy for bilinguals searching English documents, which our analysis attributes to enhanced token matching compared to native queries. 1 * This work was done when the author was at aiXplain 1 The code and data for this work are available at : https://github.com/jonghwi-kim/milq.2 In this study, code-switching, mixed-language, and codemixing are used synonymously.Was sind die Vorteile und Nachteile einer einheitlichen europäischen Währung?Was sind die Advantages und Disadvantages einer single European Currency?What are the advantages and disadvantages of a single European currency? Jonghwi Kim, Deokhyung Kang, Seonjeong Hwang, Yunsu Kim 0001, Jungseul Ok, Gary Geunbae Lee |
EMNLP | 3 |
| 2025 | KoBLEX: Open Legal Question Answering with Multi-hop ReasoningabstractLarge Language Models (LLM) have achieved remarkable performances in general domains and are now extending into the expert domain of law.Several benchmarks have been proposed to evaluate LLMs' legal capabilities.However, these benchmarks fail to evaluate open-ended and provisiongrounded Question Answering (QA).To address this, we introduce a Korean Benchmark for Legal EXplainable QA (KOBLEX), designed to evaluate provision-grounded, multihop legal reasoning.KOBLEX includes 226 scenario-based QA instances and their supporting provisions, created using a hybrid LLM-human expert pipeline.We also propose a method called Parametric provisionguided Selection Retrieval (PARSER), which uses LLM-generated parametric provisions to guide legally grounded and reliable answers.PARSER facilitates multi-hop reasoning on complex legal questions by generating parametric provisions and employing a three-stage sequential retrieval process.Furthermore, to better evaluate the legal fidelity of the generated answers, we propose Legal Fidelity Evaluation (LF-EVAL).LF-EVAL is an automatic metric that jointly considers the question, answer, and supporting provisions and shows a high correlation with human judgments.Experimental results show that PARSER consistently outperforms strong baselines, achieving the best results across multiple LLMs.Notably, compared to standard retrieval with GPT-4o, PARSER achieves 37.91 higher F-1 and 30.81 higher LF-EVAL.Further analyses reveal that PARSER efficiently delivers consistent performance across reasoning depths, with ablations confirming the effectiveness of PARSER. 1 * Equal Contribution. 1 The code and dataset are available at https://github. com/daehuikim/ Jihyung Lee, Daehui Kim, Seonjeong Hwang, Hyounghun Kim, Gary Geunbae Lee |
EMNLP | 3 |
| 2024 | Explainable Multi-hop Question Generation: An End-to-End Approach without Intermediate Question LabelingabstractIn response to the increasing use of interactive artificial intelligence, the demand for the capacity to handle complex questions has increased. Multi-hop question generation aims to generate complex questions that requires multi-step reasoning over several documents. Previous studies have predominantly utilized end-to-end models, wherein questions are decoded based on the representation of context documents. However, these approaches lack the ability to explain the reasoning process behind the generated multi-hop questions. Additionally, the question rewriting approach, which incrementally increases the question complexity, also has limitations due to the requirement of labeling data for intermediate-stage questions. In this paper, we introduce an end-to-end question rewriting model that increases question complexity through sequential rewriting. The proposed model has the advantage of training with only the final multi-hop questions, without intermediate questions. Experimental results demonstrate the effectiveness of our model in generating complex questions, particularly 3- and 4-hop questions, which are appropriately paired with input answers. We also prove that our model logically and incrementally increases the complexity of questions, and the generated multi-hop questions are also beneficial for training question answering models. Seonjeong Hwang, Yunsu Kim 0001, Gary Geunbae Lee |
LREC/COLING | 1 |
| 2024 | Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target LanguagesabstractAutomatic question generation (QG) serves a wide range of purposes, such as augmenting question-answering (QA) corpora, enhancing chatbot systems, and developing educational materials.Despite its importance, most existing datasets predominantly focus on English, resulting in a considerable gap in data availability for other languages.Cross-lingual transfer for QG (XLT-QG) addresses this limitation by allowing models trained on high-resource language datasets to generate questions in lowresource languages.In this paper, we propose a simple and efficient XLT-QG method that operates without the need for monolingual, parallel, or labeled data in the target language, utilizing a small language model.Our model, trained solely on English QA datasets, learns interrogative structures from a limited set of question exemplars, which are then applied to generate questions in the target language.Experimental results show that our method outperforms several XLT-QG baselines and achieves performance comparable to GPT-3.5-turbo across different languages.Additionally, the synthetic data generated by our model proves beneficial for training multilingual QA models.With significantly fewer parameters than large language models and without requiring additional training for target languages, our approach offers an effective solution for QG and QA tasks across various languages 1 . Seonjeong Hwang, Yunsu Kim 0001, Gary Geunbae Lee |
EMNLP | 1 |
| 2024 | Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic ParsingabstractRecent efforts have aimed to utilize multilingual pretrained language models (mPLMs) to extend semantic parsing (SP) across multiple languages without requiring extensive annotations.However, achieving zero-shot cross-lingual transfer for SP remains challenging, leading to a performance gap between source and target languages.In this study, we propose Cross-lingual Back-Parsing (CBP), a novel data augmentation methodology designed to enhance cross-lingual transfer for SP.Leveraging the representation geometry of the mPLMs, CBP synthesizes target language utterances from source meaning representations.Our methodology effectively performs cross-lingual data augmentation in challenging zero-resource settings, by utilizing only labeled data in the source language and monolingual corpora.Extensive experiments on two cross-lingual SP benchmarks (Mschema2QA and Xspider) demonstrate that CBP brings substantial gains in the target language.Further analysis of the synthesized utterances shows that our method successfully generates target language utterances with high slot value alignment rates while preserving semantic integrity. Deokhyung Kang, Seonjeong Hwang, Yunsu Kim 0001, Gary Geunbae Lee |
EMNLP | 2 |
| 2022 | Conversational QA Dataset Generation with Answer RevisionabstractConversational question-answer generation is a task that automatically generates a large-scale conversational question answering dataset based on input passages. In this paper, we introduce a novel framework that extracts question-worthy phrases from a passage and then generates corresponding questions considering previous conversations. In particular, our framework revises the extracted answers after generating questions so that answers exactly match paired questions. Experimental results show that our simple answer revision approach leads to significant improvement in the quality of synthetic data. Moreover, we prove that our framework can be effectively utilized for domain adaptation of conversational question answering. Seonjeong Hwang, Gary Geunbae Lee |
COLING | 1 |