Kwangwook Seo

dblp:371/4619 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 33% Information extraction and text analysis · 25% Reinforcement learning · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model evaluation
checklist-based evaluation
1.012026
P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist · ACL (1) 2026
Machine learning › Trustworthy machine learning
interpretability
1.012026
P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model › large language model adaptation › personalization
personalized reward modeling
1.012026
P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist · ACL (1) 2026
Machine learning › Reinforcement learning › reward learning
reward modeling
1.012026
P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist · ACL (1) 2026
Information retrieval
evaluation
0.912025
MT-RAIG: Novel Benchmark and Evaluation Framework for Retrieval-Augmented Insight Generation over Multiple Tables · ACL (1) 2025
Information retrieval
retrieval-augmented generation
0.912025
MT-RAIG: Novel Benchmark and Evaluation Framework for Retrieval-Augmented Insight Generation over Multiple Tables · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › named entity recognition
biomedical named entity recognition
0.812024
VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models · ACL (1) 2024
Natural language and speech › Information extraction and text analysis
named entity recognition
0.812024
VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models · ACL (1) 2024
Natural language and speech › Question answering and dialogue systems › table question answering
table reasoning
0.312025
MT-RAIG: Novel Benchmark and Evaluation Framework for Retrieval-Augmented Insight Generation over Multiple Tables · ACL (1) 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge-based systems
knowledge-grounded reasoning
0.212024
VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models · ACL (1) 2024

Methods — techniques the papers use, named apart from their topics

large language model · 2.7retrieval-augmented generation · 1.7preference-contrastive criterion weighting · 1.0large language model reasoning · 0.8
YearPublicationVenuePosition
2026 P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist
abstract
Recent approaches in personalized reward modeling have primarily focused on leveraging user interaction history to align model judgments with individual preferences.However, existing approaches largely treat user context as a static or implicit conditioning signal, failing to capture the dynamic and multi-faceted nature of human judgment.In this paper, we propose P-CHECK, a novel personalized reward modeling framework, designed to train a plugand-play checklist generator that synthesizes dynamic evaluation criteria for guiding the reward prediction.To better align these checklists with personalized nuances, we introduce Preference-Contrastive Criterion Weighting, a training strategy that assigns saliency scores to criteria based on their discriminative power for personalized judgment.We conduct extensive experiments and demonstrate that P-CHECK not only improves reward accuracy but also enhances downstream personalized generation, and remains robust in OOD scenarios.[CODE]
Kwangwook Seo, Dongha Lee 0003
ACL (1)1
2025 MT-RAIG: Novel Benchmark and Evaluation Framework for Retrieval-Augmented Insight Generation over Multiple Tables
abstract
Recent advancements in table-based reasoning have expanded beyond factoid-level QA to address insight-level tasks, where systems should synthesize implicit knowledge in the table to provide explainable analyses. Although effective, existing studies remain confined to scenarios where a single gold table is given alongside the user query, failing to address cases where users seek comprehensive insights from multiple unknown tables. To bridge these gaps, we propose MT-RAIG Bench, design to evaluate systems on Retrieval-Augmented Insight Generation over Mulitple-Tables. Additionally, to tackle the suboptimality of existing automatic evaluation methods in the table domain, we further introduce a fine-grained evaluation framework MT-RAIG Eval, which achieves better alignment with human quality judgments on the generated insights. We conduct extensive experiments and reveal that even frontier LLMs still struggle with complex multi-table reasoning, establishing our MT-RAIG Bench as a challenging testbed for future research.
Kwangwook Seo, Donguk Kwon, Dongha Lee 0003
ACL (1)1
2024 VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models
abstract
Recent approaches in domain-specific named entity recognition (NER), such as biomedical NER, have shown remarkable advances.However, they still lack of faithfulness, producing erroneous predictions.We assume that knowledge of entities can be useful in verifying the correctness of the predictions.Despite the usefulness of knowledge, resolving such errors with knowledge is nontrivial, since the knowledge itself does not directly indicate the groundtruth label.To this end, we propose VER-IFINER, a post-hoc verification framework that identifies errors from existing NER methods using knowledge and revises them into more faithful predictions.Our framework leverages the reasoning abilities of large language models to adequately ground on knowledge and the contextual information in the verification process.We validate effectiveness of VERIFINER through extensive experiments on biomedical datasets.The results suggest that VERIFINER can successfully verify errors from existing models as a model-agnostic approach.Further analyses on out-of-domain and low-resource settings show the usefulness of VERIFINER on real-world applications.1
Kwangwook Seo, Hyungjoo Chae, Jinyoung Yeo, Dongha Lee 0003
ACL (1)2