EDBT 2026 Demo / reviewers in the wild / expert
Ki Hwan Kim
dblp:52/4596
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Information extraction and text analysis · 50% Vision and language · 50% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › document analysis
document parsing |
1.0 | 1 | 2026 | Layout-Aware Document Parsing with Visual-Linguistic Fusion: The DATA-LUX with Academic Content Service Provider · AAAI 2026 |
Computer vision › Vision and language › multimodal fusion
text-image fusion |
1.0 | 1 | 2026 | Layout-Aware Document Parsing with Visual-Linguistic Fusion: The DATA-LUX with Academic Content Service Provider · AAAI 2026 |
Information retrieval
document processing |
0.3 | 1 | 2026 | Layout-Aware Document Parsing with Visual-Linguistic Fusion: The DATA-LUX with Academic Content Service Provider · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
transformer · 2.0layout refinement · 2.0OCR · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Layout-Aware Document Parsing with Visual-Linguistic Fusion: The DATA-LUX with Academic Content Service ProviderabstractMany organizations are increasingly relying on unstructured documents such as PDFs and scanned forms to support downstream large language model (LLM) services, including search, summarization, and recommendation. However, traditional OCR systems struggle with diverse layouts of documents, leading to frequent errors and high costs of labor. So, this study developed DATALUX - a robust document layout system that trans-forms unstructured documents into structured, machine-readable data suitable for automation. Built on a trans-former-based detector, DATALUX incorporates several modules for layout refinement, text-visual fusion, and layer-wise optimization to improve coherence and generalization across diverse layouts. Around January 2025, we successfully deployed DATALUX into one of the largest academic content service firms (Nurimedia) in South Korea. This firm faced the challenge of extracting metadata and references from thousands of academic pa-pers submitted in various formats. Also, the existing LLM-based tools provided unreliable results. So, they needed to process them manually, creating bottlenecks in both labor and time. However, DATALUX enabled the automatic structuring of over 100,000 research papers a year, improving extraction accuracy to over 97%, reducing costs by more than USD 185K annually, and accelerating processing speed by 8.7 times. These deployment results suggest that DATALUX enables scalable and efficient document automation in complex and high-volume environments successfully. We thus believe that our DATALUX has a significant impact on both academia and industry practices. Min Chan Kim, Yeonkyung Kim, Jae Won Lee, Ki Hwan Kim, Ji Woo Kwak, Jae Hong Park |
AAAI | 4 |