EDBT 2026 Demo / reviewers in the wild / expert
Jae Won Lee
dblp:85/4031
· DBLP profile ↗
27ranked-venue papers
6as first author
1since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7Databases, data management, data science and information retrieval · 4 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Information extraction and text analysis · 48% Vision and language · 48% Reinforcement learning · 2% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 83% Computational finance and economics · 17% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › document analysis
document parsing |
1.0 | 1 | 2026 | Layout-Aware Document Parsing with Visual-Linguistic Fusion: The DATA-LUX with Academic Content Service Provider · AAAI 2026 |
Computer vision › Vision and language › multimodal fusion
text-image fusion |
1.0 | 1 | 2026 | Layout-Aware Document Parsing with Visual-Linguistic Fusion: The DATA-LUX with Academic Content Service Provider · AAAI 2026 |
Information retrieval
document processing |
0.3 | 1 | 2026 | Layout-Aware Document Parsing with Visual-Linguistic Fusion: The DATA-LUX with Academic Content Service Provider · AAAI 2026 |
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis |
0.1 | 1 | 2008 | OutlierD: an R package for outlier detection using quantile regression on mass spectrometry data · Bioinform. 2008 |
Bioinformatics and computational biology
proteomics |
0.1 | 1 | 2008 | OutlierD: an R package for outlier detection using quantile regression on mass spectrometry data · Bioinform. 2008 |
Bioinformatics and computational biology › computational oncology
cancer classification from gene expression |
0.1 | 1 | 2006 | Structured polychotomous machine diagnosis of multiple cancer types using gene expression · Bioinform. 2006 |
Bioinformatics and computational biology
gene expression analysis |
0.1 | 1 | 2006 | Structured polychotomous machine diagnosis of multiple cancer types using gene expression · Bioinform. 2006 |
Bioinformatics and computational biology › gene expression analysis › gene expression classification
marker gene selection |
0.1 | 1 | 2006 | Structured polychotomous machine diagnosis of multiple cancer types using gene expression · Bioinform. 2006 |
Knowledge, reasoning and agents › Multi-agent systems
cooperative agents |
0.0 | 1 | 2002 | Stock Trading System Using Reinforcement Learning with Cooperative Agents · ICML 2002 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.0 | 1 | 2002 | Stock Trading System Using Reinforcement Learning with Cooperative Agents · ICML 2002 |
Computational finance and economics
algorithmic trading |
0.0 | 1 | 2002 | Stock Trading System Using Reinforcement Learning with Cooperative Agents · ICML 2002 |
Computational finance and economics › algorithmic trading
reinforcement learning for trading |
0.0 | 1 | 2002 | Stock Trading System Using Reinforcement Learning with Cooperative Agents · ICML 2002 |
Methods — techniques the papers use, named apart from their topics
transformer · 2.0layout refinement · 2.0OCR · 2.0quantile regression · 0.1reinforcement learning · 0.1wald test · 0.1support vector machine · 0.1structured kernels · 0.1rao test · 0.1newton-raphson · 0.1analysis of variance decomposition · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Layout-Aware Document Parsing with Visual-Linguistic Fusion: The DATA-LUX with Academic Content Service ProviderabstractMany organizations are increasingly relying on unstructured documents such as PDFs and scanned forms to support downstream large language model (LLM) services, including search, summarization, and recommendation. However, traditional OCR systems struggle with diverse layouts of documents, leading to frequent errors and high costs of labor. So, this study developed DATALUX - a robust document layout system that trans-forms unstructured documents into structured, machine-readable data suitable for automation. Built on a trans-former-based detector, DATALUX incorporates several modules for layout refinement, text-visual fusion, and layer-wise optimization to improve coherence and generalization across diverse layouts. Around January 2025, we successfully deployed DATALUX into one of the largest academic content service firms (Nurimedia) in South Korea. This firm faced the challenge of extracting metadata and references from thousands of academic pa-pers submitted in various formats. Also, the existing LLM-based tools provided unreliable results. So, they needed to process them manually, creating bottlenecks in both labor and time. However, DATALUX enabled the automatic structuring of over 100,000 research papers a year, improving extraction accuracy to over 97%, reducing costs by more than USD 185K annually, and accelerating processing speed by 8.7 times. These deployment results suggest that DATALUX enables scalable and efficient document automation in complex and high-volume environments successfully. We thus believe that our DATALUX has a significant impact on both academia and industry practices. Min Chan Kim, Yeonkyung Kim, Jae Won Lee, Ki Hwan Kim, Ji Woo Kwak, Jae Hong Park |
AAAI | 3 |
| 2019 | Performance of MMSE-based Symbol-level Precoding in Multi-user MISO SystemabstractThe data-aided symbol-level precoding scheme has been known as a useful means of improving the system efficiency in multi-user MISO system. Even if it was mostly evaluated under the condition of an ideal channel state information (CSI), it is conjectured that its performance is critically governed by imperfect received CSI (CSIR). Furthermore, as the downlink channel can be estimated only in a block level, i.e., without tracking an effective channel on each symbol, its performance would be more seriously degraded by symbol-by-symbol channel variation. Meanwhile, we also investigate a channel coding gain for symbol-level preceding, which is not clear when the minimum distance between received signals can be increased by a nature of cumulative constructive interference gain (CCIG). The objective of this paper is to provide a complete performance comparison between the conventional block-level precoding and symbol-level precoding schemes under more realistic system assumptions. Jae Won Lee, Chung Gu Kang 0001 |
APCC | 1 |
| 2019 | System-level Performance Evaluation with 5G K-SimSys for 5G URLLC SystemabstractIn 5G mobile communication system design and standard specification, a system-level simulator is an indispensable element of research & development (R&D). Due to various usage scenarios, 5G systems require the development of different simulators that deal with individual evaluation scenarios and environments. To minimize the time and effort required for the implementation, we have designed and implemented 5G K-SimSys, which is a system-level simulator with modular & flexible structure, facilitating its reuse thorough open-source code sharing. This paper describes the 5G ultra-reliable low latency communication (URLLC) specification in 3GPP and its system-level evaluation to illustrate that 5G K-SimSys can provide reconfiguration between various usage scenarios. Minsig Han, Jae Won Lee, Chung Gu Kang 0001, Minjoong Rim |
CCNC | 2 |
| 2019 | 5G K-SimSys for System-level Evaluation of Massive MIMOabstractMassive MIMO is a key feature for improving spectral efficiency in 5G system. Since massive MIMO has been standardized in 3GPP Re1-13 (LTE-Advanced Pro), a next generation of massive MIMO standard is now available in Rel-15, a.k.a New Radio (NR), for 5G, which is aiming at further improving the average spectral efficiency with more antenna elements over the higher frequency band. In particular, beam-based air interface for above 6GHz band involves with various antenna configurations and feedback schemes, requiring a more complex testbed for system-level performance evaluation. In this paper, we examine the multi-antenna technologies in 3GPP NR specification, so as to develop its system-level model in 5G K-SimSys, which has been designed and implemented to provide an open platform and testbed for evaluating the system-level performance for 5G system. Its actual implementation is detailed along with the overall platform architecture for modular and flexible design. Then, its actual performance is demonstrated for one of the precoding matrices in the NR specification as an example. Jae Won Lee, Minsig Han, Chung Gu Kang 0001, Minjoong Rim |
CCNC | 1 |
| 2019 | A study on novel filtering and relationship between input-features and target-vectors in a deep learning model for stock price prediction
Yoojeong Song, Jae Won Lee, Jongwoo Lee |
Appl. Intell. | 2 |
| 2019 | Constructive Interference Optimization for Data-Aided Precoding in Multi-User MISO SystemsabstractUnlike the general concept of eliminating or avoiding inter-user interference in the downlink multi-user multiple-input single-output (MISO) system, the data-aided precoding scheme attempts to exploit the constructive interference at symbol level. Positive interference can be constructed to enhance the received signal gain by predicting the phase and magnitude of the inter-user interference. In this paper, we formulate a constructive interference optimization problem that minimizes a sum of minimum mean square error (MMSE) for all users, while ensuring the minimum required constructive interference gain with a fixed total power constraint. As opposed to the existing scheme, such as constructive zero-forcing precoding or minimum power precoding subject to strict phase conservation for phase-shift keying (PSK), our proposed scheme exploits a full range of relaxation for the constructive interference region, while ensuring the link performance by minimizing the sum of mean-square error (MSE) for all users. In fact, the relaxed requirements lead to more degrees of freedom for improving the cumulative constructive interference gain (CCIG) under the varying channel conditions. Furthermore, our MMSE-based optimization approach allows for a semi-closed-form optimal solution to the data-aided precoding scheme, providing a more CCIG without incurring unacceptable complexity than other state-of-the-art schemes. Yongin Choi, Jae Won Lee, Minjoong Rim, Chung Gu Kang 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2016 | Detecting high-dimensional genetic associations using a Markov-Blanket in a family-based studyabstractIn recent years, detecting interactions between different genes has become a hot topic, for better understanding multigenic, complex diseases. For population-based genome-wide association studies (GWAS), a number of methods to detect gene-gene interactions such as logistic regression, multifactor dimensionality reduction (MDR) and support vector machine (SVM), have been applied. Bayesian approaches such as BEAM (Bayesian marker partition model) and DASSO-MB (detection of association using Markov Blanket) have also been suggested. However, the studies for family-based GWAS have been limited. In this study, we developed a new Markov Blanket-based algorithm called MB-TDT to find gene-gene interactions for pedigree data. A transmission disequilibrium test statistic was used as an association measure and the incremental association a Markov Blanket (IAMB) algorithm was applied to find Markov Blanket. This proposed MB-TDT method can identify a minimal set of causal SNPs, associated with a specific disease, thus avoiding an exhaustive search. By conducting a simulation study to compare MB-TDT with current methods, we show its superior high power in many cases, and lower false positive rates, in others. Hyo Jung Lee, Jae Won Lee, Seohoon Jin, Hee Jeong Yoo, Mira Park 0002 |
BIBM | 2 |
| 2015 | Practical approach to determine sample size for building logistic prediction models using high-throughput data
Dae-Soon Son, DongHyuk Lee, Kyusang Lee, Sin-Ho Jung, TaeJin Ahn, Eunjin Lee, Insuk Sohn, Jongsuk Chung, Woong-Yang Park, Nam Huh, Jae Won Lee |
J. Biomed. Informatics | 11 |
| 2008 | OutlierD: an R package for outlier detection using quantile regression on mass spectrometry dataabstractUNLABELLED: It is important to preprocess high-throughput data generated from mass spectrometry experiments in order to obtain a successful proteomics analysis. Outlier detection is an important preprocessing step. A naive outlier detection approach may miss many true outliers and instead select many non-outliers because of the heterogeneity of the variability observed commonly in high-throughput data. Because of this issue, we developed a outlier detection software program accounting for the heterogeneous variability by utilizing linear, non-linear and non-parametric quantile regression techniques. Our program was developed using the R computer language. As a consequence, it can be used interactively and conveniently in the R environment. AVAILABILITY: An R package, OutlierD, is available at the Bioconductor project at http://www.bioconductor.org HyungJun Cho, Yang-jin Kim, Hee Jung Jung, Sang-Won Lee 0006, Jae Won Lee |
Bioinform. | 5 |
| 2007 | A Multiagent Approach to Q-Learning for Daily Stock TradingabstractThe portfolio management for trading in the stock market poses a challenging stochastic control problem of significant commercial interests to finance industry. To date, many researchers have proposed various methods to build an intelligent portfolio management system that can recommend financial decisions for daily stock trading. Many promising results have been reported from the supervised learning community on the possibility of building a profitable trading system. More recently, several studies have shown that even the problem of integrating stock price prediction results with trading strategies can be successfully addressed by applying reinforcement learning algorithms. Motivated by this, we present a new stock trading framework that attempts to further enhance the performance of reinforcement learning-based systems. The proposed approach incorporates multiple Q-learning agents, allowing them to effectively divide and conquer the stock trading problem by defining necessary roles for cooperatively carrying out stock pricing and selection decisions. Furthermore, in an attempt to address the complexity issue when considering a large amount of data to obtain long-term dependence among the stock prices, we present a representation scheme that can succinctly summarize the history of price changes. Experimental results on a Korean stock market show that the proposed trading framework outperforms those trained by other alternative approaches both in terms of profit and risk management. Jae Won Lee, Jangmin O, Jongwoo Lee, Euyseok Hong |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2006 | CSTallocator: Call-Site Tracing Based Shared Memory Allocator for False Sharing Reduction in Page-Based DSM Systems
Jongwoo Lee, Sung Dong Kim, Jae Won Lee, Jangmin O |
HPCC | 3 |
| 2006 | A Mutated Intrusion Detection System Using Principal Component Analysis and Time Delay Neural Network
Byoung-Doo Kang, Jae Won Lee, Jong-Ho Kim, O-Hwa Kwon, Chi-Young Seong, Se-Myung Park, Sang-Kyoon Kim |
ISNN (2) | 2 |
| 2006 | Hierarchical Classification of Object Images Using Neural Networks
Jong-Ho Kim, Jae Won Lee, Byoung-Doo Kang, O-Hwa Kwon, Chi-Young Seong, Sang-Kyoon Kim, Se-Myung Park |
ISNN (2) | 2 |
| 2006 | Music Genre Classification Using a Time-Delay Neural Network
Jae Won Lee, Soo Beom Park, Sang-Kyoon Kim |
ISNN (2) | 1 |
| 2006 | Structured polychotomous machine diagnosis of multiple cancer types using gene expressionabstractMOTIVATION: The problem of class prediction has received a tremendous amount of attention in the literature recently. In the context of DNA microarrays, where the task is to classify and predict the diagnostic category of a sample on the basis of its gene expression profile, a problem of particular importance is the diagnosis of cancer type based on microarray data. One method of classification which has been very successful in cancer diagnosis is the support vector machine (SVM). The latter has been shown (through simulations) to be superior in comparison with other methods, such as classical discriminant analysis, however, SVM suffers from the drawback that the solution is implicit and therefore is difficult to interpret. In order to remedy this difficulty, an analysis of variance decomposition using structured kernels is proposed and is referred to as the structured polychotomous machine. This technique utilizes Newton-Raphson to find estimates of coefficients followed by the Rao and Wald tests, respectively, for addition and deletion of import vectors. RESULTS: The proposed method is applied to microarray data and simulation data. The major breakthrough of our method is efficiency in that only a minimal number of genes that accurately predict the classes are selected. It has been verified that the selected genes serve as legitimate markers for cancer classification from a biological point of view. AVAILABILITY: All source codes used are available on request from the authors. Ja-Yong Koo, Insuk Sohn, Sujong Kim, Jae Won Lee |
Bioinform. | 4 |
| 2006 | Effect of data normalization on fuzzy clustering of DNA microarray dataabstractBACKGROUND: Microarray technology has made it possible to simultaneously measure the expression levels of large numbers of genes in a short time. Gene expression data is information rich; however, extensive data mining is required to identify the patterns that characterize the underlying mechanisms of action. Clustering is an important tool for finding groups of genes with similar expression patterns in microarray data analysis. However, hard clustering methods, which assign each gene exactly to one cluster, are poorly suited to the analysis of microarray datasets because in such datasets the clusters of genes frequently overlap. RESULTS: In this study we applied the fuzzy partitional clustering method known as Fuzzy C-Means (FCM) to overcome the limitations of hard clustering. To identify the effect of data normalization, we used three normalization methods, the two common scale and location transformations and Lowess normalization methods, to normalize three microarray datasets and three simulated datasets. First we determined the optimal parameters for FCM clustering. We found that the optimal fuzzification parameter in the FCM analysis of a microarray dataset depended on the normalization method applied to the dataset during preprocessing. We additionally evaluated the effect of normalization of noisy datasets on the results obtained when hard clustering or FCM clustering was applied to those datasets. The effects of normalization were evaluated using both simulated datasets and microarray datasets. A comparative analysis showed that the clustering results depended on the normalization method used and the noisiness of the data. In particular, the selection of the fuzzification parameter value for the FCM method was sensitive to the normalization method used for datasets with large variations across samples. CONCLUSION: Lowess normalization is more robust for clustering of genes from general microarray data than the two common scale and location adjustment methods when samples have varying expression patterns or are noisy. In particular, the FCM method slightly outperformed the hard clustering methods when the expression patterns of genes overlapped and was advantageous in finding co-regulated genes. Thus, the FCM approach offers a convenient method for finding subsets of genes that are strongly associated to a given cluster. Seo Young Kim, Jae Won Lee, Jong Sung Bae |
BMC Bioinform. | 2 |
| 2006 | Adaptive stock trading with dynamic asset allocation using reinforcement learning
Jangmin O, Jongwoo Lee, Jae Won Lee, Byoung-Tak Zhang |
Inf. Sci. | 3 |
| 2004 | Dynamic Asset Allocation Exploiting Predictors in Reinforcement Learning Framework
Jangmin O, Jae Won Lee, Jongwoo Lee, Byoung-Tak Zhang |
ECML | 2 |
| 2004 | Stock Trading by Modelling Price Trend with Dynamic Bayesian Networks
Jangmin O, Jae Won Lee, Sung-Bae Park, Byoung-Tak Zhang |
IDEAL | 2 |
| 2004 | Content-based image classification using a neural network
Soo Beom Park, Jae Won Lee, Sang-Kyoon Kim |
Pattern Recognit. Lett. | 2 |
| 2003 | Internet advertising strategy by comparison challenge approachabstractA comparison challenge approach is proposed as a form of challenger-activated, just-in-time advertising. To develop a framework for a comparison challenge, we propose a theory of comparison. Based on this theory, the CompareMe and CompareThem strategies are devised, and comparable objects are classified in terms of price and performance dominance as well as the scope of proximity. The idea is demonstrated with a comparison of PCs from five leading manufacturers. To assist in the planning of the comparison challenge, a mathematical programming model was formulated to maximize the value of comparison under the constraints of the comparison opportunity and budget. The model is applied to eight scenarios in terms of the range of comparing objects. We found the ad effect of comparison challenge to be substantially better than banners (4.75 times) and similarity-based comparisons (2.77 times), providing customers with better performance and reduced prices. Jae Kyu Lee, Jae Won Lee |
ICEC | 2 |
| 2002 | A Two-Phase Stock Trading System Using Distributional Differences
Sung Dong Kim, Jae Won Lee, Jongwoo Lee, Jinseok Chae |
DEXA | 2 |
| 2002 | A Multi-agent Q-learning Framework for Optimizing Stock Trading Systems
Jae Won Lee, Jangmin O |
DEXA | 1 |
| 2002 | Stock Trading System Using Reinforcement Learning with Cooperative Agents
Jangmin O, Jae Won Lee, Byoung-Tak Zhang |
ICML | 2 |
| 2002 | Topic Extraction from Text Documents Using Multiple-Cause Networks
Jeong Ho Chang, Jae Won Lee, Yuseop Kim, Byoung-Tak Zhang |
PRICAI | 2 |
| 2002 | Construction of Large-Scale Bayesian Networks by Local to Global Search
Kyu-Baek Hwang, Jae Won Lee, Seung-Woo Chung, Byoung-Tak Zhang |
PRICAI | 2 |
| 2002 | PATI: An Approach for Identifying and Resolving Ambiguities
Jae Won Lee, Sung Dong Kim |
PRICAI | 1 |