EDBT 2026 Demo / reviewers in the wild / expert
Suyong Kwon
dblp:266/6114
· DBLP profile ↗
4ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-8972-5766ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Query processing and optimization · 84% Machine learning and data management · 7% Information retrieval · 5% | |
| Network and information security
2 papers |
Privacy and data protection · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
cardinality estimation |
1.4 | 2 | 2025 | Cardinality Estimation of LIKE Predicate Queries using Deep Learning · Proc. ACM Manag. Data 2025 Cardinality Estimation of Approximate Substring Queries using Deep Learning · Proc. VLDB Endow. 2022 |
Query processing and optimization › cardinality estimation
learned cardinality estimation |
1.4 | 2 | 2025 | Cardinality Estimation of LIKE Predicate Queries using Deep Learning · Proc. ACM Manag. Data 2025 Cardinality Estimation of Approximate Substring Queries using Deep Learning · Proc. VLDB Endow. 2022 |
Privacy and data protection
differential privacy |
0.9 | 2 | 2021 | TIDY: Publishing a Time Interval Dataset With Differential Privacy · IEEE Trans. Knowl. Data Eng. 2021 TIDY: Publishing a Time Interval Dataset with Differential Privacy (Extended abstract) · ICDE 2020 |
Privacy and data protection › differential privacy
differentially private data release |
0.5 | 1 | 2021 | TIDY: Publishing a Time Interval Dataset With Differential Privacy · IEEE Trans. Knowl. Data Eng. 2021 |
Privacy and data protection › differential privacy › differentially private data release
differentially private histogram |
0.4 | 1 | 2020 | TIDY: Publishing a Time Interval Dataset with Differential Privacy (Extended abstract) · ICDE 2020 |
Query processing and optimization
query workload |
0.3 | 1 | 2025 | Cardinality Estimation of LIKE Predicate Queries using Deep Learning · Proc. ACM Manag. Data 2025 |
Machine learning and data management › training data management
training data generation |
0.3 | 1 | 2025 | Cardinality Estimation of LIKE Predicate Queries using Deep Learning · Proc. ACM Manag. Data 2025 |
Information retrieval › similarity search
approximate string matching |
0.2 | 1 | 2022 | Cardinality Estimation of Approximate Substring Queries using Deep Learning · Proc. VLDB Endow. 2022 |
Methods — techniques the papers use, named apart from their topics
deep learning · 1.4maximum likelihood estimation · 1.0laplace mechanism · 1.0n-gram table · 0.9conditional regression · 0.9train data generation · 0.6partitioning · 0.4frequency vector representation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cardinality Estimation of LIKE Predicate Queries using Deep LearningabstractCardinality estimation of LIKE predicate queries has an important role in the query optimization of database systems. Traditional approaches generally use a summary of text data with some statistical assumptions. Recently, the deep learning model for cardinality estimation of LIKE predicate queries has been investigated. To provide more accurate cardinality estimates and reduce the maximum estimation errors, we propose a deep learning model that utilizes the extended N -gram table and the conditional regression header. We next investigate how to efficiently generate training data. Our LEADER (LikE predicate trAining Data gEneRation) algorithms utilize the shareable results across the relational queries corresponding to the LIKE predicates. By analyzing the queries corresponding to LIKE predicates, we develop an efficient join method and utilize the join order for fast query execution and maximal sharing of shareable results . Extensive experiments with real-life datasets confirm the efficiency of the proposed training data generation algorithms and the effectiveness of the proposed model. Suyong Kwon, Kyuseok Shim, Woohwan Jung |
Proc. ACM Manag. Data | 1 |
| 2022 | Cardinality Estimation of Approximate Substring Queries using Deep LearningabstractCardinality estimation of an approximate substring query is an important problem in database systems. Traditional approaches build a summary from the text data and estimate the cardinality using the summary with some statistical assumptions. Since deep learning models can learn underlying complex data patterns effectively, they have been successfully applied and shown to outperform traditional methods for cardinality estimations of queries in database systems. However, since they are not yet applied to approximate substring queries, we investigate a deep learning approach for cardinality estimation of such queries. Although the accuracy of deep learning models tends to improve as the train data size increases, producing a large train data is computationally expensive for cardinality estimation of approximate substring queries. Thus, we develop efficient train data generation algorithms by avoiding unnecessary computations and sharing common computations. We also propose a deep learning model as well as a novel learning method to quickly obtain an accurate deep learning-based estimator. Extensive experiments confirm the superiority of our data generation algorithms and deep learning model with the novel learning method. Suyong Kwon, Woohwan Jung, Kyuseok Shim |
Proc. VLDB Endow. | 1 |
| 2021 | TIDY: Publishing a Time Interval Dataset With Differential PrivacyabstractLog data from mobile devices generally contain a series of events with temporal information including time intervals which consist of the start and finish times. However, the problem of releasing differentially private time interval datasets has not been tackled yet. A time interval dataset can be represented by a two dimensional (2D) histogram. Most of the methods to publish 2D histograms partition the data into rectangular spaces to reduce the aggregated noise error for range queries. However, the existing algorithms to publish 2D histograms suffer from the structural error when applied to time interval datasets. To reduce the aggregated noise errors and suppress the increase in the structural error, we propose the TIDY (publishing Time Intervals via Differential privacY) algorithm. We use the frequency vectors as a compact representation of the time interval dataset. After applying the Laplace mechanism to the frequency vectors, we improve the utility of the frequency vectors based on a maximum likelihood estimation. We also develop a new partitioning method adapted for the frequency vectors to balance the trade-off between the noise and structural errors. Our empirical study on real-life and synthetic datasets confirms that TIDY outperforms the existing algorithms for 2D histograms. Woohwan Jung, Suyong Kwon, Kyuseok Shim |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | TIDY: Publishing a Time Interval Dataset with Differential Privacy (Extended abstract)abstractLog data from mobile devices usually contain a series of events with time intervals. However, the problem of releasing differentially private time interval data has not been tackled yet. We propose the TIDY (publishing Time Intervals via Differential privacY) algorithm to release time interval data under differential privacy. We use the frequency vectors as a compact representation of the time interval data to reduce the aggregated noise. We also develop a new partitioning method adapted for the frequency vectors to balance the trade-off between the noise and structural errors. Our experiments confirm that TIDY outperforms the existing algorithms for releasing 2D histograms. Woohwan Jung, Suyong Kwon, Kyuseok Shim |
ICDE | 2 |