EDBT 2026 Demo / reviewers in the wild / expert
Byung-Chul Tak
dblp:10/6711 · also Byungchul Tak
· DBLP profile ↗
4ranked-venue papers in the field
0as first author
4since 2021 · last 2026
0000-0002-8204-6816ORCID · corroborated
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Batcher: Learning to Construct Cost-Efficient Batches of Small Queries in Big Data Processing Platforms
Yeonsu Park 0001, Taesung Lee, Byung-Chul Tak, Wook-Shin Han |
ICDE | 3 |
| 2026 | TurboLynx: Schemaless Graph Engine Strikes Back for General-Purpose Analytics
Taesung Lee, Jaehyun Ha, Byung-Chul Tak, Wook-Shin Han |
Proc. VLDB Endow. | 3 |
| 2024 | ReCG: Bottom-Up JSON Schema Discovery Using a Repetitive Cluster-and-Generalize FrameworkabstractThe schemalessness, one of the major advantages of JSON representation format, comes with high penalties in querying and operations by denying various critical functions such as query optimizations, indexing, or data verification. There have been continuous efforts to develop an accurate JSON schema discovery algorithm from a bag of JSON documents. Unfortunately, existing schema discovery techniques, being top-down algorithms, face challenges from the lack of visibility into children nodes of JSON tree. With absence of the information about lower-level JSON elements, top-down algorithms need to employ assumptions and heuristics to decide the schema type of nodes. However, such static decisions are often violated in datasets which causes top-down algorithms to perform poorly. To overcome this, we propose an algorithm, called ReCG, that processes JSON documents in a bottom-up manner. It builds up schemas from leaf elements upward in the JSON document tree and, thus, can make more informed decisions of the schema node types. In addition, we adopt MDL (Minimum Description Length) principles systematically while building up the schemas to choose among candidate schemas the most concise yet accurate one with well-balanced generality. Evaluations show that our technique improves the recall and precision of found schemas by as high as 47%, resulting in 46% better F1 score while also performing 2.11× faster on average against the state-of-the-art. Joohyung Yun, Byung-Chul Tak, Wook-Shin Han |
Proc. VLDB Endow. | 2 |
| 2023 | QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in SparkabstractSpark big data processing platform is heavily used in today's IT services for various critical applications such as machine learning tasks for service recommendations or massive volumes of raw sales data analysis. Spark is designed to deliver high performance by enabling a high degree of parallelism while processing various heavy-weight queries that require homogeneous operations on large data. However, it has been observed that workloads made of small and short-running queries coming from various sources are becoming dominant in practice. Unfortunately, the current Spark architecture is unfit to process workloads made of a large number of small queries optimally due to excessive I/Os with small computations. We present a technique, called QaaD, that addresses this problem fundamentally by applying i) transparent conversion of workloads made of small queries into one with large queries and ii) dynamic partition size adjustment for runtime overhead minimization. For this, we introduce a new abstraction, microRDD, to support our design of query merging, the embedding of queries as part of data, and an opportunistic sharing of common input data among queries. Comprehensive evaluation using real-world data shows that QaaD is able to deliver 10.6x to 36.6x speed-up against standard Spark executions for small query workloads. Yeonsu Park 0001, Byung-Chul Tak, Wook-Shin Han |
Proc. ACM Manag. Data | 2 |