Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wooseong Kwak

dblp:42/3511 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 53% Data integration and cleaning · 23% Spatial and temporal data management · 23%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Spatial and temporal data management › spatial query processing
spatial join
0.112010
On supporting effective web extraction · ICDE 2010
Data integration and cleaning › data extraction
web data extraction
0.112010
On supporting effective web extraction · ICDE 2010
Query processing and optimization › query optimization
join enumeration
0.112008
Parallelizing query optimization · Proc. VLDB Endow. 2008
Query processing and optimization › query optimization
parallel query optimization
0.112008
Parallelizing query optimization · Proc. VLDB Endow. 2008
Query processing and optimization
query optimization
0.112008
Parallelizing query optimization · Proc. VLDB Endow. 2008
Parallel and multicore computing › parallel algorithms › dynamic programming
parallel dynamic programming
0.112008
Parallelizing query optimization · Proc. VLDB Endow. 2008
Parallel and multicore computing
parallel programming models
0.112008
Parallelizing query optimization · Proc. VLDB Endow. 2008

Methods — techniques the papers use, named apart from their topics

skip vector array · 0.2dynamic programming · 0.2spatial relationship analysis · 0.1XPath · 0.1
YearPublicationVenuePosition
2014 Leveraging spatial join for robust tuple extraction from web pages
Wook-Shin Han, Wooseong Kwak, Hwanjo Yu, Jeonghoon Lee 0004, Min-Soo Kim 0002
Inf. Sci.2
2010 On supporting effective web extraction
abstract
Commercial tuple extraction systems have enjoyed some success to extract tuples by regarding HTML pages as tree structures and exploiting XPath queries to find attributes of tuples in the HTML pages. However, such systems would be vulnerable to small changes on the web pages. In this paper, we propose a robust tuple extraction system which utilizes spatial relationships among elements rather than the XPath queries of the elements. Our system regards elements in the rendered page as spatial objects in the 2-D space and executes spatial joins to extract target elements. Since humans also identify an element in a web page by its relative spatial location, our system extracting elements by their spatial relationships could possibly be as robust as manual extraction and is far more robust than existing tuple extraction systems.
Wook-Shin Han, Wooseong Kwak, Hwanjo Yu
ICDE2
2008 Parallelizing query optimization
abstract
Many commercial RDBMSs employ cost-based query optimization exploiting dynamic programming (DP) to efficiently generate the optimal query execution plan. However, optimization time increases rapidly for queries joining more than 10 tables. Randomized or heuristic search algorithms reduce query optimization time for large join queries by considering fewer plans, sacrificing plan optimality. Though commercial systems executing query plans in parallel have existed for over a decade, the optimization of such plans still occurs serially. While modern microprocessors employ multiple cores to accelerate computations, parallelizing query optimization to exploit multi-core parallelism is not as straightforward as it may seem. The DP used in join enumeration belongs to the challenging nonserial polyadic DP class because of its non-uniform data dependencies. In this paper, we propose a comprehensive and practical solution for parallelizing query optimization in the multi-core processor architecture, including a parallel join enumeration algorithm and several alternative ways to allocate work to threads to balance their load. We also introduce a novel data structure called skip vector array to significantly reduce the generation of join partitions that are infeasible. This solution has been prototyped in PostgreSQL. Extensive experiments using various query graph topologies confirm that our algorithms allocate the work evenly, thereby achieving almost linear speed-up. Our parallel join enumeration algorithm enhanced with our skip vector array outperforms the conventional generate-and-filter DP algorithm by up to two orders of magnitude for star queries-linear speedup due to parallelism and an order of magnitude performance improvement due to the skip vector array.
Wook-Shin Han, Wooseong Kwak, Jinsoo Lee, Guy M. Lohman, Volker Markl
Proc. VLDB Endow.2