Fen-Ling Ling

dblp:83/1583 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
0since 2021 · last 2008
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 43% Distributed and cloud data management · 30% Information retrieval · 26%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed and cloud data management
query offloading
0.112008
Integration of Server, Storage and Database Stack: Moving Processing Towards Data · ICDE 2008
Query processing and optimization
adaptive query processing
0.112007
Lazy, adaptive rid-list intersection, and its application to index anding · SIGMOD Conference 2007
Information retrieval › query processing
list intersection
0.112007
Lazy, adaptive rid-list intersection, and its application to index anding · SIGMOD Conference 2007
Query processing and optimization › OLAP
star query
0.012008
Integration of Server, Storage and Database Stack: Moving Processing Towards Data · ICDE 2008
Storage systems › computational storage
in-storage computing
0.012008
Integration of Server, Storage and Database Stack: Moving Processing Towards Data · ICDE 2008
Query processing and optimization
cardinality estimation
0.012007
Lazy, adaptive rid-list intersection, and its application to index anding · SIGMOD Conference 2007

Methods — techniques the papers use, named apart from their topics

value proposition analysis · 0.2lazy evaluation · 0.1adaptive algorithms · 0.1
YearPublicationVenuePosition
2008 Integration of Server, Storage and Database Stack: Moving Processing Towards Data
abstract
Storage architecture includes more and more processing power for increasing requirement of reliability, managibility and scalability. For example, an IBM storage server is equipped with 4 or 8 state-of-the-art processors and gigabytes of memories. This trend enables analyzing data locally inside a storage server. Processing data locally is appealing under the following circumstances: (1) huge reduction of data flowing to the host, (2) reduction of CPU consumption on host. Accordingly, the benefits are (1) less data traffic through IO channel to the host, (2) better utilization of host bufferpool, and (3) enabling more workload on the host. One crucial task is to understand how DBMS can benefit from such hardware. That is to identify which database operations are beneficial to be offloaded given a query workload in a particular setting. For certain operations, we establish value proposition via various approaches and show the analytical and experimental results. In particular, starjoin queries are commonly used in business warehouses. We propose to offload a portion of a starjoin query from host to the POWER5 P processors on a storage server, which dramatically reduces the amount of channel IO and host CPU consumption. Moreover, the query elapsed time is improved via the exploitation of the state-of-the-art P processors on a storage server.
Lin Qiao 0001, Vijayshankar Raman, Inderpal Narang, Prashant Pandey 0005, David D. Chambliss, Gene Fuh, James A. Ruddy, Ying-Lin Chen, Kou-Horng Yang, Fen-Ling Ling
ICDE10
2007 Lazy, adaptive rid-list intersection, and its application to index anding
abstract
RID-List (row id list) intersection is a common strategy in query processing, used in star joins, column stores, and even search engines. To apply a conjunction of predicates on a table, a query process ordoes index lookups to form sorted RID-lists (or bitmap) of the rows matching each predicate, then intersects the RID-lists via an AND-tree, and finally fetches the corresponding rows to apply any residual predicates and aggregates. This process can be expensive when the RID-lists are large. Furthermore, the performance is sensitive to the order in which RID lists are intersected together, and to treating the right predicates as residuals. If the optimizer chooses a wrong order or a wrong residual, due to a poor cardinality estimate, the resulting plan can run orders of magnitude slower than expected. We present a new algorithm for RID-list intersection that is both more efficient and more robust than this standard algorithm. First, we avoid forming the RID-lists up front, and instead form this lazily as part of the intersection. This reduces the associated IO and sort cost significantly, especially when the data distribution is skewed. It also ameliorates the problem of wrong residual table selection. Second, we do not intersect the RID-lists via an AND-tree, because this is vulnerable to cardinality mis-estimations. Instead, we use an adaptive set intersection algorithm that performs well even when the cardinality estimates are wrong. We present detailed experiments of this algorithm on data with varying distributions to validate its efficiency and predictability.
Vijayshankar Raman, Lin Qiao 0001, Inderpal Narang, Ying-Lin Chen, Kou-Horng Yang, Fen-Ling Ling
SIGMOD Conference7