EDBT 2026 Demo / reviewers in the wild / expert
Liang Jeff Chen
dblp:17/7993
· DBLP profile ↗
7ranked-venue papers
4as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 4 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
7 papers |
Query processing and optimization · 40% Information retrieval · 14% Database system architecture and tuning · 13% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 100% |
Topics — the 20 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
query rewriting |
0.3 | 2 | 2014 | Mapping XML to a Wide Sparse Table · IEEE Trans. Knowl. Data Eng. 2014 Mapping XML to a Wide Sparse Table · ICDE 2012 |
Data integration and cleaning › schema mapping
XML-to-relational mapping |
0.3 | 2 | 2014 | Mapping XML to a Wide Sparse Table · IEEE Trans. Knowl. Data Eng. 2014 Mapping XML to a Wide Sparse Table · ICDE 2012 |
Database system architecture and tuning › database design
physical database design |
0.3 | 1 | 2017 | Wide Table Layout Optimization based on Column Ordering and Duplication · SIGMOD Conference 2017 |
Storage systems › data management › database storage
columnar storage |
0.3 | 1 | 2017 | Wide Table Layout Optimization based on Column Ordering and Duplication · SIGMOD Conference 2017 |
Storage systems
storage engine |
0.3 | 1 | 2017 | Wide Table Layout Optimization based on Column Ordering and Duplication · SIGMOD Conference 2017 |
Graph data management › graph algorithms
graph traversal |
0.2 | 1 | 2016 | G-SQL: Fast Query Processing via Graph Exploration · Proc. VLDB Endow. 2016 |
Query processing and optimization › join processing
multi-way join |
0.2 | 1 | 2016 | G-SQL: Fast Query Processing via Graph Exploration · Proc. VLDB Endow. 2016 |
Query processing and optimization
approximate query processing |
0.2 | 1 | 2014 | Error-bounded Sampling for Analytics on Big Sparse Data · Proc. VLDB Endow. 2014 |
Query processing and optimization › approximate query processing
stratified sampling |
0.2 | 1 | 2014 | Error-bounded Sampling for Analytics on Big Sparse Data · Proc. VLDB Endow. 2014 |
Data models and query languages
XML |
0.2 | 1 | 2014 | Mapping XML to a Wide Sparse Table · IEEE Trans. Knowl. Data Eng. 2014 |
Data models and query languages
XML data management |
0.1 | 1 | 2012 | Mapping XML to a Wide Sparse Table · ICDE 2012 |
Information retrieval › ranking
context-aware ranking |
0.1 | 1 | 2011 | Context-sensitive ranking for document retrieval · SIGMOD Conference 2011 |
Information retrieval
ranking |
0.1 | 1 | 2011 | Context-sensitive ranking for document retrieval · SIGMOD Conference 2011 |
Information retrieval
keyword search |
0.1 | 1 | 2010 | Supporting top-K keyword search in XML databases · ICDE 2010 |
Query processing and optimization › top-k query processing
top-k keyword search |
0.1 | 1 | 2010 | Supporting top-K keyword search in XML databases · ICDE 2010 |
Query processing and optimization
top-k query processing |
0.1 | 1 | 2010 | Supporting top-K keyword search in XML databases · ICDE 2010 |
Information retrieval › keyword search
XML keyword search |
0.1 | 1 | 2010 | Supporting top-K keyword search in XML databases · ICDE 2010 |
Query processing and optimization
aggregate query processing |
0.1 | 1 | 2014 | Error-bounded Sampling for Analytics on Big Sparse Data · Proc. VLDB Endow. 2014 |
Data mining
big data analytics |
0.1 | 1 | 2014 | Error-bounded Sampling for Analytics on Big Sparse Data · Proc. VLDB Endow. 2014 |
Query processing and optimization
materialized view |
0.0 | 1 | 2011 | Context-sensitive ranking for document retrieval · SIGMOD Conference 2011 |
Methods — techniques the papers use, named apart from their topics
cost model · 0.8in-memory graph engine · 0.2stratified sampling · 0.2join reduction · 0.2error bound analysis · 0.2XPath-to-SQL translation · 0.2sparse table mapping · 0.1materialized views · 0.1top-k join · 0.1lowest common ancestor semantics · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Wide Table Layout Optimization based on Column Ordering and DuplicationabstractModern data analytical tasks often witness very wide tables, from a few hundred columns to a few thousand. While it is commonly agreed that column stores are an appropriate data format for wide tables and analytical workloads, the physical order of columns has not been investigated. Column ordering plays a critical role in I/O performance, because in wide tables accessing the columns in a single horizontal partition may involve multiple disk seeks. An optimal column ordering will incur minimal cumulative disk seek costs for the set of queries applied to the data. In this paper, we aim to find such an optimal column layout to maximize I/O performance. Specifically, we study two problems for column stores on HDFS: column ordering and column duplication. Column ordering seeks an approximately optimal order of columns; column duplication complements column ordering in that some columns may be duplicated multiple times to reduce contention among the queries' diverse requirements on the column order. We consider an actual fine-grained cost model for column accesses and propose algorithms that take a query workload as input and output a column ordering strategy with or without storage redundancy that significantly improves the overall I/O performance. Experimental results over real-life data and production query workloads confirm the effectiveness of the proposed algorithms in diverse settings. Haoqiong Bian, Ying Yan 0006, Wenbo Tao, Liang Jeff Chen, Yueguo Chen, Xiaoyong Du 0001, Thomas Moscibroda |
SIGMOD Conference | 4 |
| 2016 | G-SQL: Fast Query Processing via Graph ExplorationabstractA lot of real-life data are of graph nature. However, it is not until recently that business begins to exploit data's connectedness for business insights. On the other hand, RDBMSs are a mature technology for data management, but they are not for graph processing. Take graph traversal, a common graph operation for example, it heavily relies on a graph primitive that accesses a given node's neighborhood. We need to join tables following foreign keys to access the nodes in the neighborhood if an RDBMS is used to manage graph data. Graph exploration is a fundamental building block of many graph algorithms. But this simple operation is costly due to a large volume of I/O caused by the massive amount of table joins. In this paper, we present G-SQL, our effort toward the integration of a RDBMS and a native in-memory graph processing engine. G-SQL leverages the fast graph exploration capability provided by the graph engine to answer multi-way join queries. Meanwhile, it uses RDBMSs to provide mature data management functionalities, such as reliable data storage and additional data access methods. Specifically, G-SQL is a SQL dialect augmented with graph exploration functionalities and it dispatches query tasks to the in-memory graph engine and its underlying RDMBS. The G-SQL runtime coordinates the two query processors via a unified cost model to ensure the entire query is processed efficiently. Experimental results show that our approach greatly expands capabilities of RDBMs and delivers exceptional performance for SQL-graph hybrid queries. Hongbin Ma, Bin Shao 0002, Yanghua Xiao, Liang Jeff Chen, Haixun Wang |
Proc. VLDB Endow. | 4 |
| 2014 | Error-bounded Sampling for Analytics on Big Sparse DataabstractAggregation queries are at the core of business intelligence and data analytics. In the big data era, many scalable shared-nothing systems have been developed to process aggregation queries over massive amount of data. Microsoft's SCOPE is a well-known instance in this category. Nevertheless, aggregation queries are still expensive, because query processing needs to consume the entire data set, which is often hundreds of terabytes. Data sampling is a technique that samples a small portion of data to process and returns an approximate result with an error bound, thereby reducing the query's execution time. While similar problems were studied in the database literature, we encountered new challenges that disable most of prior efforts: (1) error bounds are dictated by end users and cannot be compromised, (2) data is sparse , meaning data has a limited population but a wide range. For such cases, conventional uniform sampling often yield high sampling rates and thus deliver limited or no performance gains. In this paper, we propose error-bounded stratified sampling to reduce sample size. The technique relies on the insight that we may only reduce the sampling rate with the knowledge of data distributions. The technique has been implemented into Microsoft internal search query platform. Results show that the proposed approach can reduce up to 99% sample size comparing with uniform sampling, and its performance is robust against data volume and other key performance metrics. Ying Yan 0006, Liang Jeff Chen, Zheng Zhang 0001 |
Proc. VLDB Endow. | 2 |
| 2014 | Mapping XML to a Wide Sparse TableabstractXML is commonly supported by SQL database systems. However, existing mappings of XML to tables can only deliver satisfactory query performance for limited use cases. In this paper, we propose a novel mapping of XML data into one wide table whose columns are sparsely populated. This mapping provides good performance for document types and queries that are observed in enterprise applications but are not supported efficiently by existing work. XML queries are evaluated by translating them into SQL queries over the wide sparsely-populated table. We show how to translate full XPath 1.0 into SQL. Based on the characteristics of the new mapping, we present rewriting optimizations that dramatically reduce the number of joins. Experiments demonstrate that query evaluation over the new mapping delivers considerable improvements over existing techniques for the target use cases. Liang Jeff Chen, Philip A. Bernstein, Peter Carlin, Dimitrije Filipovic, Michael Rys, Nikita Shamgunov, James F. Terwilliger, Milos Todic, Sasa Tomasevic, Dragan Tomic |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Mapping XML to a Wide Sparse TableabstractXML is commonly supported by SQL database systems. However, existing mappings of XML to tables can only deliver satisfactory query performance for limited use cases. In this paper, we propose a novel mapping of XML data into one wide table whose columns are sparsely populated. This mapping provides good performance for document types and queries that are observed in enterprise applications but are not supported efficiently by existing work. XML queries are evaluated by translating them into SQL queries over the wide sparsely-populated table. We show how to translate full XPath 1.0 into SQL. Based on the characteristics of the new mapping, we present rewriting optimizations that minimize the number of joins. Experiments demonstrate that query evaluation over the new mapping delivers considerable improvements over existing techniques for the target use cases. Liang Jeff Chen, Philip A. Bernstein, Peter Carlin, Dimitrije Filipovic, Michael Rys, Nikita Shamgunov, James F. Terwilliger, Milos Todic, Sasa Tomasevic, Dragan Tomic |
ICDE | 1 |
| 2011 | Context-sensitive ranking for document retrievalabstractWe study the problem of context-sensitive ranking for document retrieval, where a context is defined as a sub-collection of documents, and is specified by queries provided by domain-interested users. The motivation of context-sensitive search is that the ranking of the same keyword query generally depends on the context. The reason is that the underlying keyword statistics differ significantly from one context to another. The query evaluation challenge is the computation of keyword statistics at runtime, which involves expensive online aggregations. We appropriately leverage and extend materialized view research in order to deliver algorithms and data structures that evaluate context-sensitive queries efficiently. Specifically, a number of views are selected and materialized, each corresponding to one or more large contexts. Materialized views are used at query time to compute statistics which are used to compute ranking scores. Experimental results show that the contextsensitive ranking generally improves the ranking quality, while our materialized view-based technique improves the query efficiency. Liang Jeff Chen, Yannis Papakonstantinou |
SIGMOD Conference | 1 |
| 2010 | Supporting top-K keyword search in XML databasesabstractKeyword search is considered to be an effective information discovery method for both structured and semi-structured data. In XML keyword search, query semantics is based on the concept of Lowest Common Ancestor (LCA). However, naive LCA-based semantics leads to exponential computation and result size. In the literature, LCA-based semantic variants (e.g., ELCA and SLCA) were proposed, which define a subset of all the LCAs as the results. While most existing work focuses on algorithmic efficiency, top-K processing for XML keyword search is an important issue that has received very little attention. Existing algorithms focusing on efficiency are designed to optimize the semantic pruning and are incapable of supporting top-K processing. On the other hand, straightforward applications of top-K techniques from other areas (e.g., relational databases) generate LCAs that may not be the results and unnecessarily expand efforts in the semantic pruning. In this paper, we propose a series of join-based algorithms that combine the semantic pruning and the top-K processing to support top-K keyword search in XML databases. The algorithms essentially reduce the keyword query evaluation to relational joins, and incorporate the idea of the top-K join from relational databases. Extensive experimental evaluations show the performance advantages of our algorithms. Liang Jeff Chen, Yannis Papakonstantinou |
ICDE | 1 |