Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jiexing Li

dblp:79/796 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
2since 2021 · last 2025
0000-0001-7449-526XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Query processing and optimization · 54% Database system architecture and tuning · 22% Distributed and cloud data management · 19%
Computer architecture, parallel and distributed computing, and storage systems
7 papers
Performance modeling and evaluation · 50% Cloud and datacenter computing · 36% Parallel and multicore computing · 8%
Network and information security
4 papers
Privacy and data protection · 100%

Topics — the 19 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
adaptive query processing
0.812024
Adaptive and Robust Query Execution for Lakehouses At Scale · Proc. VLDB Endow. 2024
Distributed and cloud data management › federated database
federated query processing
0.312018
F1 Query: Declarative Querying at Scale · Proc. VLDB Endow. 2018
Privacy and data protection
anonymization
0.332010
Correlation hiding by independence masking · ICDE 2010
Privacy Preserving Publishing on Multiple Quasi-identifiers · ICDE 2009
On Anti-Corruption Privacy Preserving Publication · ICDE 2008
Query processing and optimization › query execution
query progress estimation
0.212016
Operator and Query Progress Estimation in Microsoft SQL Server Live Query Statistics · SIGMOD Conference 2016
Distributed and cloud data management
data partitioning
0.212014
Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014
Query processing and optimization
parallel query processing
0.212014
Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014
Query processing and optimization
cost estimation
0.112012
Robust Estimation of Resource Consumption for SQL Queries using Statistical Techniques · Proc. VLDB Endow. 2012
Query processing and optimization › query execution
query progress indicator
0.112012
GSLPI: A Cost-Based Query Progress Indicator · ICDE 2012
Performance modeling and evaluation › performance prediction
query performance prediction
0.112012
GSLPI: A Cost-Based Query Progress Indicator · ICDE 2012
Information retrieval › keyword search
keyword search over relational databases
0.112010
Toward Scalable Keyword Search over Relational Data · Proc. VLDB Endow. 2010
Privacy and data protection
privacy-preserving data analysis
0.112010
Correlation hiding by independence masking · ICDE 2010
Privacy and data protection › anonymization
k-anonymity
0.112009
Privacy Preserving Publishing on Multiple Quasi-identifiers · ICDE 2009
Parallel and multicore computing › parallel computing
parallel database systems
0.112017
Resource bricolage and resource selection for parallel database systems · VLDB J. 2017
Privacy and data protection › data publishing
anonymized data publication
0.112008
Preservation of proximity privacy in publishing numerical sensitive data · SIGMOD Conference 2008
Privacy and data protection
data publishing
0.112008
On Anti-Corruption Privacy Preserving Publication · ICDE 2008
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.112014
Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014
GPUs and heterogeneous computing
heterogeneous resources
0.112014
Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014
Query processing and optimization › interactive data exploration
answer space exploration
0.012010
Toward Scalable Keyword Search over Relational Data · Proc. VLDB Endow. 2010
Information retrieval
search engines
0.012010
Toward Scalable Keyword Search over Relational Data · Proc. VLDB Endow. 2010

Methods — techniques the papers use, named apart from their topics

runtime statistics collection · 1.5declarative querying · 0.7SQL · 0.7resource selection · 0.6progress estimation · 0.5linear programming · 0.4statistical model · 0.3operator-level modeling · 0.3cost model · 0.3asymptotic behavior modeling · 0.3independence masking · 0.1data distortion · 0.1stratified sampling · 0.1perturbation · 0.1generalization · 0.1
YearPublicationVenuePosition
2025 Periodic-noise-tolerant neurodynamic approach for kWTA operation applied to opinions evolution
Jiexing Li, Yongji Guan, Tiantai Deng, Long Jin 0001
Neural Networks1
2024 Adaptive and Robust Query Execution for Lakehouses At Scale
abstract
Many organizations have embraced the "Lakehouse" data management paradigm, which involves constructing structured data warehouses on top of open, unstructured data lakes. This approach stands in stark contrast to traditional, closed, relational databases and introduces challenges for performance and stability of distributed query processors. Firstly, in large-scale, open Lakehouses with uncurated data, high ingestion rates, external tables, or deeply nested schemas, it is often costly or wasteful to maintain perfect and up-to-date table and column statistics. Secondly, inherently imperfect cardinality estimates with conjunctive predicates, joins and user-defined functions can lead to bad query plans. Thirdly, for the sheer magnitude of data involved, strictly relying on static query plan decisions can result in performance and stability issues such as excessive data movement, substantial disk spillage, or high memory pressure. To address these challenges, this paper presents our design, implementation, evaluation and practice of the Adaptive Query Execution (AQE) framework, which exploits natural execution pipeline breakers in query plans to collect accurate statistics and re-optimize them at runtime for both performance and robustness. In the TPC-DS benchmark, the technique demonstrates up to 25× per query speedup. At Databricks, AQE has been successfully deployed in production for multiple years. It powers billions of queries and ETL jobs to process exabytes of data per day, through key enterprise products such as Databricks Runtime, Databricks SQL, and Delta Live Tables.
Maryann Xue, Yingyi Bu, Abhishek Somani, Wenchen Fan, Steven Chen, Herman Van Hövell, Bart Samwel, Mostafa Mokhtar, Rk Korlapati, Andy Lam, Yunxiao Ma, Vuk Ercegovac, Jiexing Li, Alexander Behm, Yuanjian Li, Xiao Li 0087, Sriram Krishnamurthy, Amit Shukla 0001, Michalis Petropoulos, Sameer Paranjpye, Reynold Xin, Matei Zaharia
Proc. VLDB Endow.14
2018 F1 Query: Declarative Querying at Scale
abstract
F1 Query is a stand-alone, federated query processing platform that executes SQL queries against data stored in different file-based formats as well as different storage systems at Google (e.g., Bigtable, Spanner, Google Spreadsheets, etc.). F1 Query eliminates the need to maintain the traditional distinction between different types of data processing workloads by simultaneously supporting: (i) OLTP-style point queries that affect only a few records; (ii) low-latency OLAP querying of large amounts of data; and (iii) large ETL pipelines. F1 Query has also significantly reduced the need for developing hard-coded data processing pipelines by enabling declarative queries integrated with custom business logic. F1 Query satisfies key requirements that are highly desirable within Google: (i) it provides a unified view over data that is fragmented and distributed over multiple data sources; (ii) it leverages datacenter resources for performant query processing with high throughput and low latency; (iii) it provides high scalability for large data sizes by increasing computational parallelism; and (iv) it is extensible and uses innovative approaches to integrate complex business logic in declarative query processing. This paper presents the end-to-end design of F1 Query. Evolved out of F1, the distributed database originally built to manage Google's advertising data, F1 Query has been in production for multiple years at Google and serves the querying needs of a large number of users and systems.
Bart Samwel, John Cieslewicz, Ben Handy, Jason Govig, Petros Venetis, Chanjun Yang, Keith Peters, Jeff Shute, Daniel Tenedorio, Himani Apte, Felix Weigel, David Wilhite, Jiexing Li, Zhan Yuan, Craig Chasseur, Ian Rae, Anurag Biyani, Andrew Harn, Andrey Gubichev, Amr El-Helw, Orri Erling, Zhepeng Yan, Mohan Yang, Yiqun Wei, Thanh Do, Colin Zheng, Goetz Graefe, Somayeh Sardashti, Ahmed M. Aly, Divyakant Agrawal, Shivakumar Venkataraman
Proc. VLDB Endow.15
2017 Resource bricolage and resource selection for parallel database systems
Jiexing Li, Jeffrey F. Naughton, Rimma V. Nehme
VLDB J.1
2016 Operator and Query Progress Estimation in Microsoft SQL Server Live Query Statistics
abstract
We describe the design and implementation of the new Live Query Statistics (LQS) feature in Microsoft SQL Server 2016. The functionality includes the display of overall query progress as well as progress of individual operators in the query execution plan. We describe the overall functionality of LQS, give usage examples and detail all areas where we had to extend the current state-of-the-art to build the complete LQS feature. Finally, we evaluate the effect these extensions have on progress estimation accuracy with a series of experiments using a large set of synthetic and real workloads.
Kukjin Lee, Arnd Christian König, Vivek R. Narasayya, Bolin Ding, Surajit Chaudhuri, Brent Ellwein, Alexey Eksarevskiy, Manbeen Kohli, Jacob Wyant, Praneeta Prakash, Rimma V. Nehme, Jiexing Li, Jeffrey F. Naughton
SIGMOD Conference12
2014 Resource Bricolage for Parallel Database Systems
abstract
Running parallel database systems in an environment with heterogeneous resources has become increasingly common, due to cluster evolution and increasing interest in moving applications into public clouds. For database systems running in a heterogeneous cluster, the default uniform data partitioning strategy may overload some of the slow machines while at the same time it may under-utilize the more powerful machines. Since the processing time of a parallel query is determined by the slowest machine, such an allocation strategy may result in a significant query performance degradation. We take a first step to address this problem by introducing a technique we call resource bricolage that improves database performance in heterogeneous environments. Our approach quantifies the performance differences among machines with various resources as they process workloads with diverse resource requirements. We formalize the problem of minimizing workload execution time and view it as an optimization problem, and then we employ linear programming to obtain a recommended data partitioning scheme. We verify the effectiveness of our technique with an extensive experimental study on a commercial database system.
Jiexing Li, Jeffrey F. Naughton, Rimma V. Nehme
Proc. VLDB Endow.1
2013 Toward Progress Indicators on Steroids for Big Data Systems
Jiexing Li, Rimma V. Nehme, Jeffrey F. Naughton
CIDR1
2012 GSLPI: A Cost-Based Query Progress Indicator
abstract
Progress indicators for SQL queries were first published in 2004 with the simultaneous and independent proposals from Chaudhuri et al. and Luo et al. In this paper, we implement both progress indicators in the same commercial RDBMS to investigate their performance. We summarize common cases in which they are both accurate and cases in which they fail to provide reliable estimates. Although there are differences in their performance, much more striking is the similarity in the errors they make due to a common simplifying uniform future speed assumption. While the developers of these progress indicators were aware that this assumption could cause errors, they neither explored how large the errors might be nor did they investigate the feasibility of removing the assumption. To rectify this we propose a new query progress indicator, similar to these early progress indicators but without the uniform speed assumption. Experiments show that on the TPC-H benchmark, on queries for which the original progress indicators have errors up to 30X the query running time, the new progress indicator is accurate to within 10 percent. We also discuss the sources of the errors that still remain and shed some light on what would need to be done to eliminate them.
Jiexing Li, Rimma V. Nehme, Jeffrey F. Naughton
ICDE1
2012 Robust Estimation of Resource Consumption for SQL Queries using Statistical Techniques
abstract
The ability to estimate resource consumption of SQL queries is crucial for a number of tasks in a database system such as admission control, query scheduling and costing during query optimization. Recent work has explored the use of statistical techniques for resource estimation in place of the manually constructed cost models used in query optimization. Such techniques, which require as training data examples of resource usage in queries, offer the promise of superior estimation accuracy since they can account for factors such as hardware characteristics of the system or bias in cardinality estimates. However, the proposed approaches lack robustness in that they do not generalize well to queries that are different from the training examples, resulting in significant estimation errors. Our approach aims to address this problem by combining knowledge of database query processing with statistical models. We model resource-usage at the level of individual operators, with different models and features for each operator type, and explicitly model the asymptotic behavior of each operator. This results in significantly better estimation accuracy and the ability to estimate resource usage of arbitrary plans, even when they are very different from the training instances. We validate our approach using various large scale real-life and benchmark workloads on Microsoft SQL Server.
Jiexing Li, Arnd Christian König, Vivek R. Narasayya, Surajit Chaudhuri
Proc. VLDB Endow.1
2010 Correlation hiding by independence masking
abstract
Extracting useful correlation from a dataset has been extensively studied. In this paper, we deal with the opposite, namely, a problem we call correlation hiding (CH), which is fundamental in numerous applications that need to disseminate data containing sensitive information. In this problem, we are given a relational table T whose attributes can be classified into three disjoint sets A, B, and C. The objective is to distort some values in T so that A becomes independent from B, and yet, their correlation with C is preserved as much as possible. CH is different from all the problems studied previously in the area of data privacy, in that CH demands complete elimination of the correlation between two sets of attributes, whereas the previous research focuses on partial elimination up to a certain level. A new operator called independence masking is proposed to solve the CH problem. Implementations of the operator with good worst case guarantees are described in the full version of this short note.
Yufei Tao 0001, Jian Pei 0001, Jiexing Li, Xiaokui Xiao, Ke Yi 0001, Zhengzheng Xing
ICDE3
2010 Toward Scalable Keyword Search over Relational Data
abstract
Keyword search (KWS) over relational databases has recently received significant attention. Many solutions and many prototypes have been developed. This task requires addressing many issues, including robustness, accuracy, reliability, and privacy. An emerging issue, however, appears to be performance related: current KWS systems have unpredictable running times. In particular, for certain queries it takes too long to produce answers, and for others the system may even fail to return (e.g., after exhausting memory). In this paper we argue that as today's users have been "spoiled" by the performance of Internet search engines, KWS systems should return whatever answers they can produce quickly and then provide users with options for exploring any portion of the answer space not covered by these answers. Our basic idea is to produce answers that can be generated quickly as in today's KWS systems, then to show users query forms that characterize the unexplored portion of the answer space. Combining KWS systems with forms allows us to bypass the performance problems inherent to KWS without compromising query coverage. We provide a proof of concept for this proposed approach, and discuss the challenges encountered in building this hybrid system. Finally, we present experiments over real-world datasets to demonstrate the feasibility of the proposed solution.
Akanksha Baid, Ian Rae, Jiexing Li, AnHai Doan, Jeffrey F. Naughton
Proc. VLDB Endow.3
2009 Privacy Preserving Publishing on Multiple Quasi-identifiers
abstract
In some applications of privacy preserving data publishing, a practical demand is to publish a data set on multiple quasi-identifiers for multiple users simultaneously, which poses several challenges. Can we generate one anonymized version of the data so that the privacy preservation requirement like k-anonymity is satisfied for all users and the information loss is reduced as much as possible? In this paper, we identify and tackle the novel problem by an elegant solution.The full paper is available at http://www.cs.sfu.ca/~jpei/publications/butterfly-tr.pdf.
Jian Pei 0001, Yufei Tao 0001, Jiexing Li, Xiaokui Xiao
ICDE3
2009 A brief survey of computational approaches in Social Computing
abstract
Web 2.0 technologies have brought new ways of connecting people in social networks for collaboration in various on-line communities. Social Computing is a novel and emerging computing paradigm that involves a multi-disciplinary approach in analyzing and modeling social behaviors on different media and platforms to produce intelligent and interactive applications and results. In this paper, we give a brief survey of the various machine learning and computational techniques used in Social Computing by first examining the social platforms, e.g., social network sites, social media, social games, social bookmarking, and social knowledge sites, where computational methodology is required to collect, extract, process, mine, and visualize the data. We then present surveys on more specific instances of computation tasks and techniques, e.g., social network analysis, link modeling and mining, ranking, sentiment analysis, etc., that are being used on these social platforms to obtain desirable results. Lastly, we present a small subset of an extensive reference list, which contains over 140 highly relevant references relating to the recent development in the computational aspects of Social Computing.
Irwin King, Jiexing Li, Kam Tong Chan
IJCNN2
2008 On Anti-Corruption Privacy Preserving Publication
abstract
This paper deals with a new type of privacy threat, called "corruption" in anonymized data publication. Specifically, an adversary is said to have corrupted some individuals, if s/he has already obtained their sensitive values before consulting the released information. Conventional generalization may lead to severe privacy disclosure in the presence of corruption. Motivated by this, we advocate an alternative anonymization technique that integrates generalization with perturbation and stratified sampling. The integration provides strong privacy guarantees, even if an adversary has corrupted any number of individuals. We verify the effectiveness of the proposed technique through experiments with real data.
Yufei Tao 0001, Xiaokui Xiao, Jiexing Li
ICDE3
2008 Preservation of proximity privacy in publishing numerical sensitive data
abstract
We identify proximity breach as a privacy threat specific to numerical sensitive attributes in anonymized data publication. Such breach occurs when an adversary concludes with high confidence that the sensitive value of a victim individual must fall in a short interval --- even though the adversary may have low confidence about the victim's actual value.
Jiexing Li, Yufei Tao 0001, Xiaokui Xiao
SIGMOD Conference1