VLDB 2026 Research / reviewers in the wild / expert
Jiexing Li
dblp:79/796
· DBLP profile ↗
15ranked-venue papers
7as first author
2since 2021 · last 2025
0000-0001-7449-526XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
8 papers |
Query processing and optimization · 54% Database system architecture and tuning · 22% Distributed and cloud data management · 19% | |
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Performance modeling and evaluation · 50% Cloud and datacenter computing · 36% Parallel and multicore computing · 8% | |
| Network and information security
4 papers |
Privacy and data protection · 100% |
Topics — the 19 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
adaptive query processing |
0.8 | 1 | 2024 | Adaptive and Robust Query Execution for Lakehouses At Scale · Proc. VLDB Endow. 2024 |
Distributed and cloud data management › federated database
federated query processing |
0.3 | 1 | 2018 | F1 Query: Declarative Querying at Scale · Proc. VLDB Endow. 2018 |
Privacy and data protection
anonymization |
0.3 | 3 | 2010 | Correlation hiding by independence masking · ICDE 2010 Privacy Preserving Publishing on Multiple Quasi-identifiers · ICDE 2009 On Anti-Corruption Privacy Preserving Publication · ICDE 2008 |
Query processing and optimization › query execution
query progress estimation |
0.2 | 1 | 2016 | Operator and Query Progress Estimation in Microsoft SQL Server Live Query Statistics · SIGMOD Conference 2016 |
Distributed and cloud data management
data partitioning |
0.2 | 1 | 2014 | Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014 |
Query processing and optimization
parallel query processing |
0.2 | 1 | 2014 | Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014 |
Query processing and optimization
cost estimation |
0.1 | 1 | 2012 | Robust Estimation of Resource Consumption for SQL Queries using Statistical Techniques · Proc. VLDB Endow. 2012 |
Query processing and optimization › query execution
query progress indicator |
0.1 | 1 | 2012 | GSLPI: A Cost-Based Query Progress Indicator · ICDE 2012 |
Performance modeling and evaluation › performance prediction
query performance prediction |
0.1 | 1 | 2012 | GSLPI: A Cost-Based Query Progress Indicator · ICDE 2012 |
Information retrieval › keyword search
keyword search over relational databases |
0.1 | 1 | 2010 | Toward Scalable Keyword Search over Relational Data · Proc. VLDB Endow. 2010 |
Privacy and data protection
privacy-preserving data analysis |
0.1 | 1 | 2010 | Correlation hiding by independence masking · ICDE 2010 |
Privacy and data protection › anonymization
k-anonymity |
0.1 | 1 | 2009 | Privacy Preserving Publishing on Multiple Quasi-identifiers · ICDE 2009 |
Parallel and multicore computing › parallel computing
parallel database systems |
0.1 | 1 | 2017 | Resource bricolage and resource selection for parallel database systems · VLDB J. 2017 |
Privacy and data protection › data publishing
anonymized data publication |
0.1 | 1 | 2008 | Preservation of proximity privacy in publishing numerical sensitive data · SIGMOD Conference 2008 |
Privacy and data protection
data publishing |
0.1 | 1 | 2008 | On Anti-Corruption Privacy Preserving Publication · ICDE 2008 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.1 | 1 | 2014 | Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014 |
GPUs and heterogeneous computing
heterogeneous resources |
0.1 | 1 | 2014 | Resource Bricolage for Parallel Database Systems · Proc. VLDB Endow. 2014 |
Query processing and optimization › interactive data exploration
answer space exploration |
0.0 | 1 | 2010 | Toward Scalable Keyword Search over Relational Data · Proc. VLDB Endow. 2010 |
Information retrieval
search engines |
0.0 | 1 | 2010 | Toward Scalable Keyword Search over Relational Data · Proc. VLDB Endow. 2010 |
Methods — techniques the papers use, named apart from their topics
runtime statistics collection · 1.5declarative querying · 0.7SQL · 0.7resource selection · 0.6progress estimation · 0.5linear programming · 0.4statistical model · 0.3operator-level modeling · 0.3cost model · 0.3asymptotic behavior modeling · 0.3independence masking · 0.1data distortion · 0.1stratified sampling · 0.1perturbation · 0.1generalization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Periodic-noise-tolerant neurodynamic approach for kWTA operation applied to opinions evolution
Jiexing Li, Yongji Guan, Tiantai Deng, Long Jin 0001 |
Neural Networks | 1 |
| 2024 | Adaptive and Robust Query Execution for Lakehouses At ScaleabstractMany organizations have embraced the "Lakehouse" data management paradigm, which involves constructing structured data warehouses on top of open, unstructured data lakes. This approach stands in stark contrast to traditional, closed, relational databases and introduces challenges for performance and stability of distributed query processors. Firstly, in large-scale, open Lakehouses with uncurated data, high ingestion rates, external tables, or deeply nested schemas, it is often costly or wasteful to maintain perfect and up-to-date table and column statistics. Secondly, inherently imperfect cardinality estimates with conjunctive predicates, joins and user-defined functions can lead to bad query plans. Thirdly, for the sheer magnitude of data involved, strictly relying on static query plan decisions can result in performance and stability issues such as excessive data movement, substantial disk spillage, or high memory pressure. To address these challenges, this paper presents our design, implementation, evaluation and practice of the Adaptive Query Execution (AQE) framework, which exploits natural execution pipeline breakers in query plans to collect accurate statistics and re-optimize them at runtime for both performance and robustness. In the TPC-DS benchmark, the technique demonstrates up to 25× per query speedup. At Databricks, AQE has been successfully deployed in production for multiple years. It powers billions of queries and ETL jobs to process exabytes of data per day, through key enterprise products such as Databricks Runtime, Databricks SQL, and Delta Live Tables. Maryann Xue, Yingyi Bu, Abhishek Somani, Wenchen Fan, Steven Chen, Herman Van Hövell, Bart Samwel, Mostafa Mokhtar, Rk Korlapati, Andy Lam, Yunxiao Ma, Vuk Ercegovac, Jiexing Li, Alexander Behm, Yuanjian Li, Xiao Li 0087, Sriram Krishnamurthy, Amit Shukla 0001, Michalis Petropoulos, Sameer Paranjpye, Reynold Xin, Matei Zaharia |
Proc. VLDB Endow. | 14 |
| 2018 | F1 Query: Declarative Querying at ScaleabstractF1 Query is a stand-alone, federated query processing platform that executes SQL queries against data stored in different file-based formats as well as different storage systems at Google (e.g., Bigtable, Spanner, Google Spreadsheets, etc.). F1 Query eliminates the need to maintain the traditional distinction between different types of data processing workloads by simultaneously supporting: (i) OLTP-style point queries that affect only a few records; (ii) low-latency OLAP querying of large amounts of data; and (iii) large ETL pipelines. F1 Query has also significantly reduced the need for developing hard-coded data processing pipelines by enabling declarative queries integrated with custom business logic. F1 Query satisfies key requirements that are highly desirable within Google: (i) it provides a unified view over data that is fragmented and distributed over multiple data sources; (ii) it leverages datacenter resources for performant query processing with high throughput and low latency; (iii) it provides high scalability for large data sizes by increasing computational parallelism; and (iv) it is extensible and uses innovative approaches to integrate complex business logic in declarative query processing. This paper presents the end-to-end design of F1 Query. Evolved out of F1, the distributed database originally built to manage Google's advertising data, F1 Query has been in production for multiple years at Google and serves the querying needs of a large number of users and systems. Bart Samwel, John Cieslewicz, Ben Handy, Jason Govig, Petros Venetis, Chanjun Yang, Keith Peters, Jeff Shute, Daniel Tenedorio, Himani Apte, Felix Weigel, David Wilhite, Jiexing Li, Zhan Yuan, Craig Chasseur, Ian Rae, Anurag Biyani, Andrew Harn, Andrey Gubichev, Amr El-Helw, Orri Erling, Zhepeng Yan, Mohan Yang, Yiqun Wei, Thanh Do, Colin Zheng, Goetz Graefe, Somayeh Sardashti, Ahmed M. Aly, Divyakant Agrawal, Shivakumar Venkataraman |
Proc. VLDB Endow. | 15 |
| 2017 | Resource bricolage and resource selection for parallel database systems
Jiexing Li, Jeffrey F. Naughton, Rimma V. Nehme |
VLDB J. | 1 |
| 2016 | Operator and Query Progress Estimation in Microsoft SQL Server Live Query StatisticsabstractWe describe the design and implementation of the new Live Query Statistics (LQS) feature in Microsoft SQL Server 2016. The functionality includes the display of overall query progress as well as progress of individual operators in the query execution plan. We describe the overall functionality of LQS, give usage examples and detail all areas where we had to extend the current state-of-the-art to build the complete LQS feature. Finally, we evaluate the effect these extensions have on progress estimation accuracy with a series of experiments using a large set of synthetic and real workloads. Kukjin Lee, Arnd Christian König, Vivek R. Narasayya, Bolin Ding, Surajit Chaudhuri, Brent Ellwein, Alexey Eksarevskiy, Manbeen Kohli, Jacob Wyant, Praneeta Prakash, Rimma V. Nehme, Jiexing Li, Jeffrey F. Naughton |
SIGMOD Conference | 12 |
| 2014 | Resource Bricolage for Parallel Database SystemsabstractRunning parallel database systems in an environment with heterogeneous resources has become increasingly common, due to cluster evolution and increasing interest in moving applications into public clouds. For database systems running in a heterogeneous cluster, the default uniform data partitioning strategy may overload some of the slow machines while at the same time it may under-utilize the more powerful machines. Since the processing time of a parallel query is determined by the slowest machine, such an allocation strategy may result in a significant query performance degradation. We take a first step to address this problem by introducing a technique we call resource bricolage that improves database performance in heterogeneous environments. Our approach quantifies the performance differences among machines with various resources as they process workloads with diverse resource requirements. We formalize the problem of minimizing workload execution time and view it as an optimization problem, and then we employ linear programming to obtain a recommended data partitioning scheme. We verify the effectiveness of our technique with an extensive experimental study on a commercial database system. Jiexing Li, Jeffrey F. Naughton, Rimma V. Nehme |
Proc. VLDB Endow. | 1 |
| 2013 | Toward Progress Indicators on Steroids for Big Data Systems
Jiexing Li, Rimma V. Nehme, Jeffrey F. Naughton |
CIDR | 1 |
| 2012 | GSLPI: A Cost-Based Query Progress IndicatorabstractProgress indicators for SQL queries were first published in 2004 with the simultaneous and independent proposals from Chaudhuri et al. and Luo et al. In this paper, we implement both progress indicators in the same commercial RDBMS to investigate their performance. We summarize common cases in which they are both accurate and cases in which they fail to provide reliable estimates. Although there are differences in their performance, much more striking is the similarity in the errors they make due to a common simplifying uniform future speed assumption. While the developers of these progress indicators were aware that this assumption could cause errors, they neither explored how large the errors might be nor did they investigate the feasibility of removing the assumption. To rectify this we propose a new query progress indicator, similar to these early progress indicators but without the uniform speed assumption. Experiments show that on the TPC-H benchmark, on queries for which the original progress indicators have errors up to 30X the query running time, the new progress indicator is accurate to within 10 percent. We also discuss the sources of the errors that still remain and shed some light on what would need to be done to eliminate them. Jiexing Li, Rimma V. Nehme, Jeffrey F. Naughton |
ICDE | 1 |
| 2012 | Robust Estimation of Resource Consumption for SQL Queries using Statistical TechniquesabstractThe ability to estimate resource consumption of SQL queries is crucial for a number of tasks in a database system such as admission control, query scheduling and costing during query optimization. Recent work has explored the use of statistical techniques for resource estimation in place of the manually constructed cost models used in query optimization. Such techniques, which require as training data examples of resource usage in queries, offer the promise of superior estimation accuracy since they can account for factors such as hardware characteristics of the system or bias in cardinality estimates. However, the proposed approaches lack robustness in that they do not generalize well to queries that are different from the training examples, resulting in significant estimation errors. Our approach aims to address this problem by combining knowledge of database query processing with statistical models. We model resource-usage at the level of individual operators, with different models and features for each operator type, and explicitly model the asymptotic behavior of each operator. This results in significantly better estimation accuracy and the ability to estimate resource usage of arbitrary plans, even when they are very different from the training instances. We validate our approach using various large scale real-life and benchmark workloads on Microsoft SQL Server. Jiexing Li, Arnd Christian König, Vivek R. Narasayya, Surajit Chaudhuri |
Proc. VLDB Endow. | 1 |
| 2010 | Correlation hiding by independence maskingabstractExtracting useful correlation from a dataset has been extensively studied. In this paper, we deal with the opposite, namely, a problem we call correlation hiding (CH), which is fundamental in numerous applications that need to disseminate data containing sensitive information. In this problem, we are given a relational table T whose attributes can be classified into three disjoint sets A, B, and C. The objective is to distort some values in T so that A becomes independent from B, and yet, their correlation with C is preserved as much as possible. CH is different from all the problems studied previously in the area of data privacy, in that CH demands complete elimination of the correlation between two sets of attributes, whereas the previous research focuses on partial elimination up to a certain level. A new operator called independence masking is proposed to solve the CH problem. Implementations of the operator with good worst case guarantees are described in the full version of this short note. Yufei Tao 0001, Jian Pei 0001, Jiexing Li, Xiaokui Xiao, Ke Yi 0001, Zhengzheng Xing |
ICDE | 3 |
| 2010 | Toward Scalable Keyword Search over Relational DataabstractKeyword search (KWS) over relational databases has recently received significant attention. Many solutions and many prototypes have been developed. This task requires addressing many issues, including robustness, accuracy, reliability, and privacy. An emerging issue, however, appears to be performance related: current KWS systems have unpredictable running times. In particular, for certain queries it takes too long to produce answers, and for others the system may even fail to return (e.g., after exhausting memory). In this paper we argue that as today's users have been "spoiled" by the performance of Internet search engines, KWS systems should return whatever answers they can produce quickly and then provide users with options for exploring any portion of the answer space not covered by these answers. Our basic idea is to produce answers that can be generated quickly as in today's KWS systems, then to show users query forms that characterize the unexplored portion of the answer space. Combining KWS systems with forms allows us to bypass the performance problems inherent to KWS without compromising query coverage. We provide a proof of concept for this proposed approach, and discuss the challenges encountered in building this hybrid system. Finally, we present experiments over real-world datasets to demonstrate the feasibility of the proposed solution. Akanksha Baid, Ian Rae, Jiexing Li, AnHai Doan, Jeffrey F. Naughton |
Proc. VLDB Endow. | 3 |
| 2009 | Privacy Preserving Publishing on Multiple Quasi-identifiersabstractIn some applications of privacy preserving data publishing, a practical demand is to publish a data set on multiple quasi-identifiers for multiple users simultaneously, which poses several challenges. Can we generate one anonymized version of the data so that the privacy preservation requirement like k-anonymity is satisfied for all users and the information loss is reduced as much as possible? In this paper, we identify and tackle the novel problem by an elegant solution.The full paper is available at http://www.cs.sfu.ca/~jpei/publications/butterfly-tr.pdf. Jian Pei 0001, Yufei Tao 0001, Jiexing Li, Xiaokui Xiao |
ICDE | 3 |
| 2009 | A brief survey of computational approaches in Social ComputingabstractWeb 2.0 technologies have brought new ways of connecting people in social networks for collaboration in various on-line communities. Social Computing is a novel and emerging computing paradigm that involves a multi-disciplinary approach in analyzing and modeling social behaviors on different media and platforms to produce intelligent and interactive applications and results. In this paper, we give a brief survey of the various machine learning and computational techniques used in Social Computing by first examining the social platforms, e.g., social network sites, social media, social games, social bookmarking, and social knowledge sites, where computational methodology is required to collect, extract, process, mine, and visualize the data. We then present surveys on more specific instances of computation tasks and techniques, e.g., social network analysis, link modeling and mining, ranking, sentiment analysis, etc., that are being used on these social platforms to obtain desirable results. Lastly, we present a small subset of an extensive reference list, which contains over 140 highly relevant references relating to the recent development in the computational aspects of Social Computing. Irwin King, Jiexing Li, Kam Tong Chan |
IJCNN | 2 |
| 2008 | On Anti-Corruption Privacy Preserving PublicationabstractThis paper deals with a new type of privacy threat, called "corruption" in anonymized data publication. Specifically, an adversary is said to have corrupted some individuals, if s/he has already obtained their sensitive values before consulting the released information. Conventional generalization may lead to severe privacy disclosure in the presence of corruption. Motivated by this, we advocate an alternative anonymization technique that integrates generalization with perturbation and stratified sampling. The integration provides strong privacy guarantees, even if an adversary has corrupted any number of individuals. We verify the effectiveness of the proposed technique through experiments with real data. Yufei Tao 0001, Xiaokui Xiao, Jiexing Li |
ICDE | 3 |
| 2008 | Preservation of proximity privacy in publishing numerical sensitive dataabstractWe identify proximity breach as a privacy threat specific to numerical sensitive attributes in anonymized data publication. Such breach occurs when an adversary concludes with high confidence that the sensitive value of a victim individual must fall in a short interval --- even though the adversary may have low confidence about the victim's actual value. Jiexing Li, Yufei Tao 0001, Xiaokui Xiao |
SIGMOD Conference | 1 |