VLDB 2026 Research / reviewers in the wild / expert
Ziqiao Ao
dblp:268/2514
· DBLP profile ↗
6ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0003-1266-5790ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
3 papers |
Information theory · 50% Mathematical optimization · 46% Algorithms and data structures · 4% | |
| Artificial intelligence
3 papers |
Probabilistic and Bayesian machine learning · 69% Generative modeling · 31% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information theory › estimation theory
entropy estimation |
1.2 | 2 | 2023 | Entropy estimation via uniformization · Artif. Intell. 2023 Entropy Estimation via Normalizing Flow · AAAI 2022 |
Information theory › estimation theory › entropy estimation
high-dimensional entropy estimation |
1.2 | 2 | 2023 | Entropy estimation via uniformization · Artif. Intell. 2023 Entropy Estimation via Normalizing Flow · AAAI 2022 |
Machine learning › Generative modeling
normalizing flow |
0.8 | 2 | 2023 | Entropy Estimation via Normalizing Flow · AAAI 2022 Entropy estimation via uniformization · Artif. Intell. 2023 |
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design |
0.8 | 1 | 2024 | On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design · AAAI 2024 |
Machine learning › Probabilistic and Bayesian machine learning › experimental design › bayesian experimental design
expected information gain |
0.8 | 1 | 2024 | On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design · AAAI 2024 |
Mathematical optimization
gradient estimation |
0.8 | 1 | 2024 | On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design · AAAI 2024 |
Mathematical optimization › stochastic optimization › stochastic gradient methods
stochastic gradient descent |
0.8 | 1 | 2024 | On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design · AAAI 2024 |
Mathematical optimization
stochastic optimization |
0.8 | 1 | 2024 | On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design · AAAI 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation |
0.2 | 1 | 2023 | Entropy estimation via uniformization · Artif. Intell. 2023 |
Algorithms and data structures › similarity search
nearest neighbor search |
0.2 | 1 | 2022 | Entropy Estimation via Normalizing Flow · AAAI 2022 |
Methods — techniques the papers use, named apart from their topics
normalizing flow · 2.5stochastic gradient descent · 1.5markov chain monte carlo · 1.5k-nearest neighbors entropy estimation · 1.3k-NN entropy estimation · 1.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Predictive Method for Estimating the Limits of Lossless Data CompressionabstractIn this paper, we address the limitations of current measures for estimating the lossless compression limits. Shannon entropy, while practical, assumes a known data distribution and does not account for the complexity of representing this distribution. Kolmogorov Complexity (KC) [2], on the other hand, offers a more complete measure by considering both data and model complexity, but it is uncomputable in practice. To bridge these gaps, we propose a novel framework that estimates lower and upper bounds for lossless compression limits, leveraging neural scaling laws [1] to balance model and data complexity. Our experiments demonstrate the accuracy of our approach on synthetic datasets, with an average estimation error of 1.18%, and highlight its effectiveness as a tool for evaluating lossless compression methods on real-world datasets. Ziqiao Ao, Zhaoyi Sun, Jie Sun 0007 |
DCC | 1 |
| 2024 | On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental DesignabstractBayesian Experimental Design (BED), which aims to find the optimal experimental conditions for Bayesian inference, is usually posed as to optimize the expected information gain (EIG). The gradient information is often needed for efficient EIG optimization, and as a result the ability to estimate the gradient of EIG is essential for BED problems. The primary goal of this work is to develop methods for estimating the gradient of EIG, which, combined with the stochastic gradient descent algorithms, result in efficient optimization of EIG. Specifically, we first introduce a posterior expected representation of the EIG gradient with respect to the design variables. Based on this, we propose two methods for estimating the EIG gradient, UEEG-MCMC that leverages posterior samples generated through Markov Chain Monte Carlo (MCMC) to estimate the EIG gradient, and BEEG-AP that focuses on achieving high simulation efficiency by repeatedly using parameter samples. Theoretical analysis and numerical studies illustrate that UEEG-MCMC is robust agains the actual EIG value, while BEEG-AP is more efficient when the EIG value to be optimized is small. Moreover, both methods show superior performance compared to several popular benchmarks in our numerical experiments. Ziqiao Ao, Jinglai Li |
AAAI | 1 |
| 2023 | Entropy estimation via uniformizationabstractEntropy estimation is of practical importance in information theory and statistical science. Many existing entropy estimators suffer from fast growing estimation bias with respect to dimensionality, rendering them unsuitable for high-dimensional problems. In this work we propose a transform-based method for high-dimensional entropy estimation, which consists of the following two main ingredients. Firstly, we provide a modified k-nearest neighbors (k-NN) entropy estimator that can reduce estimation bias for samples closely resembling a uniform distribution. Second we design a normalizing flow based mapping that pushes samples toward the uniform distribution, and the relation between the entropy of the original samples and the transformed ones is also derived. As a result the entropy of a given set of samples is estimated by first transforming them toward the uniform distribution and then applying the proposed estimator to the transformed samples. The performance of the proposed method is compared against several existing entropy estimators, with both mathematical examples and real-world applications. Ziqiao Ao, Jinglai Li |
Artif. Intell. | 1 |
| 2023 | Skill requirements in job advertisements: A comparison of skill-categorization methods based on wage regressions
Ziqiao Ao, Gergely Horváth, Chunyuan Sheng |
Inf. Process. Manag. | 1 |
| 2022 | Entropy Estimation via Normalizing FlowabstractEntropy estimation is an important problem in information theory and statistical science. Many popular entropy estimators suffer from fast growing estimation bias with respect to dimensionality, rendering them unsuitable for high dimensional problems. In this work we propose a transformbased method for high dimensional entropy estimation, which consists of the following two main ingredients. First by modifying the k-NN based entropy estimator, we propose a new estimator which enjoys small estimation bias for samples that are close to a uniform distribution. Second we design a normalizing flow based mapping that pushes samples toward a uniform distribution, and the relation between the entropy of the original samples and the transformed ones is also derived. As a result the entropy of a given set of samples is estimated by first transforming them toward a uniform distribution and then applying the proposed estimator to the transformed samples. Numerical experiments demonstrate the effectiveness of the method for high dimensional entropy estimation problems. Ziqiao Ao, Jinglai Li |
AAAI | 1 |
| 2020 | An approximate KLD based experimental design for models with intractable likelihoodsabstractData collection is a critical step in statistical inference and data science,and the goal of statistical experimental design (ED) is to find the data collection setupthat can provide most information for the inference. In this work we consider a special type of ED problems where the likelihoods are not available in a closed form. In this case, the popular information-theoretic Kullback-Leibler divergence (KLD) based design criterioncan not be used directly, as it requires to evaluate the likelihood function. To address the issue, we derive a new utility function,which is a lower bound of the original KLD utility. This lower bound is expressed in terms of the summation of two or more entropies in the data space, and thus can be evaluated efficiently via entropy estimation methods.We provide several numerical examples to demonstrate the performance of the proposed method. Ziqiao Ao, Jinglai Li |
AISTATS | 1 |