Ziqiao Ao

dblp:268/2514 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0003-1266-5790ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
3 papers
Information theory · 50% Mathematical optimization · 46% Algorithms and data structures · 4%
Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 69% Generative modeling · 31%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information theory › estimation theory
entropy estimation
1.222023
Entropy estimation via uniformization · Artif. Intell. 2023
Entropy Estimation via Normalizing Flow · AAAI 2022
Information theory › estimation theory › entropy estimation
high-dimensional entropy estimation
1.222023
Entropy estimation via uniformization · Artif. Intell. 2023
Entropy Estimation via Normalizing Flow · AAAI 2022
Machine learning › Generative modeling
normalizing flow
0.822023
Entropy Estimation via Normalizing Flow · AAAI 2022
Entropy estimation via uniformization · Artif. Intell. 2023
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design
0.812024
On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design · AAAI 2024
Machine learning › Probabilistic and Bayesian machine learning › experimental design › bayesian experimental design
expected information gain
0.812024
On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design · AAAI 2024
Mathematical optimization
gradient estimation
0.812024
On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design · AAAI 2024
Mathematical optimization › stochastic optimization › stochastic gradient methods
stochastic gradient descent
0.812024
On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design · AAAI 2024
Mathematical optimization
stochastic optimization
0.812024
On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design · AAAI 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation
0.212023
Entropy estimation via uniformization · Artif. Intell. 2023
Algorithms and data structures › similarity search
nearest neighbor search
0.212022
Entropy Estimation via Normalizing Flow · AAAI 2022

Methods — techniques the papers use, named apart from their topics

normalizing flow · 2.5stochastic gradient descent · 1.5markov chain monte carlo · 1.5k-nearest neighbors entropy estimation · 1.3k-NN entropy estimation · 1.1
YearPublicationVenuePosition
2025 A Predictive Method for Estimating the Limits of Lossless Data Compression
abstract
In this paper, we address the limitations of current measures for estimating the lossless compression limits. Shannon entropy, while practical, assumes a known data distribution and does not account for the complexity of representing this distribution. Kolmogorov Complexity (KC) [2], on the other hand, offers a more complete measure by considering both data and model complexity, but it is uncomputable in practice. To bridge these gaps, we propose a novel framework that estimates lower and upper bounds for lossless compression limits, leveraging neural scaling laws [1] to balance model and data complexity. Our experiments demonstrate the accuracy of our approach on synthetic datasets, with an average estimation error of 1.18%, and highlight its effectiveness as a tool for evaluating lossless compression methods on real-world datasets.
Ziqiao Ao, Zhaoyi Sun, Jie Sun 0007
DCC1
2024 On Estimating the Gradient of the Expected Information Gain in Bayesian Experimental Design
abstract
Bayesian Experimental Design (BED), which aims to find the optimal experimental conditions for Bayesian inference, is usually posed as to optimize the expected information gain (EIG). The gradient information is often needed for efficient EIG optimization, and as a result the ability to estimate the gradient of EIG is essential for BED problems. The primary goal of this work is to develop methods for estimating the gradient of EIG, which, combined with the stochastic gradient descent algorithms, result in efficient optimization of EIG. Specifically, we first introduce a posterior expected representation of the EIG gradient with respect to the design variables. Based on this, we propose two methods for estimating the EIG gradient, UEEG-MCMC that leverages posterior samples generated through Markov Chain Monte Carlo (MCMC) to estimate the EIG gradient, and BEEG-AP that focuses on achieving high simulation efficiency by repeatedly using parameter samples. Theoretical analysis and numerical studies illustrate that UEEG-MCMC is robust agains the actual EIG value, while BEEG-AP is more efficient when the EIG value to be optimized is small. Moreover, both methods show superior performance compared to several popular benchmarks in our numerical experiments.
Ziqiao Ao, Jinglai Li
AAAI1
2023 Entropy estimation via uniformization
abstract
Entropy estimation is of practical importance in information theory and statistical science. Many existing entropy estimators suffer from fast growing estimation bias with respect to dimensionality, rendering them unsuitable for high-dimensional problems. In this work we propose a transform-based method for high-dimensional entropy estimation, which consists of the following two main ingredients. Firstly, we provide a modified k-nearest neighbors (k-NN) entropy estimator that can reduce estimation bias for samples closely resembling a uniform distribution. Second we design a normalizing flow based mapping that pushes samples toward the uniform distribution, and the relation between the entropy of the original samples and the transformed ones is also derived. As a result the entropy of a given set of samples is estimated by first transforming them toward the uniform distribution and then applying the proposed estimator to the transformed samples. The performance of the proposed method is compared against several existing entropy estimators, with both mathematical examples and real-world applications.
Ziqiao Ao, Jinglai Li
Artif. Intell.1
2023 Skill requirements in job advertisements: A comparison of skill-categorization methods based on wage regressions
Ziqiao Ao, Gergely Horváth, Chunyuan Sheng
Inf. Process. Manag.1
2022 Entropy Estimation via Normalizing Flow
abstract
Entropy estimation is an important problem in information theory and statistical science. Many popular entropy estimators suffer from fast growing estimation bias with respect to dimensionality, rendering them unsuitable for high dimensional problems. In this work we propose a transformbased method for high dimensional entropy estimation, which consists of the following two main ingredients. First by modifying the k-NN based entropy estimator, we propose a new estimator which enjoys small estimation bias for samples that are close to a uniform distribution. Second we design a normalizing flow based mapping that pushes samples toward a uniform distribution, and the relation between the entropy of the original samples and the transformed ones is also derived. As a result the entropy of a given set of samples is estimated by first transforming them toward a uniform distribution and then applying the proposed estimator to the transformed samples. Numerical experiments demonstrate the effectiveness of the method for high dimensional entropy estimation problems.
Ziqiao Ao, Jinglai Li
AAAI1
2020 An approximate KLD based experimental design for models with intractable likelihoods
abstract
Data collection is a critical step in statistical inference and data science,and the goal of statistical experimental design (ED) is to find the data collection setupthat can provide most information for the inference. In this work we consider a special type of ED problems where the likelihoods are not available in a closed form. In this case, the popular information-theoretic Kullback-Leibler divergence (KLD) based design criterioncan not be used directly, as it requires to evaluate the likelihood function. To address the issue, we derive a new utility function,which is a lower bound of the original KLD utility. This lower bound is expressed in terms of the summation of two or more entropies in the data space, and thus can be evaluated efficiently via entropy estimation methods.We provide several numerical examples to demonstrate the performance of the proposed method.
Ziqiao Ao, Jinglai Li
AISTATS1