VLDB 2026 Research / reviewers in the wild / expert
Aonan Zhang
dblp:86/40
· DBLP profile ↗
15ranked-venue papers
9as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Language models and text generation · 43% Probabilistic and Bayesian machine learning · 34% Representation and self-supervised learning · 8% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 49% Recommender systems · 33% Graph data management · 10% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 22 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model |
1.1 | 4 | 2019 | Random Function Priors for Correlation Modeling · ICML 2019 Deep Bayesian Nonparametric Tracking · ICML 2018 Markov Mixed Membership Models · ICML 2015 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.9 | 1 | 2025 | Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model reasoning
multi-step reasoning |
0.9 | 1 | 2025 | Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo · ICLR 2025 |
Recommender systems
collaborative filtering |
0.7 | 1 | 2023 | Graph-Based Model-Agnostic Data Subsampling for Recommendation Systems · KDD 2023 |
Data mining › sampling
subsampling |
0.7 | 1 | 2023 | Graph-Based Model-Agnostic Data Subsampling for Recommendation Systems · KDD 2023 |
Mathematical optimization
statistical estimation |
0.5 | 1 | 2021 | Nonuniform Negative Sampling and Log Odds Correction with Rare Events Data · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning
correlation modeling |
0.4 | 1 | 2019 | Random Function Priors for Correlation Modeling · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
factor analysis |
0.4 | 1 | 2019 | Random Function Priors for Correlation Modeling · ICML 2019 |
Machine learning › Time series and sequential data › time series analysis
deep learning for time series |
0.3 | 1 | 2018 | Deep Bayesian Nonparametric Tracking · ICML 2018 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent feature model |
0.2 | 1 | 2016 | Markov Latent Feature Models · ICML 2016 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.2 | 1 | 2016 | Markov Latent Feature Models · ICML 2016 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding |
0.2 | 1 | 2016 | Markov Latent Feature Models · ICML 2016 |
Natural language and speech › Information extraction and text analysis › topic model
correlated topic model |
0.2 | 1 | 2015 | Markov Mixed Membership Models · ICML 2015 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixed membership models |
0.2 | 1 | 2015 | Markov Mixed Membership Models · ICML 2015 |
Natural language and speech › Information extraction and text analysis
topic model |
0.2 | 1 | 2015 | Markov Mixed Membership Models · ICML 2015 |
Natural language and speech › Speech recognition and synthesis › acoustic model training
discriminative training |
0.2 | 1 | 2014 | Max-Margin Infinite Hidden Markov Models · ICML 2014 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › hidden markov model
infinite hidden markov model |
0.2 | 1 | 2014 | Max-Margin Infinite Hidden Markov Models · ICML 2014 |
Machine learning › Optimization for machine learning › regularized risk minimization
max-margin learning |
0.2 | 1 | 2014 | Max-Margin Infinite Hidden Markov Models · ICML 2014 |
Data mining › text mining › topic modeling
sparse topic model |
0.2 | 1 | 2013 | Sparse online topic models · WWW 2013 |
Information retrieval
text analysis |
0.2 | 1 | 2013 | Sparse online topic models · WWW 2013 |
Data mining › text mining
topic model |
0.2 | 1 | 2013 | Sparse online topic models · WWW 2013 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
likelihood estimation |
0.1 | 1 | 2021 | Nonuniform Negative Sampling and Log Odds Correction with Rare Events Data · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
maximum likelihood estimation · 1.0inverse probability weighting · 1.0verification · 0.9twisted sequential monte carlo · 0.9reward estimation · 0.9multimodal instruction tuning · 0.8large-scale pretraining · 0.8model-agnostic subsampling · 0.7importance propagation · 0.7graph conductance · 0.7representation theorem · 0.4neural network · 0.4amortized variational inference · 0.4sparsity-inducing regularization · 0.2online learning · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte CarloabstractAugmenting the multi-step reasoning abilities of Large Language Models (LLMs) has been a persistent challenge. Recently, verification has shown promise in improving solution consistency by evaluating generated outputs. However, current verification approaches suffer from sampling inefficiencies, requiring a large number of samples to achieve satisfactory performance. Additionally, training an effective verifier often depends on extensive process supervision, which is costly to acquire. In this paper, we address these limitations by introducing a novel verification method based on Twisted Sequential Monte Carlo (TSMC). TSMC sequentially refines its sampling effort to focus exploration on promising candidates, resulting in more efficient generation of high-quality solutions. We apply TSMC to LLMs by estimating the expected future rewards at partial solutions. This approach results in a more straightforward training target that eliminates the need for step-wise human annotations. We empirically demonstrate the advantages of our method across multiple math benchmarks, and also validate our theoretical analysis of both our approach and existing verification methods. Shengyu Feng, Xiang Kong, Aonan Zhang, Ruoming Pang, Yiming Yang 0002 |
ICLR | 4 |
| 2024 | MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang 0002, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Anton Belyi, Haotian Zhang 0005, Karanjeet Singh 0003, Doug Kang, Hongyu Hè, Max Schwarzer, Tom Gunter, Xiang Kong, Aonan Zhang, Nan Du 0002, Tao Lei 0001, Sam Wiseman, Mark Lee 0003, Ruoming Pang, Peter Grasch, Alexander Toshev, Yinfei Yang |
ECCV (29) | 18 |
| 2023 | Graph-Based Model-Agnostic Data Subsampling for Recommendation SystemsabstractData subsampling is widely used to speed up the training of large-scale recommendation systems. Most subsampling methods are model-based and often require a pre-trained pilot model to measure data importance via e.g. sample hardness. However, when the pilot model is misspecified, model-based subsampling methods deteriorate. Since model misspecification is persistent in real recommendation systems, we instead propose model-agnostic data subsampling methods by only exploring input data structure represented by graphs. Specifically, we study the topology of the user-item graph to estimate the importance of each user-item interaction (an edge in the user-item graph) via graph conductance, followed by a propagation step on the network to smooth out the estimated importance value. Since our proposed method is model-agnostic, we can marry the merits of both model-agnostic and model-based subsampling methods. Empirically, we show that combing the two consistently improves over any single method on the used datasets. Experimental results on KuaiRec and MIND datasets demonstrate that our proposed methods achieve superior results compared to baseline approaches. Jiankai Sun, Taiqing Wang, Ruocheng Guo, Liping Liu 0001, Aonan Zhang |
KDD | 6 |
| 2021 | Nonuniform Negative Sampling and Log Odds Correction with Rare Events DataabstractWe investigate the issue of parameter estimation with nonuniform negative sampling for imbalanced data. We first prove that, with imbalanced data, the available information about unknown parameters is only tied to the relatively small number of positive instances, which justifies the usage of negative sampling. However, if the negative instances are subsampled to the same level of the positive cases, there is information loss. To maintain more information, we derive the asymptotic distribution of a general inverse probability weighted (IPW) estimator and obtain the optimal sampling probability that minimizes its variance. To further improve the estimation efficiency over the IPW method, we propose a likelihood-based estimator by correcting log odds for the sampled data and prove that the improved estimator has the smallest asymptotic variance among a large class of estimators. It is also more robust to pilot misspecification. We validate our approach on simulated data as well as a real click-through rate dataset with more than 0.3 trillion instances, collected over a period of a month. Both theoretical and empirical results demonstrate the effectiveness of our method. Aonan Zhang, Chong Wang 0002 |
NeurIPS | 2 |
| 2019 | Fully Supervised Speaker DiarizationabstractIn this paper, we propose a fully supervised speaker diarization approach, named unbounded interleaved-state recurrent neural networks (UIS-RNN). Given extracted speaker-discriminative embeddings (a.k.a. d-vectors) from input utterances, each individual speaker is modeled by a parameter-sharing RNN, while the RNN states for different speakers interleave in the time domain. This RNN is naturally integrated with a distance-dependent Chinese restaurant process (ddCRP) to accommodate an unknown number of speakers. Our system is fully supervised and is able to learn from examples where time-stamped speaker labels are annotated. We achieved a 7.6% diarization error rate on NIST SRE 2000 CALLHOME, which is better than the state-of-the-art method using spectral clustering. Moreover, our method decodes in an online fashion while most state-of-the-art systems rely on offline clustering. Aonan Zhang, Zhenyao Zhu, John W. Paisley, Chong Wang 0002 |
ICASSP | 1 |
| 2019 | Random Function Priors for Correlation ModelingabstractThe likelihood model of high dimensional data $X_n$ can often be expressed as $p(X_n|Z_n,\theta)$, where $\theta\mathrel{\mathop:}=(\theta_k)_{k\in[K]}$ is a collection of hidden features shared across objects, indexed by $n$, and $Z_n$ is a non-negative factor loading vector with $K$ entries where $Z_{nk}$ indicates the strength of $\theta_k$ used to express $X_n$. In this paper, we introduce random function priors for $Z_n$ for modeling correlations among its $K$ dimensions $Z_{n1}$ through $Z_{nK}$, which we call population random measure embedding (PRME). Our model can be viewed as a generalized paintbox model \cite{Broderick13} using random functions, and can be learned efficiently with neural networks via amortized variational inference. We derive our Bayesian nonparametric method by applying a representation theorem on separately exchangeable discrete random measures. Aonan Zhang, John W. Paisley |
ICML | 1 |
| 2018 | Asymptotic Simulated Annealing for Variational InferenceabstractVariational inference (VI) is an effective deterministic method for approximate posterior inference, which arises in many practical applications. However, it typically suffers from non-convexity issues. This paper proposes a novel optimization tool called asymptotically-annealed variational inference (AVI), for better local optimal convergence of VI by using ideas from small-variance asymptotics to efficiently search for better solutions. The algorithm entails a simple modification to the basic VI algorithm, has little additional computational cost and is very simple. Furthermore, our algorithm can be viewed as an asymptotic limit of simulated annealing, connecting it to a recent literature in machine learning on deterministic versions of stochastic algorithms. Experiments show better convergence performance than VI and other annealing methods for models such as LDA and the HMM, as well as on stochastic variational inference problems for big data. San Gultekin, Aonan Zhang, John W. Paisley |
GLOBECOM | 2 |
| 2018 | Deep Bayesian Nonparametric TrackingabstractTime-series data often exhibit irregular behavior, making them hard to analyze and explain with a simple dynamic model. For example, information in social networks may show change-point-like bursts that then diffuse with smooth dynamics. Powerful models such as deep neural networks learn smooth functions from data, but are not as well-suited (in off-the-shelf form) for discovering and explaining sparse, discrete and bursty dynamic patterns. Bayesian models can do this well by encoding the appropriate probabilistic assumptions in the model prior. We propose an integration of Bayesian nonparametric methods within deep neural networks for modeling irregular patterns in time-series data. We use a Bayesian nonparametrics to model change-point behavior in time, and a deep neural network to model nonlinear latent space dynamics. We compare with a non-deep linear version of the model also proposed here. Empirical evaluations demonstrates improved performance and interpretable results when tracking stock prices and Twitter trends. Aonan Zhang, John W. Paisley |
ICML | 1 |
| 2016 | Stochastic Variational Inference for the HDP-HMMabstractWe derive a variational inference algorithm for the HDP-HMM based on the two-level stick breaking construction. This construction has previously been applied to the hierarchical Dirichlet processes (HDP) for mixed membership models, allowing for efficient handling of the coupled weight parameters. However, the same algorithm is not directly applicable to HDP-based infinite hidden Markov models (HDP-HMM) because of extra sequential dependencies in the Markov chain. In this paper we provide a solution to this problem by deriving a variational inference algorithm for the HDP-HMM, as well as its stochastic extension, for which all parameter updates are in closed form. We apply our algorithm to sequential text analysis and audio signal analysis, comparing our results with the beam-sampled iHMM, the parametric HMM, and other variational inference approximations. Aonan Zhang, San Gultekin, John W. Paisley |
AISTATS | 1 |
| 2016 | Markov Latent Feature ModelsabstractWe introduce Markov latent feature models (MLFM), a sparse latent feature model that arises naturally from a simple sequential construction. The key idea is to interpret each state of a sequential process as corresponding to a latent feature, and the set of states visited between two null-state visits as picking out features for an observation. We show that, given some natural constraints, we can represent this stochastic process as a mixture of recurrent Markov chains. In this way we can perform correlated latent feature modeling for the sparse coding problem. We demonstrate two cases in which we define finite and infinite latent feature models constructed from first-order Markov chains, and derive their associated scalable inference algorithms. We show empirical results on a genome analysis task and an image denoising task. Aonan Zhang, John W. Paisley |
ICML | 1 |
| 2015 | Markov Mixed Membership ModelsabstractWe present a Markov mixed membership model (Markov M3) for grouped data that learns a fully connected graph structure among mixing components. A key feature of Markov M3 is that it interprets the mixed membership assignment as a Markov random walk over this graph of nodes. This is in contrast to tree-structured models in which the assignment is done according to a tree structure on the mixing components. The Markov structure results in a simple parametric model that can learn a complex dependency structure between nodes, while still maintaining full conjugacy for closed-form stochastic variational inference. Empirical results demonstrate that Markov M3 performs well compared with tree structured topic models, and can learn meaningful dependency structure between topics. Aonan Zhang, John W. Paisley |
ICML | 1 |
| 2014 | Max-Margin Infinite Hidden Markov ModelsabstractInfinite hidden Markov models (iHMMs) are nonparametric Bayesian extensions of hidden Markov models (HMMs) with an infinite number of states. Though flexible in describing sequential data, the generative formulation of iHMMs could limit their discriminative ability in sequential prediction tasks. Our paper introduces max-margin infinite HMMs (M2iHMMs), new infinite HMMs that explore the max-margin principle for discriminative learning. By using the theory of Gibbs classifiers and data augmentation, we develop efficient beam sampling algorithms without making restricting mean-field assumptions or truncated approximation. For single variate classification, M2iHMMs reduce to a new formulation of DP mixtures of max-margin machines. Empirical results on synthetic and real data sets show that our methods obtain superior performance than other competitors in both single variate classification and sequential prediction tasks. Aonan Zhang, Jun Zhu 0001, Bo Zhang 0010 |
ICML | 1 |
| 2014 | Max-margin latent feature relational models for entity-attribute networksabstractLink prediction is a fundamental task in statistical analysis of network data. Though much research has concentrated on predicting entity-entity relationships in homogeneous networks, it has attracted increasing attentions to predict relationships in heterogeneous networks, which consist of multiple types of nodes and relational links. Existing work on heterogeneous network link prediction mainly focuses on using input features that are explicitly extracted by humans. This paper presents an approach to automatically learn latent features from partially observed heterogeneous networks, with a particular focus on entity-attribute networks (EANs), and making predictions for unseen pairs. To make the latent features discriminative, we adopt the max-margin idea under the framework of maximum entropy discrimination (MED). Our maximum entropy discrimination joint relational model (MED-JRM) can jointly predict entity-entity relationships as well as the missing attributes of entities in EANs. Experimental results on several real networks demonstrate that our model has improved performance over state-of-the-art homogeneous and heterogeneous network link prediction algorithms. Fei Xia 0005, Ning Chen 0002, Jun Zhu 0001, Aonan Zhang, Xiaoming Jin |
IJCNN | 4 |
| 2013 | Sparse Relational Topic Models for Document Networks
Aonan Zhang, Jun Zhu 0001, Bo Zhang 0010 |
ECML/PKDD (1) | 1 |
| 2013 | Sparse online topic modelsabstractTopic models have shown great promise in discovering latent semantic structures from complex data corpora, ranging from text documents and web news articles to images, videos, and even biological data. In order to deal with massive data collections and dynamic text streams, probabilistic online topic models such as online latent Dirichlet allocation (OLDA) have recently been developed. However, due to normalization constraints, OLDA can be ineffective in controlling the sparsity of discovered representations, a desirable property for learning interpretable semantic patterns, especially when the total number of topics is large. In contrast, sparse topical coding (STC) has been successfully introduced as a non-probabilistic topic model for effectively discovering sparse latent patterns by using sparsity-inducing regularization. But, unfortunately STC cannot scale to very large datasets or deal with online text streams, partly due to its batch learning procedure. In this paper, we present a sparse online topic model, which directly controls the sparsity of latent semantic patterns by imposing sparsity-inducing regularization and learns the topical dictionary by an online algorithm. The online algorithm is efficient and guaranteed to converge. Extensive empirical results of the sparse online topic model as well as its collapsed and supervised extensions on a large-scale Wikipedia dataset and the medium-sized 20Newsgroups dataset demonstrate appealing performance. Aonan Zhang, Jun Zhu 0001, Bo Zhang 0010 |
WWW | 1 |