Aonan Zhang

dblp:86/40 · DBLP profile ↗
← Back
15ranked-venue papers
9as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 43% Probabilistic and Bayesian machine learning · 34% Representation and self-supervised learning · 8%
Databases, data mining, and information retrieval
2 papers
Data mining · 49% Recommender systems · 33% Graph data management · 10%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 22 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
1.142019
Random Function Priors for Correlation Modeling · ICML 2019
Deep Bayesian Nonparametric Tracking · ICML 2018
Markov Mixed Membership Models · ICML 2015
Natural language and speech › Language models and text generation
mathematical reasoning
0.912025
Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo · ICLR 2025
Natural language and speech › Language models and text generation › large language model reasoning
multi-step reasoning
0.912025
Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo · ICLR 2025
Recommender systems
collaborative filtering
0.712023
Graph-Based Model-Agnostic Data Subsampling for Recommendation Systems · KDD 2023
Data mining › sampling
subsampling
0.712023
Graph-Based Model-Agnostic Data Subsampling for Recommendation Systems · KDD 2023
Mathematical optimization
statistical estimation
0.512021
Nonuniform Negative Sampling and Log Odds Correction with Rare Events Data · NeurIPS 2021
Machine learning › Representation and self-supervised learning
correlation modeling
0.412019
Random Function Priors for Correlation Modeling · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
factor analysis
0.412019
Random Function Priors for Correlation Modeling · ICML 2019
Machine learning › Time series and sequential data › time series analysis
deep learning for time series
0.312018
Deep Bayesian Nonparametric Tracking · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent feature model
0.212016
Markov Latent Feature Models · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.212016
Markov Latent Feature Models · ICML 2016
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.212016
Markov Latent Feature Models · ICML 2016
Natural language and speech › Information extraction and text analysis › topic model
correlated topic model
0.212015
Markov Mixed Membership Models · ICML 2015
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixed membership models
0.212015
Markov Mixed Membership Models · ICML 2015
Natural language and speech › Information extraction and text analysis
topic model
0.212015
Markov Mixed Membership Models · ICML 2015
Natural language and speech › Speech recognition and synthesis › acoustic model training
discriminative training
0.212014
Max-Margin Infinite Hidden Markov Models · ICML 2014
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › hidden markov model
infinite hidden markov model
0.212014
Max-Margin Infinite Hidden Markov Models · ICML 2014
Machine learning › Optimization for machine learning › regularized risk minimization
max-margin learning
0.212014
Max-Margin Infinite Hidden Markov Models · ICML 2014
Data mining › text mining › topic modeling
sparse topic model
0.212013
Sparse online topic models · WWW 2013
Information retrieval
text analysis
0.212013
Sparse online topic models · WWW 2013
Data mining › text mining
topic model
0.212013
Sparse online topic models · WWW 2013
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
likelihood estimation
0.112021
Nonuniform Negative Sampling and Log Odds Correction with Rare Events Data · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

maximum likelihood estimation · 1.0inverse probability weighting · 1.0verification · 0.9twisted sequential monte carlo · 0.9reward estimation · 0.9multimodal instruction tuning · 0.8large-scale pretraining · 0.8model-agnostic subsampling · 0.7importance propagation · 0.7graph conductance · 0.7representation theorem · 0.4neural network · 0.4amortized variational inference · 0.4sparsity-inducing regularization · 0.2online learning · 0.2
YearPublicationVenuePosition
2025 Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo
abstract
Augmenting the multi-step reasoning abilities of Large Language Models (LLMs) has been a persistent challenge. Recently, verification has shown promise in improving solution consistency by evaluating generated outputs. However, current verification approaches suffer from sampling inefficiencies, requiring a large number of samples to achieve satisfactory performance. Additionally, training an effective verifier often depends on extensive process supervision, which is costly to acquire. In this paper, we address these limitations by introducing a novel verification method based on Twisted Sequential Monte Carlo (TSMC). TSMC sequentially refines its sampling effort to focus exploration on promising candidates, resulting in more efficient generation of high-quality solutions. We apply TSMC to LLMs by estimating the expected future rewards at partial solutions. This approach results in a more straightforward training target that eliminates the need for step-wise human annotations. We empirically demonstrate the advantages of our method across multiple math benchmarks, and also validate our theoretical analysis of both our approach and existing verification methods.
Shengyu Feng, Xiang Kong, Aonan Zhang, Ruoming Pang, Yiming Yang 0002
ICLR4
2024 MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang 0002, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Anton Belyi, Haotian Zhang 0005, Karanjeet Singh 0003, Doug Kang, Hongyu Hè, Max Schwarzer, Tom Gunter, Xiang Kong, Aonan Zhang, Nan Du 0002, Tao Lei 0001, Sam Wiseman, Mark Lee 0003, Ruoming Pang, Peter Grasch, Alexander Toshev, Yinfei Yang
ECCV (29)18
2023 Graph-Based Model-Agnostic Data Subsampling for Recommendation Systems
abstract
Data subsampling is widely used to speed up the training of large-scale recommendation systems. Most subsampling methods are model-based and often require a pre-trained pilot model to measure data importance via e.g. sample hardness. However, when the pilot model is misspecified, model-based subsampling methods deteriorate. Since model misspecification is persistent in real recommendation systems, we instead propose model-agnostic data subsampling methods by only exploring input data structure represented by graphs. Specifically, we study the topology of the user-item graph to estimate the importance of each user-item interaction (an edge in the user-item graph) via graph conductance, followed by a propagation step on the network to smooth out the estimated importance value. Since our proposed method is model-agnostic, we can marry the merits of both model-agnostic and model-based subsampling methods. Empirically, we show that combing the two consistently improves over any single method on the used datasets. Experimental results on KuaiRec and MIND datasets demonstrate that our proposed methods achieve superior results compared to baseline approaches.
Jiankai Sun, Taiqing Wang, Ruocheng Guo, Liping Liu 0001, Aonan Zhang
KDD6
2021 Nonuniform Negative Sampling and Log Odds Correction with Rare Events Data
abstract
We investigate the issue of parameter estimation with nonuniform negative sampling for imbalanced data. We first prove that, with imbalanced data, the available information about unknown parameters is only tied to the relatively small number of positive instances, which justifies the usage of negative sampling. However, if the negative instances are subsampled to the same level of the positive cases, there is information loss. To maintain more information, we derive the asymptotic distribution of a general inverse probability weighted (IPW) estimator and obtain the optimal sampling probability that minimizes its variance. To further improve the estimation efficiency over the IPW method, we propose a likelihood-based estimator by correcting log odds for the sampled data and prove that the improved estimator has the smallest asymptotic variance among a large class of estimators. It is also more robust to pilot misspecification. We validate our approach on simulated data as well as a real click-through rate dataset with more than 0.3 trillion instances, collected over a period of a month. Both theoretical and empirical results demonstrate the effectiveness of our method.
Aonan Zhang, Chong Wang 0002
NeurIPS2
2019 Fully Supervised Speaker Diarization
abstract
In this paper, we propose a fully supervised speaker diarization approach, named unbounded interleaved-state recurrent neural networks (UIS-RNN). Given extracted speaker-discriminative embeddings (a.k.a. d-vectors) from input utterances, each individual speaker is modeled by a parameter-sharing RNN, while the RNN states for different speakers interleave in the time domain. This RNN is naturally integrated with a distance-dependent Chinese restaurant process (ddCRP) to accommodate an unknown number of speakers. Our system is fully supervised and is able to learn from examples where time-stamped speaker labels are annotated. We achieved a 7.6% diarization error rate on NIST SRE 2000 CALLHOME, which is better than the state-of-the-art method using spectral clustering. Moreover, our method decodes in an online fashion while most state-of-the-art systems rely on offline clustering.
Aonan Zhang, Zhenyao Zhu, John W. Paisley, Chong Wang 0002
ICASSP1
2019 Random Function Priors for Correlation Modeling
abstract
The likelihood model of high dimensional data $X_n$ can often be expressed as $p(X_n|Z_n,\theta)$, where $\theta\mathrel{\mathop:}=(\theta_k)_{k\in[K]}$ is a collection of hidden features shared across objects, indexed by $n$, and $Z_n$ is a non-negative factor loading vector with $K$ entries where $Z_{nk}$ indicates the strength of $\theta_k$ used to express $X_n$. In this paper, we introduce random function priors for $Z_n$ for modeling correlations among its $K$ dimensions $Z_{n1}$ through $Z_{nK}$, which we call population random measure embedding (PRME). Our model can be viewed as a generalized paintbox model \cite{Broderick13} using random functions, and can be learned efficiently with neural networks via amortized variational inference. We derive our Bayesian nonparametric method by applying a representation theorem on separately exchangeable discrete random measures.
Aonan Zhang, John W. Paisley
ICML1
2018 Asymptotic Simulated Annealing for Variational Inference
abstract
Variational inference (VI) is an effective deterministic method for approximate posterior inference, which arises in many practical applications. However, it typically suffers from non-convexity issues. This paper proposes a novel optimization tool called asymptotically-annealed variational inference (AVI), for better local optimal convergence of VI by using ideas from small-variance asymptotics to efficiently search for better solutions. The algorithm entails a simple modification to the basic VI algorithm, has little additional computational cost and is very simple. Furthermore, our algorithm can be viewed as an asymptotic limit of simulated annealing, connecting it to a recent literature in machine learning on deterministic versions of stochastic algorithms. Experiments show better convergence performance than VI and other annealing methods for models such as LDA and the HMM, as well as on stochastic variational inference problems for big data.
San Gultekin, Aonan Zhang, John W. Paisley
GLOBECOM2
2018 Deep Bayesian Nonparametric Tracking
abstract
Time-series data often exhibit irregular behavior, making them hard to analyze and explain with a simple dynamic model. For example, information in social networks may show change-point-like bursts that then diffuse with smooth dynamics. Powerful models such as deep neural networks learn smooth functions from data, but are not as well-suited (in off-the-shelf form) for discovering and explaining sparse, discrete and bursty dynamic patterns. Bayesian models can do this well by encoding the appropriate probabilistic assumptions in the model prior. We propose an integration of Bayesian nonparametric methods within deep neural networks for modeling irregular patterns in time-series data. We use a Bayesian nonparametrics to model change-point behavior in time, and a deep neural network to model nonlinear latent space dynamics. We compare with a non-deep linear version of the model also proposed here. Empirical evaluations demonstrates improved performance and interpretable results when tracking stock prices and Twitter trends.
Aonan Zhang, John W. Paisley
ICML1
2016 Stochastic Variational Inference for the HDP-HMM
abstract
We derive a variational inference algorithm for the HDP-HMM based on the two-level stick breaking construction. This construction has previously been applied to the hierarchical Dirichlet processes (HDP) for mixed membership models, allowing for efficient handling of the coupled weight parameters. However, the same algorithm is not directly applicable to HDP-based infinite hidden Markov models (HDP-HMM) because of extra sequential dependencies in the Markov chain. In this paper we provide a solution to this problem by deriving a variational inference algorithm for the HDP-HMM, as well as its stochastic extension, for which all parameter updates are in closed form. We apply our algorithm to sequential text analysis and audio signal analysis, comparing our results with the beam-sampled iHMM, the parametric HMM, and other variational inference approximations.
Aonan Zhang, San Gultekin, John W. Paisley
AISTATS1
2016 Markov Latent Feature Models
abstract
We introduce Markov latent feature models (MLFM), a sparse latent feature model that arises naturally from a simple sequential construction. The key idea is to interpret each state of a sequential process as corresponding to a latent feature, and the set of states visited between two null-state visits as picking out features for an observation. We show that, given some natural constraints, we can represent this stochastic process as a mixture of recurrent Markov chains. In this way we can perform correlated latent feature modeling for the sparse coding problem. We demonstrate two cases in which we define finite and infinite latent feature models constructed from first-order Markov chains, and derive their associated scalable inference algorithms. We show empirical results on a genome analysis task and an image denoising task.
Aonan Zhang, John W. Paisley
ICML1
2015 Markov Mixed Membership Models
abstract
We present a Markov mixed membership model (Markov M3) for grouped data that learns a fully connected graph structure among mixing components. A key feature of Markov M3 is that it interprets the mixed membership assignment as a Markov random walk over this graph of nodes. This is in contrast to tree-structured models in which the assignment is done according to a tree structure on the mixing components. The Markov structure results in a simple parametric model that can learn a complex dependency structure between nodes, while still maintaining full conjugacy for closed-form stochastic variational inference. Empirical results demonstrate that Markov M3 performs well compared with tree structured topic models, and can learn meaningful dependency structure between topics.
Aonan Zhang, John W. Paisley
ICML1
2014 Max-Margin Infinite Hidden Markov Models
abstract
Infinite hidden Markov models (iHMMs) are nonparametric Bayesian extensions of hidden Markov models (HMMs) with an infinite number of states. Though flexible in describing sequential data, the generative formulation of iHMMs could limit their discriminative ability in sequential prediction tasks. Our paper introduces max-margin infinite HMMs (M2iHMMs), new infinite HMMs that explore the max-margin principle for discriminative learning. By using the theory of Gibbs classifiers and data augmentation, we develop efficient beam sampling algorithms without making restricting mean-field assumptions or truncated approximation. For single variate classification, M2iHMMs reduce to a new formulation of DP mixtures of max-margin machines. Empirical results on synthetic and real data sets show that our methods obtain superior performance than other competitors in both single variate classification and sequential prediction tasks.
Aonan Zhang, Jun Zhu 0001, Bo Zhang 0010
ICML1
2014 Max-margin latent feature relational models for entity-attribute networks
abstract
Link prediction is a fundamental task in statistical analysis of network data. Though much research has concentrated on predicting entity-entity relationships in homogeneous networks, it has attracted increasing attentions to predict relationships in heterogeneous networks, which consist of multiple types of nodes and relational links. Existing work on heterogeneous network link prediction mainly focuses on using input features that are explicitly extracted by humans. This paper presents an approach to automatically learn latent features from partially observed heterogeneous networks, with a particular focus on entity-attribute networks (EANs), and making predictions for unseen pairs. To make the latent features discriminative, we adopt the max-margin idea under the framework of maximum entropy discrimination (MED). Our maximum entropy discrimination joint relational model (MED-JRM) can jointly predict entity-entity relationships as well as the missing attributes of entities in EANs. Experimental results on several real networks demonstrate that our model has improved performance over state-of-the-art homogeneous and heterogeneous network link prediction algorithms.
Fei Xia 0005, Ning Chen 0002, Jun Zhu 0001, Aonan Zhang, Xiaoming Jin
IJCNN4
2013 Sparse Relational Topic Models for Document Networks
Aonan Zhang, Jun Zhu 0001, Bo Zhang 0010
ECML/PKDD (1)1
2013 Sparse online topic models
abstract
Topic models have shown great promise in discovering latent semantic structures from complex data corpora, ranging from text documents and web news articles to images, videos, and even biological data. In order to deal with massive data collections and dynamic text streams, probabilistic online topic models such as online latent Dirichlet allocation (OLDA) have recently been developed. However, due to normalization constraints, OLDA can be ineffective in controlling the sparsity of discovered representations, a desirable property for learning interpretable semantic patterns, especially when the total number of topics is large. In contrast, sparse topical coding (STC) has been successfully introduced as a non-probabilistic topic model for effectively discovering sparse latent patterns by using sparsity-inducing regularization. But, unfortunately STC cannot scale to very large datasets or deal with online text streams, partly due to its batch learning procedure. In this paper, we present a sparse online topic model, which directly controls the sparsity of latent semantic patterns by imposing sparsity-inducing regularization and learns the topical dictionary by an online algorithm. The online algorithm is efficient and guaranteed to converge. Extensive empirical results of the sparse online topic model as well as its collapsed and supervised extensions on a large-scale Wikipedia dataset and the medium-sized 20Newsgroups dataset demonstrate appealing performance.
Aonan Zhang, Jun Zhu 0001, Bo Zhang 0010
WWW1