XianXing Zhang

dblp:51/10623 · also XianXing "Sean" Zhang · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 46% Information extraction and text analysis · 38% Kernel, tree and ensemble methods · 16%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel computing
parallel optimization
0.212016
GLMix: Generalized Linear Mixed Models For Large-Scale Response Prediction · KDD 2016
Natural language and speech › Information extraction and text analysis
topic model
0.112012
Joint Modeling of a Matrix with Associated Text via Latent Binary Features · NIPS 2012
Recommender systems › collaborative filtering
matrix factorization
0.112012
Joint Modeling of a Matrix with Associated Text via Latent Binary Features · NIPS 2012
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
factor analysis
0.112011
Tree-Structured Infinite Sparse Factor Model · ICML 2011
Machine learning › Probabilistic and Bayesian machine learning › hierarchical modeling
hierarchical bayesian model
0.112011
Tree-Structured Infinite Sparse Factor Model · ICML 2011
Natural language and speech › Information extraction and text analysis › topic model
nonparametric topic modeling
0.112011
Hierarchical Topic Modeling for Analysis of Time-Evolving Personal Choices · NIPS 2011
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › factor analysis
sparse factor model
0.112011
Tree-Structured Infinite Sparse Factor Model · ICML 2011
Machine learning › Kernel, tree and ensemble methods
tree-based models
0.112011
Tree-Structured Infinite Sparse Factor Model · ICML 2011
Natural language and speech › Information extraction and text analysis › topic model
hierarchical topic model
0.012011
Hierarchical Topic Modeling for Analysis of Time-Evolving Personal Choices · NIPS 2011

Methods — techniques the papers use, named apart from their topics

bulk synchronous parallel · 0.5block coordinate descent · 0.5low-rank constraint · 0.3joint modeling · 0.3poisson and product-of-gammas construction · 0.1nested chinese restaurant process · 0.1infinite sparse factor model · 0.1dirichlet process · 0.1change-point stick-breaking · 0.1
YearPublicationVenuePosition
2016 GLMix: Generalized Linear Mixed Models For Large-Scale Response Prediction
abstract
Generalized linear model (GLM) is a widely used class of models for statistical inference and response prediction problems. For instance, in order to recommend relevant content to a user or optimize for revenue, many web companies use logistic regression models to predict the probability of the user's clicking on an item (e.g., ad, news article, job). In scenarios where the data is abundant, having a more fine-grained model at the user or item level would potentially lead to more accurate prediction, as the user's personal preferences on items and the item's specific attraction for users can be better captured. One common approach is to introduce ID-level regression coefficients in addition to the global regression coefficients in a GLM setting, and such models are called generalized linear mixed models (GLMix) in the statistical literature. However, for big data sets with a large number of ID-level coefficients, fitting a GLMix model can be computationally challenging. In this paper, we report how we successfully overcame the scalability bottleneck by applying parallelized block coordinate descent under the Bulk Synchronous Parallel (BSP) paradigm. We deployed the model in the LinkedIn job recommender system, and generated 20% to 40% more job applications for job seekers on LinkedIn.
XianXing Zhang, Bee-Chung Chen, Liang Zhang 0021, Deepak Agarwal
KDD1
2012 Joint Modeling of a Matrix with Associated Text via Latent Binary Features
abstract
A new methodology is developed for joint analysis of a matrix and accompanying documents, with the documents associated with the matrix rows/columns. The documents are modeled with a focused topic model, inferring latent binary features (topics) for each document. A new matrix decomposition is developed, with latent binary features associated with the rows/columns, and with imposition of a low-rank constraint. The matrix decomposition and topic model are coupled by sharing the latent binary feature vectors associated with each. The model is applied to roll-call data, with the associated documents defined by the legislation. State-of-the-art results are manifested for prediction of votes on a new piece of legislation, based only on the observed text legislation. The coupling of the text and legislation is also demonstrated to yield insight into the properties of the matrix decomposition for roll-call data.
XianXing Zhang, Lawrence Carin
NIPS1
2012 Nested Dictionary Learning for Hierarchical Organization of Imagery and Text
Lingbo Li 0002, XianXing Zhang, Mingyuan Zhou, Lawrence Carin
UAI2
2011 Tree-Structured Infinite Sparse Factor Model
XianXing Zhang, David B. Dunson, Lawrence Carin
ICML1
2011 Hierarchical Topic Modeling for Analysis of Time-Evolving Personal Choices
abstract
The nested Chinese restaurant process is extended to design a nonparametric topic-model tree for representation of human choices. Each tree branch corresponds to a type of person, and each node (topic) has a corresponding probability vector over items that may be selected. The observed data are assumed to have associated temporal covariates (corresponding to the time at which choices are made), and we wish to impose that with increasing time it is more probable that topics deeper in the tree are utilized. This structure is imposed by developing a new “change point" stick-breaking model that is coupled with a Poisson and product-of-gammas construction. To share topics across the tree nodes, topic distributions are drawn from a Dirichlet process. As a demonstration of this concept, we analyze real data on course selections of undergraduate students at Duke University, with the goal of uncovering and concisely representing structure in the curriculum and in the characteristics of the student body.
XianXing Zhang, David B. Dunson, Lawrence Carin
NIPS1