Zhixiang Eddie Xu

dblp:84/9669 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-authorDatabases, data management, data science and information retrieval · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Efficient and distributed learning · 32% Optimization for machine learning · 16% Deep learning architectures and training · 13%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › autoencoder
denoising autoencoder
0.422015
Marginalizing stacked linear denoising autoencoders · J. Mach. Learn. Res. 2015
Marginalized Denoising Autoencoders for Domain Adaptation · ICML 2012
Computer vision › Image recognition and object detection › object detection
cascade classifier
0.422014
Classifier cascades and trees for minimizing feature evaluation cost · J. Mach. Learn. Res. 2014
Cost-Sensitive Tree of Classifiers · ICML (1) 2013
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.212014
Bayesian Optimization with Inequality Constraints · ICML 2014
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
constrained bayesian optimization
0.212014
Bayesian Optimization with Inequality Constraints · ICML 2014
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection
0.212014
Gradient boosted feature selection · KDD 2014
Machine learning › Efficient and distributed learning
model compression
0.212014
Classifier cascades and trees for minimizing feature evaluation cost · J. Mach. Learn. Res. 2014
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
anytime algorithm
0.212013
Anytime Representation Learning · ICML (3) 2013
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.112012
Marginalized Denoising Autoencoders for Domain Adaptation · ICML 2012
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.112015
Marginalizing stacked linear denoising autoencoders · J. Mach. Learn. Res. 2015
Machine learning › Learning theory
classification
0.112014
Classifier cascades and trees for minimizing feature evaluation cost · J. Mach. Learn. Res. 2014
Machine learning › Kernel, tree and ensemble methods › gradient boosting
gradient boosted decision trees
0.112014
Gradient boosted feature selection · KDD 2014

Methods — techniques the papers use, named apart from their topics

denoising autoencoder · 0.4cost-sensitive learning · 0.3marginalization · 0.2submodularity · 0.2gradient boosting · 0.2gradient boosted trees · 0.2gaussian process prior · 0.2decision tree · 0.2cascade · 0.2bayesian optimization · 0.2
YearPublicationVenuePosition
2015 Marginalizing stacked linear denoising autoencoders
Minmin Chen, Kilian Q. Weinberger, Zhixiang Eddie Xu, Fei Sha
J. Mach. Learn. Res.3
2014 Feature-Cost Sensitive Learning with Submodular Trees of Classifiers
abstract
During the past decade, machine learning algorithms have become commonplace in large-scale real-world industrial applications. In these settings, the computation time to train and test machine learning algorithms is a key consideration. At training-time the algorithms must scale to very large data set sizes.At testing-time, the cost of feature extraction can dominate the CPU runtime. Recently, a promising method was proposed to account for the feature extraction cost at testing time, called Cost-sensitive Tree of Classifiers (CSTC). Although the CSTC problem is NP-hard, the authors suggest an approximation through a mixed-norm relaxation across many classifiers. This relaxation is slow to train and requires involved optimization hyperparameter tuning. We propose a different relaxation using approximate submodularity, called Approximately Submodular Tree of Classifiers (ASTC). ASTC is much simpler to implement, yields equivalent results but requires no optimization hyperparameter tuning and is up to two orders of magnitude faster to train.
Matt J. Kusner, Zhixiang Eddie Xu, Kilian Q. Weinberger, Yixin Chen 0001
AAAI4
2014 Bayesian Optimization with Inequality Constraints
abstract
Bayesian optimization is a powerful framework for minimizing expensive objective functions while using very few function evaluations. It has been successfully applied to a variety of problems, including hyperparameter tuning and experimental design. However, this framework has not been extended to the inequality-constrained optimization setting, particularly the setting in which evaluating feasibility is just as expensive as evaluating the objective. Here we present constrained Bayesian optimization, which places a prior distribution on both the objective and the constraint functions. We evaluate our method on simulated and real data, demonstrating that constrained Bayesian optimization can quickly find optimal and feasible points, even when small feasible regions cause standard methods to fail.
Jacob R. Gardner, Matt J. Kusner, Zhixiang Eddie Xu, Kilian Q. Weinberger, John P. Cunningham
ICML3
2014 Gradient boosted feature selection
abstract
A feature selection algorithm should ideally satisfy four conditions: reliably extract relevant features; be able to identify non-linear feature interactions; scale linearly with the number of features and dimensions; allow the incorporation of known sparsity structure. In this work we propose a novel feature selection algorithm, Gradient Boosted Feature Selection (GBFS), which satisfies all four of these requirements. The algorithm is flexible, scalable, and surprisingly straight-forward to implement as it is based on a modification of Gradient Boosted Trees. We evaluate GBFS on several real world data sets and show that it matches or outperforms other state of the art feature selection algorithms. Yet it scales to larger data set sizes and naturally allows for domain-specific side information.
Zhixiang Eddie Xu, Gao Huang 0001, Kilian Q. Weinberger, Alice X. Zheng
KDD1
2014 Transductive Minimax Probability Machine
Gao Huang 0001, Shiji Song, Zhixiang Eddie Xu, Kilian Q. Weinberger
ECML/PKDD (1)3
2014 Classifier cascades and trees for minimizing feature evaluation cost
Zhixiang Eddie Xu, Matt J. Kusner, Kilian Q. Weinberger, Minmin Chen, Olivier Chapelle
J. Mach. Learn. Res.1
2013 Anytime Representation Learning
abstract
Evaluation cost during test-time is becoming increasingly important as many real-world applications need fast evaluation (e.g. web search engines, email spam filtering) or use expensive features (e.g. medical diagnosis). We introduce Anytime Feature Representations (AFR), a novel algorithm that explicitly addresses this trade-off in the data representation rather than in the classifier. This enables us to turn conventional classifiers, in particular Support Vector Machines, into test-time cost sensitive anytime classifiers - combining the advantages of anytime learning and large-margin classification.
Zhixiang Eddie Xu, Matt J. Kusner, Gao Huang 0001, Kilian Q. Weinberger
ICML (3)1
2013 Cost-Sensitive Tree of Classifiers
abstract
Recently, machine learning algorithms have successfully entered large-scale real-world industrial applications (e.g. search engines and email spam filters). Here, the CPU cost during test-time must be budgeted and accounted for. In this paper, we address the challenge of balancing test-time cost and the classifier accuracy in a principled fashion. The test-time cost of a classifier is often dominated by the computation required for feature extraction-which can vary drastically across features. We incorporate this extraction time by constructing a tree of classifiers, through which test inputs traverse along individual paths. Each path extracts different features and is optimized for a specific sub-partition of the input space. By only computing features for inputs that benefit from them the most, our cost-sensitive tree of classifiers can match the high accuracies of the current state-of-the-art at a small fraction of the computational cost.
Zhixiang Eddie Xu, Matt J. Kusner, Kilian Q. Weinberger, Minmin Chen
ICML (1)1
2012 From sBoW to dCoT marginalized encoders for text representation
abstract
In text mining, information retrieval, and machine learning, text documents are commonly represented through variants of sparse Bag of Words (sBoW) vectors (e.g. TF-IDF [1]). Although simple and intuitive, sBoW style representations suffer from their inherent over-sparsity and fail to capture word-level synonymy and polysemy. Especially when labeled data is limited (e.g. in document classification), or the text documents are short (e.g. emails or abstracts), many features are rarely observed within the training corpus. This leads to overfitting and reduced generalization accuracy. In this paper we propose Dense Cohort of Terms (dCoT), an unsupervised algorithm to learn improved sBoW document features. dCoT explicitly models absent words by removing and reconstructing random sub-sets of words in the unlabeled corpus. With this approach, dCoT learns to reconstruct frequent words from co-occurring infrequent words and maps the high dimensional sparse sBoW vectors into a low-dimensional dense representation. We show that the feature removal can be marginalized out and that the reconstruction can be solved for in closed-form. We demonstrate empirically, on several benchmark datasets, that dCoT features significantly improve the classification accuracy across several document classification tasks.
Zhixiang Eddie Xu, Minmin Chen, Kilian Q. Weinberger, Fei Sha
CIKM1
2012 Marginalized Denoising Autoencoders for Domain Adaptation
Minmin Chen, Zhixiang Eddie Xu, Kilian Q. Weinberger, Fei Sha
ICML2
2012 The Greedy Miser: Learning under Test-time Budgets
Zhixiang Eddie Xu, Kilian Q. Weinberger, Olivier Chapelle
ICML1