EDBT 2026 Demo / reviewers in the wild / expert
Ji Zhu 0001
dblp:00/4868-1
· DBLP profile ↗
11ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Probabilistic and Bayesian machine learning · 39% Graph learning · 35% Learning theory · 22% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 21 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
1.1 | 2 | 2023 | Community models for networks observed through edge nominations · J. Mach. Learn. Res. 2023 High-dimensional Gaussian graphical models on network-linked data · J. Mach. Learn. Res. 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › relational model
statistical network models |
0.9 | 1 | 2025 | Latent Process Models for Functional Network Data · J. Mach. Learn. Res. 2025 |
Machine learning › Graph learning › heterogeneous graph
bipartite graph |
0.8 | 1 | 2024 | Variational Estimators of the Degree-corrected Latent Block Model for Bipartite Networks · J. Mach. Learn. Res. 2024 |
Machine learning › Graph learning
stochastic block model |
0.7 | 1 | 2023 | Community models for networks observed through edge nominations · J. Mach. Learn. Res. 2023 |
Data mining › structured data mining › graph mining
community detection |
0.7 | 1 | 2023 | Community models for networks observed through edge nominations · J. Mach. Learn. Res. 2023 |
Data mining
network analysis |
0.7 | 1 | 2023 | Community models for networks observed through edge nominations · J. Mach. Learn. Res. 2023 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
gaussian graphical model |
0.4 | 1 | 2020 | High-dimensional Gaussian graphical models on network-linked data · J. Mach. Learn. Res. 2020 |
Machine learning › Learning theory › high-dimensional statistics
high-dimensional estimation |
0.4 | 1 | 2020 | High-dimensional Gaussian graphical models on network-linked data · J. Mach. Learn. Res. 2020 |
Mathematical optimization
gradient descent |
0.3 | 1 | 2025 | Latent Process Models for Functional Network Data · J. Mach. Learn. Res. 2025 |
Data mining › sampling
network sampling |
0.2 | 1 | 2023 | Community models for networks observed through edge nominations · J. Mach. Learn. Res. 2023 |
Machine learning › Learning theory › statistical learning theory › regularization theory
regularization path |
0.1 | 2 | 2004 | The Entire Regularization Path for the Support Vector Machine · J. Mach. Learn. Res. 2004 The Entire Regularization Path for the Support Vector Machine · NIPS 2004 |
Machine learning › Kernel, tree and ensemble methods
support vector machine |
0.1 | 2 | 2004 | The Entire Regularization Path for the Support Vector Machine · NIPS 2004 1-norm Support Vector Machines · NIPS 2003 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting |
0.0 | 1 | 2004 | Boosting as a Regularized Path to a Maximum Margin Classifier · J. Mach. Learn. Res. 2004 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
method of moments |
0.0 | 1 | 2004 | A Method for Inferring Label Sampling Mechanisms in Semi-Supervised Learning · NIPS 2004 |
Machine learning › Kernel, tree and ensemble methods › support vector machine
1-norm SVM |
0.0 | 1 | 2003 | 1-norm Support Vector Machines · NIPS 2003 |
Machine learning › Learning theory › generalization
generalization analysis |
0.0 | 1 | 2003 | Margin Maximizing Loss Functions · NIPS 2003 |
Machine learning › Learning theory
margin maximization |
0.0 | 1 | 2003 | Margin Maximizing Loss Functions · NIPS 2003 |
Machine learning › Optimization for machine learning
solution path algorithm |
0.0 | 1 | 2003 | 1-norm Support Vector Machines · NIPS 2003 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › generalized linear model › logistic regression
kernel logistic regression |
0.0 | 1 | 2001 | Kernel Logistic Regression and the Import Vector Machine · NIPS 2001 |
Machine learning › Learning theory › classification
multiclass classification |
0.0 | 1 | 2001 | Kernel Logistic Regression and the Import Vector Machine · NIPS 2001 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection |
0.0 | 1 | 2003 | 1-norm Support Vector Machines · NIPS 2003 |
Methods — techniques the papers use, named apart from their topics
gradient descent · 1.7functional basis representation · 1.7method of moments · 1.4spectral clustering · 1.3variational expectation-maximization · 0.8maximum likelihood estimation · 0.4inverse covariance estimation · 0.4support vector machine · 0.0regularization path computation · 0.0regularization path algorithm · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Latent Process Models for Functional Network DataabstractNetwork data are often sampled with auxiliary information or collected through the observation of a complex system over time, leading to multiple network snapshots indexed by a continuous variable. Many methods in statistical network analysis are traditionally designed for a single network, and can be applied to an aggregated network in this setting, but that approach can miss important functional structure. Here we develop an approach to estimating the expected network explicitly as a function of a continuous index, be it time or another indexing variable. We parameterize the network expectation through low dimensional latent processes, whose components we represent with a fixed, finite-dimensional functional basis. We derive a gradient descent estimation algorithm, establish theoretical guarantees for recovery of the low dimensional structure, compare our method to competitors, and apply it to a data set of international political interactions over time, showing our proposed method to adapt well to data, outperform competitors, and provide interpretable and meaningful results. Peter W. MacDonald, Elizaveta Levina, Ji Zhu 0001 |
J. Mach. Learn. Res. | 3 |
| 2024 | Variational Estimators of the Degree-corrected Latent Block Model for Bipartite NetworksabstractBipartite graphs are ubiquitous across various scientific and engineering fields. Simultaneously grouping the two types of nodes in a bipartite graph via biclustering represents a fundamental challenge in network analysis for such graphs. The latent block model (LBM) is a commonly used model-based tool for biclustering. However, the effectiveness of the LBM is often limited by the influence of row and column sums in the data matrix. To address this limitation, we introduce the degree-corrected latent block model (DC-LBM), which accounts for the varying degrees in row and column clusters, significantly enhancing performance on real-world data sets and simulated data. We develop an efficient variational expectation-maximization algorithm by creating closed-form solutions for parameter estimates in the M steps. Furthermore, we prove the label consistency and the rate of convergence of the variational estimator under the DC-LBM, allowing the expected graph density to approach zero as long as the average expected degrees of rows and columns approach infinity when the size of the graph increases. Ji Zhu 0001 |
J. Mach. Learn. Res. | 3 |
| 2023 | Community models for networks observed through edge nominationsabstractCommunities are a common and widely studied structure in networks, typically assuming that the network is fully and correctly observed. In practice, network data are often collected by querying nodes about their connections. In some settings, all edges of a sampled node will be recorded, and in others, a node may be asked to name its connections. These sampling mechanisms introduce noise and bias, which can obscure the community structure and invalidate assumptions underlying standard community detection methods. We propose a general model for a class of network sampling mechanisms based on recording edges via querying nodes, designed to improve community detection for network data collected in this fashion. We model edge sampling probabilities as a function of both individual preferences and community parameters, and show community detection can be performed by spectral clustering under this general class of models. We also propose, as a special case of the general framework, a parametric model for directed networks we call the nomination stochastic block model, which allows for meaningful parameter interpretations and can be fitted by the method of moments. In this case, spectral clustering and the method of moments are computationally efficient and come with theoretical guarantees of consistency. We evaluate the proposed model in simulation studies on unweighted and weighted networks and under misspecified models. The method is applied to a faculty hiring dataset, discovering a meaningful hierarchy of communities among US business schools. Tianxi Li, Elizaveta Levina, Ji Zhu 0001 |
J. Mach. Learn. Res. | 3 |
| 2020 | High-dimensional Gaussian graphical models on network-linked dataabstractGraphical models are commonly used to represent conditional dependence relationships between variables. There are multiple methods available for exploring them from high-dimensional data, but almost all of them rely on the assumption that the observations are independent and identically distributed. At the same time, observations connected by a network are becoming increasingly common, and tend to violate these assumptions. Here we develop a Gaussian graphical model for observations connected by a network with potentially different mean vectors, varying smoothly over the network. We propose an efficient estimation algorithm and demonstrate its effectiveness on both simulated and real data, obtaining meaningful and interpretable results on a statistics coauthorship network. We also prove that our method estimates both the inverse covariance matrix and the corresponding graph structure correctly under the assumption of network “cohesion”, which refers to the empirically observed phenomenon of network neighbors sharing similar traits. Tianxi Li, Elizaveta Levina, Ji Zhu 0001 |
J. Mach. Learn. Res. | 4 |
| 2004 | The Entire Regularization Path for the Support Vector MachineabstractIn this paper we argue that the choice of the SVM cost parameter can be critical. We then derive an algorithm that can fit the entire path of SVM solutions for every value of the cost parameter, with essentially the same computational cost as fitting one SVM model. Trevor J. Hastie, Saharon Rosset, Robert Tibshirani, Ji Zhu 0001 |
NIPS | 4 |
| 2004 | A Method for Inferring Label Sampling Mechanisms in Semi-Supervised LearningabstractWe consider the situation in semi-supervised learning, where the "label sampling" mechanism stochastically depends on the true response (as well as potentially on the features). We suggest a method of moments for estimating this stochastic dependence using the unlabeled data. This is potentially useful for two distinct purposes: a. As an input to a super- vised learning procedure which can be used to "de-bias" its results using labeled data only and b. As a potentially interesting learning task in it- self. We present several examples to illustrate the practical usefulness of our method. 1 Introduction In semi-supervised learning, we assume we have a sample (xi, yi, si)n i=1, of i.i.d. draws from a joint distribution on (X, Y, S), where:1 xi Rp are p-vectors of features. yi is a label, or response (yi R for regression, yi {0, 1} for 2-class classifica- tion). si {0, 1} is a "labeling indicator", that is yi is observed if and only if si = 1, while xi is observed for all i. In this paper we consider the interesting case of semi-supervised learning, where the prob- ability of observing the response depends on the data through the true response, as well as 1Our notation here differs somewhat from many semi-supervised learning papers, where the un- labeled part of the sample is separated from the labeled part and sometimes called "test set". potentially through the features. Our goal is to model this unknown dependence: l(x, y) = P r(S = 1|x, y) (1) Note that the dependence on y (which is unobserved when S = 0) prevents us from using standard supervised modeling approaches to learn l. We show here that we can use the whole data-set (labeled+unlabeled data) to obtain estimates of this probability distribution within a parametric family of distributions, without needing to "impute" the unobserved responses.2 We believe this setup is of significant practical interest. Here are a couple of examples of realistic situations: 1. The problem of learning from positive examples and unlabeled data is of significant interest in document topic learning [4, 6, 8]. Consider a generalization of that problem, where we observe a sample of positive and negative examples and unlabeled data, but we believe that the positive and negative labels are supplied with different probabilities (in the document learning example, positive examples are typically more likely to be labeled than negative ones, which are much more abundant). These probabilities may also not be uniform within each class, and depend on the features as well. Our methods allow us to infer these labeling probabilities by utilizing the unlabeled data. 2. Consider a satisfaction survey, where clients of a company are requested to report their level of satisfaction, but they can choose whether or not they do so. It is reasonable to assume that their willingness to report their satisfaction depends on their actual satisfaction level. Using our methods, we can infer the dependence of the reporting probability on the actual satisfaction by utilizing the unlabeled data, i.e., the customers who declined to respond. Being able to infer the labeling mechanism is important for two distinct reasons. First, it may be useful for "de-biasing" the results of supervised learning, which uses only the labeled examples. The generic approach for achieving this is to use "inverse sampling" weights (i.e. weigh labeled examples by 1/l(x, y)). The us of this for maximum likeli- hood estimation is well established in the literature as a method for correcting sampling bias (of which semi-supervised learning is an example) [10]. We can also use the learned mechanism to post-adjust the probabilities from a probability estimation methods such as logistic regression to attain "unbiasedness" and consistency [11]. Second, understanding the labeling mechanism may be an interesting and useful learning task in itself. Consider, for example, the "satisfaction survey" scenario described above. Understanding the way in which satisfaction affects the customers' willingness to respond to the survey can be used to get a better picture of overall satisfaction and to design better future surveys, regardless of any supervised learning task which models the actual satisfaction. Our approach is described in section 2, and is based on a method of moments. Observe that for every function of the features g(x), we can get an unbiased estimate of its mean n as 1 g(x n i=1 i). We show that if we know the underlying label sampling mechanism l(x, y) we can get a different unbiased estimate of Eg(x), which uses only the labeled examples, weighted by 1/l(x, y). We suggest inferring the unknown function l(x, y) by requiring that we get identical estimates of Eg(x) using both approaches. We illustrate our method's implementation on the California Housing data-set in section 3. In section 4 we review related work in the machine learning and statistics literature, and we conclude with a discussion in section 5. 2The importance of this is that we are required to hypothesize and fit a conditional probability model for l(x, y) only, as opposed to the full probability model for (S, X, Y ) required for, say, EM. Saharon Rosset, Ji Zhu 0001, Trevor J. Hastie |
NIPS | 2 |
| 2004 | The Entire Regularization Path for the Support Vector Machine
Trevor J. Hastie, Saharon Rosset, Robert Tibshirani, Ji Zhu 0001 |
J. Mach. Learn. Res. | 4 |
| 2004 | Boosting as a Regularized Path to a Maximum Margin Classifier
Saharon Rosset, Ji Zhu 0001, Trevor J. Hastie |
J. Mach. Learn. Res. | 2 |
| 2003 | Margin Maximizing Loss FunctionsabstractMargin maximizing properties play an important role in the analysis of classi£- cation models, such as boosting and support vector machines. Margin maximiza- tion is theoretically interesting because it facilitates generalization error analysis, and practically interesting because it presents a clear geometric interpretation of the models being built. We formulate and prove a suf£cient condition for the solutions of regularized loss functions to converge to margin maximizing separa- tors, as the regularization vanishes. This condition covers the hinge loss of SVM, the exponential loss of AdaBoost and logistic regression loss. We also generalize it to multi-class classi£cation problems, and present margin maximizing multi- class versions of logistic regression and support vector machines. Saharon Rosset, Ji Zhu 0001, Trevor J. Hastie |
NIPS | 2 |
| 2003 | 1-norm Support Vector MachinesabstractThe standard 2-norm SVM is known for its good performance in two- In this paper, we consider the 1-norm SVM. We class classi£cation. argue that the 1-norm SVM may have some advantage over the standard 2-norm SVM, especially when there are redundant noise features. We also propose an ef£cient algorithm that computes the whole solution path of the 1-norm SVM, hence facilitates adaptive selection of the tuning parameter for the 1-norm SVM. Ji Zhu 0001, Saharon Rosset, Trevor J. Hastie, Robert Tibshirani |
NIPS | 1 |
| 2001 | Kernel Logistic Regression and the Import Vector MachineabstractThe support vector machine (SVM) is known for its good performance in binary classification, but its extension to multi-class classification is still an on-going research issue. In this paper, we propose a new approach for classification, called the import vector machine (IVM), which is built on kernel logistic regression (KLR). We show that the IVM not only per- forms as well as the SVM in binary classification, but also can naturally be generalized to the multi-class case. Furthermore, the IVM provides an estimate of the underlying probability. Similar to the “support points” of the SVM, the IVM model uses only a fraction of the training data to index kernel basis functions, typically a much smaller fraction than the SVM. This gives the IVM a computational advantage over the SVM, especially when the size of the training data set is large. is qualitative and assumes values in a finite set , where the output from . We , to it. Usually it is assumed that the training data are an independently and identically distributed sample from an unknown probability distribution Ji Zhu 0001, Trevor J. Hastie |
NIPS | 1 |