Yue Zhu 0001

dblp:20/4945-1 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
0since 2021 · last 2019
0000-0002-2615-1291ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-authorDatabases, data management, data science and information retrieval · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 74% Machine learning and data management · 26%
Artificial intelligence
3 papers
Learning paradigms · 54% Kernel, tree and ensemble methods · 20% Graph learning · 20%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › predictive modeling › classification
multi-label classification
0.622018
Multi-Label Learning with Emerging New Labels · IEEE Trans. Knowl. Data Eng. 2018
Multi-label Learning with Emerging New Labels · ICDM 2016
Data mining › predictive modeling
classification
0.522017
New Class Adaptation Via Instance Generation in One-Pass Class Incremental Learning · ICDM 2017
Multi-label Learning with Emerging New Labels · ICDM 2016
Machine learning and data management
open-world learning
0.522017
New Class Adaptation Via Instance Generation in One-Pass Class Incremental Learning · ICDM 2017
Multi-label Learning with Emerging New Labels · ICDM 2016
Machine learning and data management › continual learning
incremental learning
0.422017
New Class Adaptation Via Instance Generation in One-Pass Class Incremental Learning · ICDM 2017
Multi-label Learning with Emerging New Labels · ICDM 2016
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.312018
Isolation Kernel and Its Effect on SVM · KDD 2018
Machine learning › Graph learning
label correlation
0.312018
Multi-Label Learning with Global and Local Label Correlation · IEEE Trans. Knowl. Data Eng. 2018
Machine learning › Learning paradigms
multi-label classification
0.312018
Multi-Label Learning with Global and Local Label Correlation · IEEE Trans. Knowl. Data Eng. 2018
Machine learning › Learning paradigms › multiple instance learning
multi-instance multi-label learning
0.312017
Discover Multiple Novel Labels in Multi-Instance Multi-Label Learning · AAAI 2017
Machine learning › Learning paradigms
weakly supervised learning
0.312017
Discover Multiple Novel Labels in Multi-Instance Multi-Label Learning · AAAI 2017
Data mining › predictive modeling › classification › incremental classification
class-incremental learning
0.312017
New Class Adaptation Via Instance Generation in One-Pass Class Incremental Learning · ICDM 2017
Data mining
anomaly detection
0.212016
Overcoming Key Weaknesses of Distance-based Neighbourhood Methods using a Data Dependent Dissimilarity Measure · KDD 2016
Data mining
clustering
0.212016
Overcoming Key Weaknesses of Distance-based Neighbourhood Methods using a Data Dependent Dissimilarity Measure · KDD 2016
Data mining › clustering
distance-based clustering
0.212016
Overcoming Key Weaknesses of Distance-based Neighbourhood Methods using a Data Dependent Dissimilarity Measure · KDD 2016
Data mining › anomaly detection › outlier detection
distance-based outlier detection
0.212016
Overcoming Key Weaknesses of Distance-based Neighbourhood Methods using a Data Dependent Dissimilarity Measure · KDD 2016
Data mining › predictive modeling › classification › incremental classification
novel class detection
0.212016
Multi-label Learning with Emerging New Labels · ICDM 2016
Data mining
data stream mining
0.112018
Multi-Label Learning with Emerging New Labels · IEEE Trans. Knowl. Data Eng. 2018
Machine learning and data management › online learning
one-pass learning
0.112017
New Class Adaptation Via Instance Generation in One-Pass Class Incremental Learning · ICDM 2017
Mathematical optimization › constrained optimization
augmented lagrangian method
0.112017
Discover Multiple Novel Labels in Multi-Instance Multi-Label Learning · AAAI 2017

Methods — techniques the papers use, named apart from their topics

clustering regularization · 0.6augmented lagrangian optimization · 0.6random forest kernel · 0.3latent representation learning · 0.3laplacian kernel · 0.3label manifold learning · 0.3isolation kernel · 0.3dimensionality reduction · 0.3pseudo instance optimization · 0.3instance generation · 0.3new label detection · 0.2data-dependent dissimilarity measure · 0.2classifier construction · 0.2
YearPublicationVenuePosition
2019 Lowest probability mass neighbour algorithms: relaxing the metric constraint in distance-based neighbourhood algorithms
Kai Ming Ting, Ye Zhu 0002, Mark J. Carman, Yue Zhu 0001, Takashi Washio, Zhi-Hua Zhou
Mach. Learn.4
2018 Isolation Kernel and Its Effect on SVM
abstract
This paper investigates data dependent kernels that are derived directly from data. This has been an outstanding issue for about two decades which hampered the development of kernel-based methods. We introduce Isolation Kernel which is solely dependent on data distribution, requiring neither class information nor explicit learning to be a classifier. In contrast, existing data dependent kernels rely heavily on class information and explicit learning to produce a classifier. We show that Isolation Kernel approximates well to a data independent kernel function called Laplacian kernel under uniform density distribution. With this revelation, Isolation Kernel can be viewed as a data dependent kernel that adapts a data independent kernel to the structure of a dataset. We also provide a reason why the proposed new data dependent kernel enables SVM (which employs a kernel through other means) to improve its predictive accuracy. The key differences between Random Forest kernel and Isolation Kernel are discussed to examine the reasons why the latter is a more successful tree-based kernel.
Kai Ming Ting, Yue Zhu 0001, Zhi-Hua Zhou
KDD2
2018 Multi-Label Learning with Global and Local Label Correlation
abstract
It is well-known that exploiting label correlations is important to multi-label learning. Existing approaches either assume that the label correlations are global and shared by all instances; or that the label correlations are local and shared only by a data subset. In fact, in the real-world applications, both cases may occur that some label correlations are globally applicable and some are shared only in a local group of instances. Moreover, it is also a usual case that only partial labels are observed, which makes the exploitation of the label correlations much more difficult. That is, it is hard to estimate the label correlations when many labels are absent. In this paper, we propose a new multi-label approach GLOCAL dealing with both the full-label and the missing-label cases, exploiting global and local label correlations simultaneously, through learning a latent label representation and optimizing label manifolds. The extensive experimental studies validate the effectiveness of our approach on both full-label and missing-label data.
Yue Zhu 0001, James T. Kwok, Zhi-Hua Zhou
IEEE Trans. Knowl. Data Eng.1
2018 Multi-Label Learning with Emerging New Labels
abstract
In a multi-label learning task, an object possesses multiple concepts where each concept is represented by a class label. Previous studies on multi-label learning have focused on a fixed set of class labels, i.e., the class label set of test data is the same as that in the training set. In many applications, however, the environment is dynamic and new concepts may emerge in a data stream. In order to maintain a good predictive performance in this environment, a multi-label learning method must have the ability to detect and classify instances with emerging new labels. To this end, we propose a new approach called Multi-label learning with Emerging New Labels (MuENL). It has three functions: classify instances on currently known labels, detect the emergence of a new label, and construct a new classifier for each new label that works collaboratively with the classifier for known labels. In addition, we show that MuENL can be easily extended to handle sparse high dimensional data streams by simply reducing the original dimensionality, and then applying MuENL on the reduced dimensional space. Our empirical evaluation shows the effectiveness of MuENL on several benchmark datasets and MuENLHD on the sparse high dimensional Weibo dataset.
Yue Zhu 0001, Kai Ming Ting, Zhi-Hua Zhou
IEEE Trans. Knowl. Data Eng.1
2017 Discover Multiple Novel Labels in Multi-Instance Multi-Label Learning
abstract
Multi-instance multi-label learning (MIML) is a learning paradigm where an object is represented by a bag of instances and each bag is associated with multiple labels. Ordinary MIML setting assumes a fixed target label set. In real applications, multiple novel labels may exist outside this set, but hidden in the training data and unknown to the MIML learner. Existing MIML approaches are unable to discover the hidden novel labels, let alone predicting these labels in the previously unseen test data. In this paper, we propose the first approach to discover multiple novel labels in MIML problem using an efficient augmented lagrangian optimization, which has a bag-dependent loss term and a bag-independent clustering regularization term, enabling the known labels and multiple novel labels to be modeled simultaneously. The effectiveness of the proposed approach is validated in experiments.
Yue Zhu 0001, Kai Ming Ting, Zhi-Hua Zhou
AAAI1
2017 New Class Adaptation Via Instance Generation in One-Pass Class Incremental Learning
abstract
One pass learning updates a model with only a single scan of the dataset, without storing historical data. Previous studies focus on classification tasks with a fixed class set, and will perform poorly in an open dynamic environment when new classes emerge in a data stream. The performance degrades because the classifier needs to receive a sufficient number of instances from new classes to establish a good model. This can take a long period of time. In order to reduce this period to deal with any-time prediction task, we introduce a framework to handle emerging new classes called One-Pass Class Incremental Learning (OPCIL). The central issue in OPCIL is: how to effectively adapt a classifier of existing classes to incorporate emerging new classes. We call it the new class adaptation issue, and propose a new approach to address it, which requires only one new class instance. The key is to generate pseudo instances which are optimized to satisfy properties that produce a good discriminative classifier. We provide the necessary propertiesand optimization procedures required to address this issue. Experiments validate the effectiveness of this approach.
Yue Zhu 0001, Kai Ming Ting, Zhi-Hua Zhou
ICDM1
2016 Multi-label Learning with Emerging New Labels
abstract
Multi-label learning is widely applied in many tasks, where an object possesses multiple concepts with each represented by a class label. Previous studies on multi-label learning have focused on a fixed set of class labels, i.e., the class label set of test data is the same as that in the training set. In many applications, however, the environment is open and new concepts may emerge with previously unseen instances. In order to maintain good predictive performance in this environment, a multi-label learning method must have the ability to detect and classify those instances with emerging new labels. To this end, we propose a new approach called Multi-label learning with Emerging New Labels (MuENL). It builds models with three functions: classify instances on currently known labels, detect the emergence of a new label in new instances, and construct a new classifier for each new label that works collaboratively with the classifier for known labels. Our empirical evaluation shows the effectiveness of MuENL.
Yue Zhu 0001, Kai Ming Ting, Zhi-Hua Zhou
ICDM1
2016 Overcoming Key Weaknesses of Distance-based Neighbourhood Methods using a Data Dependent Dissimilarity Measure
abstract
This paper introduces the first generic version of data dependent dissimilarity and shows that it provides a better closest match than distance measures for three existing algorithms in clustering, anomaly detection and multi-label classification. For each algorithm, we show that by simply replacing the distance measure with the data dependent dissimilarity measure, it overcomes a key weakness of the otherwise unchanged algorithm.
Kai Ming Ting, Ye Zhu 0002, Mark J. Carman, Yue Zhu 0001, Zhi-Hua Zhou
KDD4
2015 One-Pass Multi-View Learning
Yue Zhu 0001, Wei Gao 0008, Zhi-Hua Zhou
ACML1
2014 Learning with Augmented Multi-Instance View
Yue Zhu 0001, Jianxin Wu 0001, Yuan Jiang 0001, Zhi-Hua Zhou
ACML1