Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shipeng Yu

dblp:y/ShipengYu · DBLP profile ↗
← Back
49ranked-venue papers
16as first author
0since 2021 · last 2019
0000-0002-0262-4031ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 10 first-authorDatabases, data management, data science and information retrieval · 23 · 8 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
17 papers
Probabilistic and Bayesian machine learning · 20% Trustworthy machine learning · 18% Learning paradigms · 17%
Databases, data mining, and information retrieval
16 papers
Data mining · 28% Recommender systems · 27% Web and social media mining · 25%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Medical and health informatics · 58% Computational social science and digital humanities · 25% Bioinformatics and computational biology · 17%
Theoretical computer science
4 papers
Mathematical optimization · 80% Algorithms and data structures · 15% Graph algorithms and graph theory · 5%

Topics — the 30 heaviest of 77, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning paradigms
multi-task learning
0.532016
Multiplicative Multitask Feature Learning · J. Mach. Learn. Res. 2016
On Multiplicative Multitask Feature Learning · NIPS 2014
Robust multi-task learning with t-processes · ICML 2007
Web and social media mining
user engagement
0.522019
A State Transition Model for Mobile Notifications via Survival Analysis · WSDM 2019
Near Real-time Optimization of Activity-based Notifications · KDD 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
survival analysis
0.412019
A State Transition Model for Mobile Notifications via Survival Analysis · WSDM 2019
Recommender systems
content recommendation
0.412019
Internal Promotion Optimization · KDD 2019
Mathematical optimization
constrained optimization
0.412019
Internal Promotion Optimization · KDD 2019
Computational social science and digital humanities › social computing
crowdsourcing
0.322012
Eliminating Spammers and Ranking Annotators for Crowdsourced Labeling Tasks · J. Mach. Learn. Res. 2012
Learning From Crowds · J. Mach. Learn. Res. 2010
Machine learning › Optimization for machine learning › coordinate descent
block coordinate descent
0.212016
Multiplicative Multitask Feature Learning · J. Mach. Learn. Res. 2016
Machine learning › Graph learning
heterogeneous graph learning
0.212016
Identifying Decision Makers from Professional Social Networks · KDD 2016
Machine learning › Deep learning architectures and training
regularization
0.212016
Multiplicative Multitask Feature Learning · J. Mach. Learn. Res. 2016
Medical and health informatics
clinical data analysis
0.212016
Going Digital: A Survey on Digitalization and Large-Scale Data Analytics in Healthcare · Proc. IEEE 2016
Web and social media mining › social network analysis
professional social network analysis
0.212016
Identifying Decision Makers from Professional Social Networks · KDD 2016
Machine learning › Trustworthy machine learning
crowdsourced annotation
0.222011
Ranking annotators for crowdsourced labeling tasks · NIPS 2011
Learning From Crowds · J. Mach. Learn. Res. 2010
Data mining
dimensionality reduction
0.242006
Multi-Output Regularized Feature Projection · IEEE Trans. Knowl. Data Eng. 2006
Supervised probabilistic principal component analysis · KDD 2006
Multi-label informed latent semantic indexing · SIGIR 2005
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.222010
Learning From Crowds · J. Mach. Learn. Res. 2010
Supervised learning from multiple experts: whom to trust when everyone lies a bit · ICML 2009
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.232007
Robust multi-task learning with t-processes · ICML 2007
Stochastic Relational Models for Discriminative Link Prediction · NIPS 2006
Collaborative ordinal regression · ICML 2006
Data mining › predictive modeling › classification
multi-label classification
0.232008
Extracting shared subspace for multi-label classification · KDD 2008
Multi-label informed latent semantic indexing · SIGIR 2005
Multi-Output Regularized Projection · CVPR (2) 2005
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.122007
Local learning projections · ICML 2007
Multi-Output Regularized Projection · CVPR (2) 2005
Data mining › predictive modeling
classification
0.122010
Extracting shared subspace for multi-label classification · KDD 2008
Designing efficient cascaded classifiers: tradeoff between accuracy and cost · KDD 2010
Information retrieval › online advertising
ad targeting
0.112019
Internal Promotion Optimization · KDD 2019
Information retrieval
online advertising
0.112019
Internal Promotion Optimization · KDD 2019
Ubiquitous computing and smart environments › mobile computing
mobile notification
0.112019
A State Transition Model for Mobile Notifications via Survival Analysis · WSDM 2019
Computer vision › Image recognition and object detection › object detection
cascade classifier
0.112010
Designing efficient cascaded classifiers: tradeoff between accuracy and cost · KDD 2010
Information retrieval › retrieval models › latent semantic models
latent semantic indexing
0.122005
Multi-label informed latent semantic indexing · SIGIR 2005
Hierarchy-Regularized Latent Semantic Indexing · ICDM 2005
Machine learning › Trustworthy machine learning › Data-centric AI
annotator modeling
0.112009
Supervised learning from multiple experts: whom to trust when everyone lies a bit · ICML 2009
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › discriminant analysis
kernel discriminant analysis
0.112009
Multiclass Probabilistic Kernel Discriminant Analysis · IJCAI 2009
Medical and health informatics › medical imaging
medical image analysis
0.112009
Hierarchical learning for tubular structure parsing in medical imaging: A study on coronary arteries using 3D CT Angiography · ICCV 2009
Web and social media mining
social tagging
0.112009
Hierarchical Bayesian Models for Collaborative Tagging Systems · ICDM 2009
Recommender systems
tag recommendation
0.112009
Hierarchical Bayesian Models for Collaborative Tagging Systems · ICDM 2009
Data mining › text mining
topic modeling
0.112009
Hierarchical Bayesian Models for Collaborative Tagging Systems · ICDM 2009
Algorithms and data structures › numerical linear algebra › dimensionality reduction
canonical correlation analysis
0.112009
On the Equivalence between Canonical Correlation Analysis and Orthonormalized Partial Least Squares · IJCAI 2009

Methods — techniques the papers use, named apart from their topics

weibull distribution · 1.1survival analysis · 1.1logistic regression · 1.1constrained optimization · 0.8a/b testing · 0.8graph neural network · 0.5data analytics · 0.5regularization · 0.2blockwise coordinate descent · 0.2generalized eigenvalue decomposition · 0.2cascade optimization · 0.2probabilistic modeling · 0.2multiplicative feature learning · 0.2linear programming · 0.2feature selection · 0.2dimensionality reduction · 0.2hierarchical bayesian model · 0.1annotator scoring · 0.1
YearPublicationVenuePosition
2019 Internal Promotion Optimization
abstract
Most large Internet companies run internal promotions to cross-promote their different products and/or to educate members on how to obtain additional value from the products that they already use. This in turn drives engagement and/or revenue for the company. However, since these internal promotions can distract a member away from the product or page where these are shown, there is a non-zero cannibalization loss incurred for showing these internal promotions. This loss has to be carefully weighed against the gain from showing internal promotions. This can be a complex problem if different internal promotions optimize for different objectives. In that case, it is difficult to compare not just the gain from a conversion through an internal promotion against the loss incurred for showing that internal promotion, but also the gains from conversions through different internal promotions. Hence, we need a principled approach for deciding which internal promotion (if any) to serve to a member in each opportunity to serve an internal promotion. This approach should optimize not just for the net gain to the company, but also for the member's experience. In this paper, we discuss our approach for optimization of internal promotions at LinkedIn. In particular, we present a cost-benefit analysis of showing internal promotions, our formulation of internal promotion optimization as a constrained optimization problem, the architecture of the system for solving the optimization problem and serving internal promotions in real-time, and experimental results from online A/B tests.
Rupesh Gupta, Guangde Chen, Shipeng Yu
KDD3
2019 A State Transition Model for Mobile Notifications via Survival Analysis
abstract
Mobile notifications have become a major communication channel for social networking services to keep users informed and engaged. As more mobile applications push notifications to users, they constantly face decisions on what to send, when and how. A lack of research and methodology commonly leads to heuristic decision making. Many notifications arrive at an inappropriate moment or introduce too many interruptions, failing to provide value to users and spurring users' complaints. In this paper we explore unique features of interactions between mobile notifications and user engagement. We propose a state transition framework to quantitatively evaluate the effectiveness of notifications. Within this framework, we develop a survival model for badging notifications assuming a log-linear structure and a Weibull distribution. Our results show that this model achieves more flexibility for applications and superior prediction accuracy than a logistic regression model. In particular, we provide an online use case on notification delivery time optimization to show how we make better decisions, drive more user engagement, and provide more value to users.
Shaunak Chatterjee, Shipeng Yu, Rómer Rosales
WSDM4
2018 Near Real-time Optimization of Activity-based Notifications
abstract
In recent years, social media applications (e.g., Facebook, LinkedIn) have created mobile applications (apps) to give their members instant and real-time access from anywhere. To keep members informed and drive timely engagement, these mobile apps send event notifications. However, sending notifications for every possible event would result in too many notifications which would in turn annoy members and create a poor member experience.
Viral Gupta, Jinyun Yan, Changji Shi, Zhongen Tao, P. J. Xiao, Curtis Wang, Shipeng Yu, Rómer Rosales, Ajith Muralidharan, Shaunak Chatterjee
KDD8
2016 Identifying Decision Makers from Professional Social Networks
abstract
Sales professionals help organizations win clients for products and services. Generating new clients starts with identifying the right decision makers at the target organization. For the past decade, online professional networks have collected tremendous amount of data on people's identity, their network and behavior data of buyers and sellers building relationships with each other for a variety of use-cases. Sales professionals are increasingly relying on these networks to research, identify and reach out to potential prospects, but it is often hard to find the right people effectively and efficiently. In this paper we present LDMS, the LinkedIn Decision Maker Score, to quantify the ability of making a sales decision for each of the 400M+ LinkedIn members. It is the key data-driven technology underlying Sales Navigator, a proprietary LinkedIn product that is designed for sales professionals. We will specifically discuss the modeling challenges of LDMS, and present two graph-based approaches to tackle this problem by leveraging the professional network data at LinkedIn. Both approaches are able to leverage both the graph information and the contextual information on the vertices, deal with small amount of labels on the graph, and handle heterogeneous graphs among different types of vertices. We will show some offline evaluations of LDMS on historical data, and also discuss its online usage in multiple applications in live production systems as well as future use cases within the LinkedIn ecosystem.
Shipeng Yu, Evangelia Christakopoulou
KDD1
2016 Multiplicative Multitask Feature Learning
abstract
We investigate a general framework of multiplicative multitask feature learning which decomposes individual task's model parameters into a multiplication of two components. One of the components is used across all tasks and the other component is task-specific. Several previous methods can be proved to be special cases of our framework. We study the theoretical properties of this framework when different regularization conditions are applied to the two decomposed components. We prove that this framework is mathematically equivalent to the widely used multitask feature learning methods that are based on a joint regularization of all model parameters, but with a more general form of regularizers. Further, an analytical formula is derived for the across-task component as related to the task- specific component for all these regularizers, leading to a better understanding of the shrinkage effects of different regularizers. Study of this framework motivates new multitask learning algorithms. We propose two new learning formulations by varying the parameters in the proposed framework. An efficient blockwise coordinate descent algorithm is developed suitable for solving the entire family of formulations with rigorous convergence analysis. Simulation studies have identified the statistical properties of data that would be in favor of the new formulations. Extensive empirical studies on various classification and regression benchmark data sets have revealed the relative advantages of the two new formulations by comparing with the state of the art, which provides instructive insights into the feature learning problem with multiple tasks.
Xin Wang 0023, Jinbo Bi, Shipeng Yu, Jiangwen Sun, Minghu Song
J. Mach. Learn. Res.3
2016 Going Digital: A Survey on Digitalization and Large-Scale Data Analytics in Healthcare
abstract
We provide an overview of the recent trends toward digitalization and large-scale data analytics in healthcare. It is expected that these trends are instrumental in the dramatic changes in the way healthcare will be organized in the future. We discuss the recent political initiatives designed to shift care delivery processes from paper to electronic, with the goals of more effective treatments with better outcomes; cost pressure is a major driver of innovation. We describe newly developed networks of healthcare providers, research organizations, and commercial vendors to jointly analyze data for the development of decision support systems. We address the trend toward continuous healthcare where health is monitored by wearable and stationary devices; a related development is that patients increasingly assume responsibility for their own health data. Finally, we discuss recent initiatives toward a personalized medicine, based on advances in molecular medicine, data management, and data analytics.
Volker Tresp, J. Marc Overhage, Markus Bundschus, Shahrooz Rabizadeh, Peter A. Fasching, Shipeng Yu
Proc. IEEE6
2015 Predicting readmission risk with institution-specific prediction models
Shipeng Yu, Faisal Farooq, Alexander Van Esbroeck, Glenn Fung, Vikram Anand, Balaji Krishnapuram
Artif. Intell. Medicine1
2014 Operationalizing patient-generated health data: Home blood pressure monitoring as an example
Shipeng Yu, J. Marc Overhage, Paul C. Tang
AMIA1
2014 On Multiplicative Multitask Feature Learning
Xin Wang 0023, Jinbo Bi, Shipeng Yu, Jiangwen Sun
NIPS3
2012 Building Hospital-Specific Readmission Risk Prediction Models for Heart Failure, Acute Myocardial Infarction and Pneumonia patients
Shipeng Yu, Faisal Farooq, Glenn Fung, Balaji Krishnapuram, Alexander Van Esbroeck, Vikram Anand
AMIA1
2012 Eliminating Spammers and Ranking Annotators for Crowdsourced Labeling Tasks
Vikas C. Raykar, Shipeng Yu
J. Mach. Learn. Res.2
2011 Coarse-to-fine classification via parametric and nonparametric models for computer-aided diagnosis
abstract
Classification is one of the core problems in Computer-Aided Diagnosis (CAD), targeting for early cancer detection using 3D medical imaging interpretation. High detection sensitivity with desirably low false positive (FP) rate is critical for a CAD system to be accepted as a valuable or even indispensable tool in radiologists' workflow. Given various spurious imagery noises which cause observation uncertainties, this remains a very challenging task. In this paper, we propose a novel, two-tiered coarse-to-fine (CTF) classification cascade framework to tackle this problem. We first obtain classification-critical data samples (e.g., implicit samples on the decision boundary) extracted from the holistic data distributions using a robust parametric model (e.g., [13]); then we build a graph-embedding based nonparametric classifier on sampled data, which can more accurately preserve or formulate the complex classification boundary. These two steps can also be considered as effective "sample pruning" and "feature pursuing + kNN/template matching", respectively. Our approach is validated comprehensively in colorectal polyp detection and lung nodule detection CAD systems, as the top two deadly cancers, using hospital scale, multi-site clinical datasets. The results show that our method achieves overall better classification/detection performance than existing state-of-the-art algorithms using single-layer classifiers, such as the support vector machine variants [17], boosting [15], logistic regression [11], relevance vector machine [13], k-nearest neighbor [9] or spectral projections on graph [2].
Meizhu Liu, Le Lu 0001, Xiaojing Ye, Shipeng Yu, Heng Huang 0001
CIKM4
2011 Sparse Classification for Computer Aided Diagnosis Using Learned Dictionaries
Meizhu Liu, Le Lu 0001, Xiaojing Ye, Shipeng Yu, Marcos Salganicoff
MICCAI (3)4
2011 Ranking annotators for crowdsourced labeling tasks
abstract
With the advent of crowdsourcing services it has become quite cheap and reasonably effective to get a dataset labeled by multiple annotators in a short amount of time. Various methods have been proposed to estimate the consensus labels by correcting for the bias of annotators with different kinds of expertise. Often we have low quality annotators or spammers--annotators who assign labels randomly (e.g., without actually looking at the instance). Spammers can make the cost of acquiring labels very expensive and can potentially degrade the quality of the consensus labels. In this paper we formalize the notion of a spammer and define a score which can be used to rank the annotators---with the spammers having a score close to zero and the good annotators having a high score close to one.
Vikas C. Raykar, Shipeng Yu
NIPS2
2011 Matrix-variate and higher-order probabilistic projections
Shipeng Yu, Jinbo Bi, Jieping Ye
Data Min. Knowl. Discov.1
2011 Bayesian Co-Training
Shipeng Yu, Balaji Krishnapuram, Rómer Rosales, R. Bharat Rao
J. Mach. Learn. Res.1
2010 Mixture model label propagation
abstract
Usually, we can use a classification or clustering machine learning algorithm to manage knowledge and information retrieval. If we have a small size of known information with a large scale of unknown data, a semi-supervised learning (SSL) algorithm is often preferred. Under the cluster or manifold assumption, usually, the larger amount of unlabeled data are used for learning, the bigger gains of the SSL approaches are achieved. In the paper, we adopt the graph-based SSL algorithm to solve the problem. However the graph-based SSL algorithms are unable to be learnt with large-scale unlabeled samples and originally can only work in a transductive setting. In the paper, we propose a scalable graph-based SSL algorithm to attack the problems aforementioned by Gaussian mixture model label propagation. Experiments conducted on the real dataset illustrate the effectiveness of the proposed algorithm.
Mingmin Chi, Xisheng He, Shipeng Yu
CIKM3
2010 Designing efficient cascaded classifiers: tradeoff between accuracy and cost
abstract
We propose a method to train a cascade of classifiers by simultaneously optimizing all its stages. The approach relies on the idea of optimizing soft cascades. In particular, instead of optimizing a deterministic hard cascade, we optimize a stochastic soft cascade where each stage accepts or rejects samples according to a probability distribution induced by the previous stage-specific classifier. The overall system accuracy is maximized while explicitly controlling the expected cost for feature acquisition. Experimental results on three clinically relevant problems show the effectiveness of our proposed approach in achieving the desired tradeoff between accuracy and feature acquisition cost.
Vikas C. Raykar, Balaji Krishnapuram, Shipeng Yu
KDD3
2010 Learning From Crowds
Vikas C. Raykar, Shipeng Yu, Linda Zhao 0001, Gerardo Hermosillo, Charles Florin, Luca Bogoni, Linda Moy
J. Mach. Learn. Res.2
2010 A shared-subspace learning framework for multi-label classification
abstract
Multi-label problems arise in various domains such as multi-topic document categorization, protein function prediction, and automatic image annotation. One natural way to deal with such problems is to construct a binary classifier for each label, resulting in a set of independent binary classification problems. Since multiple labels share the same input space, and the semantics conveyed by different labels are usually correlated, it is essential to exploit the correlation information contained in different labels. In this paper, we consider a general framework for extracting shared structures in multi-label classification. In this framework, a common subspace is assumed to be shared among multiple labels. We show that the optimal solution to the proposed formulation can be obtained by solving a generalized eigenvalue problem, though the problem is nonconvex. For high-dimensional problems, direct computation of the solution is expensive, and we develop an efficient algorithm for this case. One appealing feature of the proposed framework is that it includes several well-known algorithms as special cases, thus elucidating their intrinsic relationships. We further show that the proposed framework can be extended to the kernel-induced feature space. We have conducted extensive experiments on multi-topic web page categorization and automatic gene expression pattern image annotation tasks, and results demonstrate the effectiveness of the proposed formulation in comparison with several representative algorithms.
Shuiwang Ji, Lei Tang 0001, Shipeng Yu, Jieping Ye
ACM Trans. Knowl. Discov. Data3
2009 Hierarchical learning for tubular structure parsing in medical imaging: A study on coronary arteries using 3D CT Angiography
abstract
Automatic coronary artery centerline extraction from 3D CT Angiography (CTA) has significant clinical importance for diagnosis of atherosclerotic heart disease. The focus of past literature is dominated by segmenting the complete coronary artery system as trees by computer. Though the labeling of different vessel branches (defined by their medical semantics) is much needed clinically, this task has been performed manually. In this paper, we propose a hierarchical machine learning approach to tackle the problem of tubular structure parsing in medical imaging. It has a progressive three-tiered classification process at volumetric voxel level, vessel segment level, and inter-segment level. Generative models are employed to project from low-level, ambiguous data to class-conditional probabilities; and discriminative classifiers are trained on the upper-level structural patterns of probabilities to label and parse the vessel segments. Our method is validated by experiments of detecting and segmenting clinically defined coronary arteries, from the initial noisy vessel segment networks generated by low-level heuristics-based tracing algorithms. The proposed framework is also generically applicable to other tubular structure parsing tasks.
Le Lu 0001, Jinbo Bi, Shipeng Yu, Zhigang Peng, Arun Krishnan, Xiang Sean Zhou
ICCV3
2009 Hierarchical Bayesian Models for Collaborative Tagging Systems
abstract
Collaborative tagging systems with user generated content have become a fundamental element of websites such as Delicious, Flickr or CiteULike. By sharing common knowledge, massively linked semantic data sets are generated that provide new challenges for data mining. In this paper, we reduce the data complexity in these systems by finding meaningful topics that serve to group similar users and serve to recommend tags or resources to users. We propose a well-founded probabilistic approach that can model every aspect of a collaborative tagging system. By integrating both user information and tag information into the well-known Latent Dirichlet Allocation framework, the developed models can be used to solve a number of important information extraction and retrieval tasks.
Markus Bundschus, Shipeng Yu, Volker Tresp, Achim Rettinger, Mathäus Dejori, Hans-Peter Kriegel
ICDM2
2009 Supervised learning from multiple experts: whom to trust when everyone lies a bit
abstract
We describe a probabilistic approach for supervised learning when we have multiple experts/annotators providing (possibly noisy) labels but no absolute gold standard. The proposed algorithm evaluates the different experts and also gives an estimate of the actual hidden labels. Experimental results indicate that the proposed method is superior to the commonly used majority voting baseline.
Vikas C. Raykar, Shipeng Yu, Linda Zhao 0001, Anna K. Jerebko, Charles Florin, Gerardo Hermosillo, Luca Bogoni, Linda Moy
ICML2
2009 Survival Prediction in Lung Cancer Treated with Radiotherapy: Bayesian Networks vs. Support Vector Machines in Handling Missing Data
abstract
Missing data is a given in the medical domain, so machine learning models should have satisfactory performance even when missing data occurs. Our previous work has focused on support vector machines (SVM), but we hypothesize that Bayesian networks (BN) can handle missing data better. To test the hypothesis, we trained a BN and SVM model for 2 year survival on 322 lung cancer patients and compared their performance in three separate external datasets (35, 47, 33 patients), each with their own characteristics in terms of missing data. The models used tumor size, clinical T and N stage, involved lymph nodes and WHO performance as prognostic features. We found that the BN model performed better than SVM (AUC 0.77, 0.72. 0.70 vs. 0.71, 0.68, 0.69), especially if tumor size was missing. We conclude that BN models are better suited for the medical domain, as they can handle missing data better.
Andre Dekker, Cary Dehing-Oberije, Dirk De Ruysscher, Philippe Lambin, Kartik Komati, Glenn Fung, Shipeng Yu, Andrew Hope, Wilfried De Neve, Yolande Lievens
ICMLA7
2009 On the Equivalence between Canonical Correlation Analysis and Orthonormalized Partial Least Squares
Liang Sun 0001, Shuiwang Ji, Shipeng Yu, Jieping Ye
IJCAI3
2009 Multiclass Probabilistic Kernel Discriminant Analysis
Zheng Zhao 0002, Liang Sun 0001, Shipeng Yu, Huan Liu 0001, Jieping Ye
IJCAI3
2008 Large Scale Diagnostic Code Classification for Medical Patient Records
Lucian Vlad Lita, Shipeng Yu, Radu Stefan Niculescu, Jinbo Bi
IJCNLP2
2008 Extracting shared subspace for multi-label classification
abstract
Multi-label problems arise in various domains such as multi-topic document categorization and protein function prediction. One natural way to deal with such problems is to construct a binary classifier for each label, resulting in a set of independent binary classification problems. Since the multiple labels share the same input space, and the semantics conveyed by different labels are usually correlated, it is essential to exploit the correlation information contained in different labels. In this paper, we consider a general framework for extracting shared structures in multi-label classification. In this framework, a common subspace is assumed to be shared among multiple labels. We show that the optimal solution to the proposed formulation can be obtained by solving a generalized eigenvalue problem, though the problem is non-convex. For high-dimensional problems, direct computation of the solution is expensive, and we develop an efficient algorithm for this case. One appealing feature of the proposed framework is that it includes several well-known algorithms as special cases, thus elucidating their intrinsic relationships. We have conducted extensive experiments on eleven multi-topic web page categorization tasks, and results demonstrate the effectiveness of the proposed formulation in comparison with several representative algorithms.
Shuiwang Ji, Lei Tang 0001, Shipeng Yu, Jieping Ye
KDD3
2008 Privacy-preserving cox regression for survival analysis
abstract
Privacy-preserving data mining (PPDM) is an emergent research area that addresses the incorporation of privacy preserving concerns to data mining techniques. In this paper we propose a privacy-preserving (PP) Cox model for survival analysis, and consider a real clinical setting where the data is horizontally distributed among different institutions. The proposed model is based on linearly projecting the data to a lower dimensional space through an optimal mapping obtained by solving a linear programming problem. Our approach differs from the commonly used random projection approach since it instead finds a projection that is optimal at preserving the properties of the data that are important for the specific problem at hand. Since our proposed approach produces an sparse mapping, it also generates a PP mapping that not only projects the data to a lower dimensional space but it also depends on a smaller subset of the original features (it provides explicit feature selection). Real data from several European healthcare institutions are used to test our model for survival prediction of non-small-cell lung cancer patients. These results are also confirmed using publicly available benchmark datasets. Our experimental results show that we are able to achieve a near-optimal performance without directly sharing the data across different data sources. This model makes it possible to conduct large-scale multi-centric survival analysis without violating privacy-preserving requirements.
Shipeng Yu, Glenn Fung, Rómer Rosales, Sriram Krishnan, R. Bharat Rao, Cary Dehing-Oberije, Philippe Lambin
KDD1
2008 An Improved Multi-task Learning Approach with Applications in Medical Diagnosis
Jinbo Bi, Shipeng Yu, Murat Dundar, R. Bharat Rao
ECML/PKDD (1)3
2007 Local learning projections
abstract
This paper presents a Local Learning Projection (LLP) approach for linear dimensionality reduction. We first point out that the well known Principal Component Analysis (PCA) essentially seeks the projection that has the minimal global estimation error. Then we propose a dimensionality reduction algorithm that leads to the projection with the minimal local estimation error, and elucidate its advantages for classification tasks. We also indicate that LLP keeps the local information in the sense that the projection value of each point can be well estimated based on its neighbors and their projection values. Experimental results are provided to validate the effectiveness of the proposed algorithm.
Mingrui Wu, Kai Yu 0001, Shipeng Yu, Bernhard Schölkopf
ICML3
2007 Robust multi-task learning with t-processes
abstract
Most current multi-task learning frameworks ignore the robustness issue, which means that the presence of "outlier" tasks may greatly reduce overall system performance. We introduce a robust framework for Bayesian multitask learning, t-processes (TP), which are a generalization of Gaussian processes (GP) for multi-task learning. TP allows the system to effectively distinguish good tasks from noisy or outlier tasks. Experiments show that TP not only improves overall system performance, but can also serve as an indicator for the "informativeness" of different tasks.
Shipeng Yu, Volker Tresp, Kai Yu 0001
ICML1
2007 Automatic medical coding of patient records via weighted ridge regression
abstract
In this paper, we apply weighted ridge regression to tackle the highly unbalanced data issue in automatic large-scale ICD-9 coding of medical patient records. Since most of the ICD-9 codes are unevenly represented in the medical records, a weighted scheme is employed to balance positive and negative examples. The weights turn out to be associated with the instance priors from a probabilistic interpretation, and an efficient EM algorithm is developed to automatically update both the weights and the regularization parameter. Experiments on a large-scale real patient database suggest that the weighted ridge regression outperforms the conventional ridge regression and linear support vector machines (SVM).
Jianwu Xu, Shipeng Yu, Jinbo Bi, Lucian Vlad Lita, Radu Stefan Niculescu, R. Bharat Rao
ICMLA2
2007 Bayesian Co-Training
abstract
We propose a Bayesian undirected graphical model for co-training, or more generally for semi-supervised multi-view learning. This makes explicit the previously unstated assumptions of a large class of co-training type algorithms, and also clarifies the circumstances under which these assumptions fail. Building upon new insights from this model, we propose an improved method for co-training, which is a novel co-training kernel for Gaussian process classifiers. The resulting approach is convex and avoids local-maxima problems, unlike some previous multi-view learning methods. Furthermore, it can automatically estimate how much each view should be trusted, and thus accommodate noisy or unreliable views. Experiments on toy data and real world data sets illustrate the benefits of this approach.
Shipeng Yu, Balaji Krishnapuram, Rómer Rosales, Harald Steck, R. Bharat Rao
NIPS1
2006 Variational Bayesian Dirichlet-Multinomial Allocation for Exponential Family Mixtures
Shipeng Yu, Kai Yu 0001, Volker Tresp, Hans-Peter Kriegel
ECML1
2006 Collaborative ordinal regression
abstract
Ordinal regression has become an effective way of learning user preferences, but most research focuses on single regression problems. In this paper we introduce collaborative ordinal regression, where multiple ordinal regression tasks are handled simultaneously. Rather than modeling each task individually, we explore the dependency between ranking functions through a hierarchical Bayesian model and assign a common Gaussian Process (GP) prior to all individual functions. Empirical studies show that our collaborative model outperforms the individual counterpart in preference learning applications.
Shipeng Yu, Kai Yu 0001, Volker Tresp, Hans-Peter Kriegel
ICML1
2006 Supervised probabilistic principal component analysis
abstract
Principal component analysis (PCA) has been extensively applied in data mining, pattern recognition and information retrieval for unsupervised dimensionality reduction. When labels of data are available, e.g., in a classification or regression task, PCA is however not able to use this information. The problem is more interesting if only part of the input data are labeled, i.e., in a semi-supervised setting. In this paper we propose a supervised PCA model called SPPCA and a semi-supervised PCA model called S2PPCA, both of which are extensions of a probabilistic PCA model. The proposed models are able to incorporate the label information into the projection phase, and can naturally handle multiple outputs (i.e., in multi-task learning problems). We derive an efficient EM learning algorithm for both models, and also provide theoretical justifications of the model behaviors. SPPCA and S2PPCA are compared with other supervised projection methods on various learning tasks, and show not only promising performance but also good scalability.
Shipeng Yu, Kai Yu 0001, Volker Tresp, Hans-Peter Kriegel, Mingrui Wu
KDD1
2006 Stochastic Relational Models for Discriminative Link Prediction
abstract
We introduce a Gaussian process (GP) framework, stochastic relational models (SRM), for learning social, physical, and other relational phenomena where interactions between entities are observed. The key idea is to model the stochastic structure of entity relationships (i.e., links) via a tensor interaction of multiple GPs, each defined on one type of entities. These models in fact define a set of nonparametric priors on infinite dimensional tensor matrices, where each element represents a relationship between a tuple of entities. By maximizing the marginalized likelihood, information is exchanged between the participating GPs through the entire relational network, so that the dependency structure of links is messaged to the dependency of entities, reflected by the adapted GP kernels. The framework offers a discriminative approach to link prediction, namely, predicting the existences, strengths, or types of relationships based on the partially observed linkage network as well as the attributes of entities (if given). We discuss properties and variants of SRM and derive an efficient learning algorithm. Very encouraging experimental results are achieved on a toy problem and a user-movie preference link prediction task. In the end we discuss extensions of SRM to general relational learning tasks.
Kai Yu 0001, Shipeng Yu, Volker Tresp, Zhao Xu 0001
NIPS3
2006 Multi-Output Regularized Feature Projection
abstract
Dimensionality reduction by feature projection is widely used in pattern recognition, information retrieval, and statistics. When there are some outputs available (e.g., regression values or classification results), it is often beneficial to consider supervised projection, which is based not only on the inputs, but also on the target values. While this applies to a single-output setting, we are more interested in applications with multiple outputs, where several tasks need to be learned simultaneously. In this paper, we introduce a novel projection approach called Multi-Output Regularized feature Projection (MORP), which preserves the information of input features and, meanwhile, captures the correlations between inputs/outputs and (if applicable) between multiple outputs. This is done by introducing a latent variable model on the joint input-output space and minimizing the reconstruction errors for both inputs and outputs. It turns out that the mappings can be found by solving a generalized eigenvalue problem and are ready to extend to nonlinear mappings. Prediction accuracy can be greatly improved by using the new features since the structure of outputs is explored. We validate our approach in two applications. In the first setting, we predict users' preferences for a set of paintings. The second is concerned with image and text categorization where each image (or document) may belong to multiple categories. The proposed algorithm produces very encouraging results in both settings.
Shipeng Yu, Kai Yu 0001, Volker Tresp, Hans-Peter Kriegel
IEEE Trans. Knowl. Data Eng.1
2005 Multi-Output Regularized Projection
abstract
Dimensionality reduction via feature projection has been widely used in pattern recognition and machine learning. It is often beneficial to derive the projections not only based on the inputs but also on the target values in the training data set. This is of particular importance in predicting multivariate or structured outputs which is an area of growing interest. In this paper we introduce a novel projection framework which is sensitive to both input features and outputs. Based on the derived features prediction accuracy can be greatly improved. We validate our approach in two applications. The first is to model users' preferences on a set of paintings. The second application is concerned with image categorization where each image may belong to multiple categories. The proposed algorithm produces very encouraging results in both settings.
Kai Yu 0001, Shipeng Yu, Volker Tresp
CVPR (2)2
2005 Hierarchy-Regularized Latent Semantic Indexing
abstract
Organizing textual documents into a hierarchical taxonomy is a common practice in knowledge management. Beside textual features, the hierarchical structure of directories reflect additional and important knowledge annotated by experts. It is generally desired to incorporate this information into text mining processes. In this paper, we propose hierarchy-regularized latent semantic indexing, which encodes the hierarchy into a similarity graph of documents and then formulates an optimization problem mapping each document into a low dimensional vector space. The new feature space preserves the intrinsic structure of the original taxonomy and thus provides a meaningful basis for various learning tasks like visualization and classification. Our approach employs the information about class proximity and class specificity, and can naturally cope with multi-labeled documents. Our empirical studies show very encouraging results on two real-world data sets, the new Reuters (RCVI) benchmark and the Swissprot protein database.
Yi Huang 0002, Kai Yu 0001, Matthias Schubert, Shipeng Yu, Volker Tresp, Hans-Peter Kriegel
ICDM4
2005 Dirichlet enhanced relational learning
abstract
We apply nonparametric hierarchical Bayesian modelling to relational learning. In a hierarchical Bayesian approach, model parameters can be "personalized", i.e., owned by entities or relationships, and are coupled via a common prior distribution. Flexibility is added in a nonparametric hierarchical Bayesian approach, such that the learned knowledge can be truthfully represented. We apply our approach to a medical domain where we form a nonparametric hierarchical Bayesian model for relations involving hospitals, patients, procedures and diagnosis. The experiments show that the additional flexibility in a nonparametric hierarchical Bayes approach results in a more accurate model of the dependencies between procedures and diagnosis and gives significantly improved estimates of the probabilities of future procedures.
Zhao Xu 0001, Volker Tresp, Kai Yu 0001, Shipeng Yu, Hans-Peter Kriegel
ICML4
2005 Soft Clustering on Graphs
abstract
We propose a simple clustering framework on graphs encoding pairwise data similarities. Unlike usual similarity-based methods, the approach softly assigns data to clusters in a probabilistic way. More importantly, a hierarchical clustering is naturally derived in this framework to gradually merge lower-level clusters into higher-level ones. A random walk analysis indicates that the algorithm exposes clustering structures in various resolutions, i.e., a higher level statistically models a longer-term diffusion on graphs and thus discovers a more global clustering structure. Finally we provide very encouraging experimental results.
Shipeng Yu, Kai Yu 0001, Volker Tresp
NIPS1
2005 A Probabilistic Clustering-Projection Model for Discrete Data
Shipeng Yu, Kai Yu 0001, Volker Tresp, Hans-Peter Kriegel
PKDD1
2005 Multi-label informed latent semantic indexing
abstract
Latent semantic indexing (LSI) is a well-known unsupervised approach for dimensionality reduction in information retrieval. However if the output information (i.e. category labels) is available, it is often beneficial to derive the indexing not only based on the inputs but also on the target values in the training data set. This is of particular importance in applications with multiple labels, in which each document can belong to several categories simultaneously. In this paper we introduce the multi-label informed latent semantic indexing (MLSI) algorithm which preserves the information of inputs and meanwhile captures the correlations between the multiple outputs. The recovered "latent semantics" thus incorporate the human-annotated category information and can be used to greatly improve the prediction accuracy. Empirical study based on two data sets, Reuters-21578 and RCV1, demonstrates very encouraging results.
Kai Yu 0001, Shipeng Yu, Volker Tresp
SIGIR2
2004 Block-based web search
abstract
Multiple-topic and varying-length of web pages are two negative factors significantly affecting the performance of web search. In this paper, we explore the use of page segmentation algorithms to partition web pages into blocks and investigate how to take advantage of block-level evidence to improve retrieval performance in the web context. Because of the special characteristics of web pages, different page segmentation method will have different impact on web search performance. We compare four types of methods, including fixed-length page segmentation, DOM-based page segmentation, vision-based page segmentation, and a combined method which integrates both semantic and fixed-length properties. Experiments on block-level query expansion and retrieval are performed. Among the four approaches, the combined method achieves the best performance for web search. Our experimental results also show that such a semantic partitioning of web pages effectively deals with the problem of multiple drifting topics and mixed lengths, and thus has great potential to boost up the performance of current web search engines.
Deng Cai 0001, Shipeng Yu, Ji-Rong Wen, Wei-Ying Ma
SIGIR2
2004 A nonparametric hierarchical bayesian framework for information filtering
abstract
Information filtering has made considerable progress in recent years. The predominant approaches are content-based methods and collaborative methods. Researchers have largely concentrated on either of the two approaches since a principled unifying framework is still lacking. This paper suggests that both approaches can be combined under a hierarchical Bayesian framework. Individual content-based user profiles are generated and collaboration between various user models is achieved via a common learned prior distribution. However, it turns out that a parametric distribution (e.g. Gaussian) is too restrictive to describe such a common learned prior distribution. We thus introduce a nonparametric common prior, which is a sample generated from a Dirichlet process which assumes the role of a hyper prior. We describe effective means to learn this nonparametric distribution, and apply it to learn users' information needs. The resultant algorithm is simple and understandable, and offers a principled solution to combine content-based filtering and collaborative filtering. Within our framework, we are now able to interpret various existing techniques from a unifying point of view. Finally we demonstrate the empirical success of the proposed information filtering methods.
Kai Yu 0001, Volker Tresp, Shipeng Yu
SIGIR3
2003 Extracting Content Structure for Web Pages Based on Visual Representation
Deng Cai 0001, Shipeng Yu, Ji-Rong Wen, Wei-Ying Ma
APWeb2
2003 Improving pseudo-relevance feedback in web information retrieval using web page segmentation
abstract
In contrast to traditional document retrieval, a web page as a whole is not a good information unit to search because it often contains multiple topics and a lot of irrelevant information from navigation, decoration, and interaction part of the page. In this paper, we propose a VIsion-based Page Segmentation (VIPS) algorithm to detect the semantic content structure in a web page. Compared with simple DOM based segmentation method, our page segmentation scheme utilizes useful visual cues to obtain a better partition of a page at the semantic level. By using our VIPS algorithm to assist the selection of query expansion terms in pseudo-relevance feedback in web information retrieval, we achieve 27% performance improvement on Web Track dataset.
Shipeng Yu, Deng Cai 0001, Ji-Rong Wen, Wei-Ying Ma
WWW1