Bokai Cao

dblp:130/3949 · DBLP profile ↗
← Back
36ranked-venue papers
11as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 9 first-author · 3 since 2021Databases, data management, data science and information retrieval · 27 · 10 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Data mining · 83% Recommender systems · 12% Knowledge graphs · 4%
Artificial intelligence
7 papers
Graph learning · 34% Efficient and distributed learning · 32% Language models and text generation · 27%
Network and information security
3 papers
Authentication and access control · 42% Privacy and data protection · 30% Biometric security · 28%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Computational finance and economics · 48% Bioinformatics and computational biology · 29% Medical and health informatics · 22%
Human-computer interaction and pervasive computing
2 papers
Health and well-being technologies · 60% Ubiquitous computing and smart environments · 21% Wearable and physiological sensing · 18%

Topics — the 30 heaviest of 46, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
multi-view learning
1.142018
Learning from Multi-View Multi-Way Data via Structural Factorization Machines · WWW 2018
Multilinear Factorization Machines for Multi-Task Multi-View Learning · WSDM 2017
Multi-view Machines · WSDM 2016
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
Financial Wind Tunnel: A Retrieval-Augmented Market Simulator · WWW 2026
Computational finance and economics › financial modeling
market simulation
1.012026
Financial Wind Tunnel: A Retrieval-Augmented Market Simulator · WWW 2026
Data mining › structured data mining
graph mining
0.732017
HitFraud: A Broad Learning Approach for Collective Fraud Detection in Heterogeneous Information Networks · ICDM 2017
Mining Brain Networks Using Multiple Side Views for Neurological Disorder Identification · ICDM 2015
Collective Prediction of Multiple Types of Links in Heterogeneous Information Networks · ICDM 2014
Machine learning › Graph learning
graph representation learning
0.622017
Cross View Link Prediction by Learning Noise-resilient Representation Consensus · WWW 2017
Structural Deep Brain Network Mining · KDD 2017
Biometric security
behavioral biometrics
0.512021
Kollector: Detecting Fraudulent Activities on Mobile Devices Using Deep Learning · IEEE Trans. Mob. Comput. 2021
Authentication and access control
continuous authentication
0.512021
Kollector: Detecting Fraudulent Activities on Mobile Devices Using Deep Learning · IEEE Trans. Mob. Comput. 2021
Authentication and access control › authentication
impersonation attack detection
0.512021
Kollector: Detecting Fraudulent Activities on Mobile Devices Using Deep Learning · IEEE Trans. Mob. Comput. 2021
Biometric security › behavioral biometrics
keystroke dynamics
0.512021
Kollector: Detecting Fraudulent Activities on Mobile Devices Using Deep Learning · IEEE Trans. Mob. Comput. 2021
Authentication and access control
user authentication
0.512021
Kollector: Detecting Fraudulent Activities on Mobile Devices Using Deep Learning · IEEE Trans. Mob. Comput. 2021
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis
tensor factorization
0.422018
Learning from Multi-View Multi-Way Data via Structural Factorization Machines · WWW 2018
Multilinear Factorization Machines for Multi-Task Multi-View Learning · WSDM 2017
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.412019
Private Model Compression via Knowledge Distillation · AAAI 2019
Machine learning › Efficient and distributed learning
model compression
0.412019
Private Model Compression via Knowledge Distillation · AAAI 2019
Privacy and data protection
differential privacy
0.412019
Private Model Compression via Knowledge Distillation · AAAI 2019
Privacy and data protection › privacy-preserving machine learning
privacy-preserving model compression
0.412019
Private Model Compression via Knowledge Distillation · AAAI 2019
Machine learning › Graph learning
network embedding
0.312018
Multi-View Multi-Graph Embedding for Brain Network Clustering Analysis · AAAI 2018
Recommender systems › factorization models
factorization machines
0.312018
Learning from Multi-View Multi-Way Data via Structural Factorization Machines · WWW 2018
Data mining › multidimensional data analysis
multiway data analysis
0.312018
Learning from Multi-View Multi-Way Data via Structural Factorization Machines · WWW 2018
Health and well-being technologies
mental health
0.312018
dpMood: Exploiting Local and Periodic Typing Dynamics for Personalized Mood Prediction · ICDM 2018
Health and well-being technologies
mood prediction
0.312018
dpMood: Exploiting Local and Periodic Typing Dynamics for Personalized Mood Prediction · ICDM 2018
Ubiquitous computing and smart environments › mobile sensing
smartphone sensing
0.312018
dpMood: Exploiting Local and Periodic Typing Dynamics for Personalized Mood Prediction · ICDM 2018
Privacy and data protection
privacy-preserving machine learning
0.312018
Not Just Privacy: Improving Performance of Private Deep Learning in Mobile Cloud · KDD 2018
Bioinformatics and computational biology › neuroscience › neuroinformatics
brain network analysis
0.322018
Mining Brain Networks Using Multiple Side Views for Neurological Disorder Identification · ICDM 2015
Multi-View Multi-Graph Embedding for Brain Network Clustering Analysis · AAAI 2018
Machine learning › Graph learning
link prediction
0.312017
Cross View Link Prediction by Learning Noise-resilient Representation Consensus · WWW 2017
Data mining
anomaly detection
0.312017
HitFraud: A Broad Learning Approach for Collective Fraud Detection in Heterogeneous Information Networks · ICDM 2017
Data mining › anomaly detection
fraud detection
0.312017
HitFraud: A Broad Learning Approach for Collective Fraud Detection in Heterogeneous Information Networks · ICDM 2017
Data mining › structured data mining › graph mining › heterogeneous information network
heterogeneous information network mining
0.312017
HitFraud: A Broad Learning Approach for Collective Fraud Detection in Heterogeneous Information Networks · ICDM 2017
Health and well-being technologies › mental health technology
mental health monitoring
0.312017
DeepMood: Modeling Mobile Phone Typing Dynamics for Mood Detection · KDD 2017
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection
0.212016
Unsupervised Feature Selection on Networks: A Generative View · AAAI 2016
Recommender systems › click-through rate prediction
feature interaction
0.212016
Multi-view Machines · WSDM 2016

Methods — techniques the papers use, named apart from their topics

deep learning · 1.7tensor decomposition · 1.4multi-view learning · 0.9knowledge distillation · 0.8hint learning · 0.8differential privacy · 0.8random noise addition · 0.7privacy budget analysis · 0.7noisy training · 0.7data nullification · 0.7multi-view bagging · 0.5meta-path · 0.5structural factorization machines · 0.3recurrent neural network · 0.3convolutional neural network · 0.3tensor factorization · 0.3subgraph mining · 0.3representation consensus learning · 0.3
YearPublicationVenuePosition
2026 Financial Wind Tunnel: A Retrieval-Augmented Market Simulator
Bokai Cao, Xueyuan Lin, Yiyan Qi, Chengjin Xu, Cehao Yang, Jian Guo 0016
WWW1
2023 MVMAFOL: A Multi-Access Three-Layer Federated Online Learning Algorithm for Internet of Vehicles
abstract
With the development of intelligent transportation system, it is urgent to transmit and analyze traffic data based on Internet of Vehicles. Federated learning has the characteristics of privacy protection, distributed learning and making full use of the computing power of local devices, which is very suitable for IoV's streaming big data analysis. Classical federated learning has only one global model, thus it can't be applied directly to IoV, where different region may have different data distribution. It is preferable that model of different region can be trained separately. Besides, the high mobility of vehicle nodes allows them to enter other areas while collecting data. This paper tries to solve the above issues. Firstly, a new vehicle-edge-cloud three-layer federated learning architecture for Io V, the Mobile Vehicle Multiple Access Federated Online Learning (MVMAFOL) algorithm, is proposed, where each vehicle terminal can upload the local model to multiple edge servers to maximize the use of its local model and effectively improve the generalization ability of the global model. Besides, a Model Similarity Vehicle Multiple Access (MSVMA) algorithm is proposed to improve the performance of MVMAFOL in the regional model aggregation phase, and an Edge Cloud Parameter Amendment (ECPA) algorithm is proposed for the global model aggregation phase. Simulation results show that this scheme can effectively improve the accuracy of the trained model.
Jieying Zhou, Junyao Zheng, Bokai Cao, Weigang Wu
IJCNN3
2022 EFL-WP: Federated Learning-Based Workload Prediction in Inter-Cloud Environments
abstract
Resource allocation has been always a major concern of cloud providers. Workload prediction can effectively improve resource utility by providing information about resources available in the future. Machine learning-based workload prediction has been widely studied and deployed in large-scale clouds. However, workload prediction in Inter-Cloud environments has not been considered. Since more and more cloud services are orchestrated by containers (or virtual machines) distributed in multiple clouds, intelligent models for workload prediction should be trained based on trace data across clouds. Due to concerns raised by data privacy, the trace data of clouds may not be shared with each other, and general distributed training is not applicable. In this paper, we propose a general and flexible framework, namely EFL-WP, for training workload prediction models in the Inter-Cloud environment. EFL-WP adopts the federated learning approach and allows cloud providers to collaborate to train prediction models without sharing traces. EFL-WP considers the difference among workloads in different cloud (i.e., the Non-IID characteristic of traces) by using two novel techniques: participant selection mechanism and multi-global models aggregation. The former can prevent some local models trained on non-IID traces from participating in global aggregation, which is beneficial to the global model. The latter allows the coordinator to aggregate several global models according to the difference among local models. To further improve accuracy, EFL-WP adopts an ensemble inference strategy. Experimental results show that the proposed framework can be superior to baselines on both Alibaba dataset and Tencent Games Traces.
Danyang Xiao, Bokai Cao, Weigang Wu
IJCNN2
2022 Reuse of Client Models in Federated Learning
abstract
Federated Learning (FL) has attracted a lot of attention from both academia and industry. In FL, user data no more need to be transmitted to the data center and each client device trains the deep model using its personal data so as to protect user privacy from being revealed. There exists two kinds of FL architecture, cloud based and edge based. Researchers have proposed the client-edge-cloud hierarchical FL system combining their advantage together to take the full advantage and avoid the defects. To improve the utilization ratio of local model parameters and data, so that enhance every single edge model accuracy, and ultimately achieve a more outstanding global model, we propose an algorithm Client Model Multiple Access (CMMA). CMMA allows clients associate with a set of edge servers, and uploads its training results to all servers in the set. That said, the local model of one client is reused by multiple edge servers. Such reuse can improve model performance particularly when the number of clients is small. Empirical experiments demonstrate the superiority of our scheme in different datasets, CNN models and data distributions. The results have not just shown the general suppress CMMA against hierarchical FL, but also validate the advantage of CMMA guaranteeing global model accuracy in unstable network or short of clients dilemma.
Bokai Cao, Weigang Wu, Congcong Zhan, Jieying Zhou
SMARTCOMP1
2022 RTGA: Robust ternary gradients aggregation for federated learning
Chengang Yang, Danyang Xiao, Bokai Cao, Weigang Wu
Inf. Sci.3
2021 Kollector: Detecting Fraudulent Activities on Mobile Devices Using Deep Learning
abstract
With the rapid growth in smartphone usage, preventing leakage of personal information and privacy has become a challenging task. One major consequence of such leakage is impersonation. This type of illegal usage is nearly impossible to prevent as existing preventive mechanisms (e.g., passcode and fingerprinting), are not capable of continuously monitoring usage and determining whether the user is authorized. Once unauthorized users can defeat the initial protection mechanisms, they would have full access to the devices including using stored passwords to access high-value websites. We present Kollector, a new framework to detect impersonation based on a multi-view bagging deep learning approach to capture sequential tapping information on the smart-phone's keyboard. We construct a sequential-tapping biometrics model to continuously authenticate the user while typing. We empirically evaluated our system using real-world phone usage sessions from 26 users over eight weeks. We then compared our model against commonly used shallow machine techniques and find that our system performs better than other approaches and can achieve an 8.42 percent equal error rate, a 94.24 percent accuracy and a 94.41 percent H-mean using only the accelerometer and only five keyboard taps. We also experiment with using only three keyboard taps and find that the system still yields high accuracy while giving additional opportunities to make more decisions that can result in more accurate final decisions.
Lichao Sun 0001, Bokai Cao, Ji Wang 0002, Witawas Srisa-an, Philip S. Yu, Alex D. Leow, Stephen Checkoway
IEEE Trans. Mob. Comput.2
2019 Private Model Compression via Knowledge Distillation
abstract
The soaring demand for intelligent mobile applications calls for deploying powerful deep neural networks (DNNs) on mobile devices. However, the outstanding performance of DNNs notoriously relies on increasingly complex models, which in turn is associated with an increase in computational expense far surpassing mobile devices’ capacity. What is worse, app service providers need to collect and utilize a large volume of users’ data, which contain sensitive information, to build the sophisticated DNN models. Directly deploying these models on public mobile devices presents prohibitive privacy risk. To benefit from the on-device deep learning without the capacity and privacy concerns, we design a private model compression framework RONA. Following the knowledge distillation paradigm, we jointly use hint learning, distillation learning, and self learning to train a compact and fast neural network. The knowledge distilled from the cumbersome model is adaptively bounded and carefully perturbed to enforce differential privacy. We further propose an elegant query sample selection method to reduce the number of queries and control the privacy loss. A series of empirical evaluations as well as the implementation on an Android mobile device show that RONA can not only compress cumbersome models efficiently but also provide a strong privacy guarantee. For example, on SVHN, when a meaningful (9.83,10−6)-differential privacy is guaranteed, the compact model trained by RONA can obtain 20× compression ratio and 19× speed-up with merely 0.97% accuracy loss.
Ji Wang 0002, Weidong Bao 0001, Lichao Sun 0001, Xiaomin Zhu 0001, Bokai Cao, Philip S. Yu
AAAI5
2018 Multi-View Multi-Graph Embedding for Brain Network Clustering Analysis
abstract
Network analysis of human brain connectivity is critically important for understanding brain function and disease states. Embedding a brain network as a whole graph instance into a meaningful low-dimensional representation can be used to investigate disease mechanisms and inform therapeutic interventions. Moreover, by exploiting information from multiple neuroimaging modalities or views, we are able to obtain an embedding that is more useful than the embedding learned from an individual view. Therefore, multi-view multi-graph embedding becomes a crucial task. Currently only a few studies have been devoted to this topic, and most of them focus on vector-based strategy which will cause structural information contained in the original graphs lost. As a novel attempt to tackle this problem, we propose Multi-view Multi-graph Embedding M2E by stacking multi-graphs into multiple partially-symmetric tensors and using tensor techniques to simultaneously leverage the dependencies and correlations among multi-view and multi-graph brain networks. Extensive experiments on real HIV and bipolar disorder brain network datasets demonstrate the superior performance of M2E on clustering brain networks by leveraging the multi-view multi-graph interactions.
Ye Liu 0006, Lifang He 0001, Bokai Cao, Philip S. Yu, Ann B. Ragin, Alex D. Leow
AAAI3
2018 Deep Learning towards Mobile Applications
abstract
Recent years have witnessed an explosive growth of mobile devices. Mobile devices are permeating every aspect of our daily lives. With the increasing usage of mobile devices and intelligent applications, there is a soaring demand for mobile applications with machine learning services. Inspired by the tremendous success achieved by deep learning in many machine learning tasks, it becomes a natural trend to push deep learning towards mobile applications. However, there exist many challenges to realize deep learning in mobile applications, including the contradiction between the miniature nature of mobile devices and the resource requirement of deep neural networks, the privacy and security concerns about individuals' data, and so on. To resolve these challenges, during the past few years, great leaps have been made in this area. In this paper, we provide an overview of the current challenges and representative achievements about pushing deep learning on mobile devices from three aspects: training with mobile data, efficient inference on mobile devices, and applications of mobile deep learning. The former two aspects cover the primary tasks of deep learning. Then, we go through our two recent applications that apply the data collected by mobile devices to inferring mood disturbance and user identification. Finally, we conclude this paper with the discussion of the future of this area.
Ji Wang 0002, Bokai Cao, Philip S. Yu, Lichao Sun 0001, Weidong Bao 0001, Xiaomin Zhu 0001
ICDCS2
2018 dpMood: Exploiting Local and Periodic Typing Dynamics for Personalized Mood Prediction
abstract
Mood disorders are common and associated with significant morbidity and mortality. Early diagnosis has the potential to greatly alleviate the burden of mental illness and the ever increasing costs to families and society. Mobile devices provide us a promising opportunity to detect the users' mood in an unobtrusive manner. In this study, we use a custom keyboard which collects keystrokes' meta-data and accelerometer values. Based on the collected time series data in multiple modalities, we propose a deep personalized mood prediction approach, called dpMood, by integrating convolutional and recurrent deep architectures as well as exploring each individual's circadian rhythm. Experimental results not only demonstrate the feasibility and effectiveness of using smart-phone meta-data to predict the presence and severity of mood disturbances in bipolar subjects, but also show the potential of personalized medical treatment for mood disorders.
He Huang 0008, Bokai Cao, Philip S. Yu, Chang-Dong Wang 0001, Alex D. Leow
ICDM2
2018 Not Just Privacy: Improving Performance of Private Deep Learning in Mobile Cloud
abstract
The increasing demand for on-device deep learning services calls for a highly efficient manner to deploy deep neural networks (DNNs) on mobile devices with limited capacity. The cloud-based solution is a promising approach to enabling deep learning applications on mobile devices where the large portions of a DNN are offloaded to the cloud. However, revealing data to the cloud leads to potential privacy risk. To benefit from the cloud data center without the privacy risk, we design, evaluate, and implement a cloud-based framework ARDEN which partitions the DNN across mobile devices and cloud data centers. A simple data transformation is performed on the mobile device, while the resource-hungry training and the complex inference rely on the cloud data center. To protect the sensitive information, a lightweight privacy-preserving mechanism consisting of arbitrary data nullification and random noise addition is introduced, which provides strong privacy guarantee. A rigorous privacy budget analysis is given. Nonetheless, the private perturbation to the original data inevitably has a negative impact on the performance of further inference on the cloud side. To mitigate this influence, we propose a noisy training method to enhance the cloud-side network robustness to perturbed data. Through the sophisticated design, ARDEN can not only preserve privacy but also improve the inference performance. To validate the proposed ARDEN, a series of experiments based on three image datasets and a real mobile application are conducted. The experimental results demonstrate the effectiveness of ARDEN. Finally, we implement ARDEN on a demo system to verify its practicality.
Ji Wang 0002, Jianguo Zhang 0005, Weidong Bao 0001, Xiaomin Zhu 0001, Bokai Cao, Philip S. Yu
KDD5
2018 Learning from Multi-View Multi-Way Data via Structural Factorization Machines
abstract
Real-world relations among entities can often be observed and determined by different perspectives/views. For example, the decision made by a user on whether to adopt an item relies on multiple aspects such as the contextual information of the decision, the item»s attributes, the user»s profile and the reviews given by other users. Different views may exhibit multi-way interactions among entities and provide complementary information. In this paper, we introduce a multi-tensor-based approach that can preserve the underlying structure of multi-view data in a generic predictive model. Specifically, we propose structural factorization machines (SFMs) that learn the common latent spaces shared by multi-view tensors and automatically adjust the importance of each view in the predictive model. Furthermore, the complexity of SFMs is linear in the number of parameters, which make SFMs suitable to large-scale problems. Extensive experiments on real-world datasets demonstrate that the proposed SFMs outperform several state-of-the-art methods in terms of prediction accuracy and computational cost.
Chun-Ta Lu, Lifang He 0001, Hao Ding 0003, Bokai Cao, Philip S. Yu
WWW4
2017 Connecting emerging relationships from news via tensor factorization
abstract
Knowledge graphs (KGs) have been widely used to represent relationships among entities, while KGs cannot capture new relationships between entities emerging along time. Since news often provides the latest information regarding the new entities and relationships, there is an opportunity to connect emerging relationships from news timely. However, it is a challenging task due to the source heterogeneity of structured KGs and unstructured news texts. In order to address the issue, we propose a tensor-based framework to capture the complex interactions among multiple types of relations, entities and text descriptions. We further develop an efficient Text-Aware MUlti-RElational learning method (TAMURE) that can learn the embedding representations of entities and relation types from both KGs and news, by jointly factorizing the interaction parameters. Furthermore, the complexity of TAMURE is linear in the number of parameters, which makes it suitable to large-scale KGs and news texts. Extensive experiments via TensorFlow demonstrate the effectiveness of the proposed TAMURE model compared with nine state-of-the-art methods on real-world datasets.
Chun-Ta Lu, Bokai Cao, Yi Chang 0001, Philip S. Yu
IEEE BigData3
2017 Hierarchical collaborative embedding for context-aware recommendations
abstract
In a variety of recommender systems, items, such as news or articles, are associated with text. Most of previous recommender systems learn item embeddings from the textual content by utilizing the bag-of-words technique. However, due to its limited ability to capture semantic meanings in the text, these methods lead to the shallow modeling of items. Recently proposed deep learning based methods try to overcome the limitation by leveraging Recurrent Neural Networks (RNN). Suffering from the problem of modeling long sequences for RNN, these methods are unable to effectively model items based on their textual content as well. In this paper, in order to overcome aforementioned limitations and accurately capture semantic meanings within the textual content, we propose Hierarchical Collaborative Embedding (HCE). HCE tightly couples a Hierarchical Recurrent Network (HRN) with Probabilistic Matrix Factorization (PMF) to provide top-N ranking lists of items for users. In the experiments, we show that HCE beats strong baselines by a wide margin on three real-world datasets.
Lei Zheng 0001, Bokai Cao, Vahid Noroozi, Philip S. Yu, Nianzu Ma
IEEE BigData2
2017 Unsupervised Feature Selection with Heterogeneous Side Information
abstract
Compared to supervised feature selection, unsupervised feature selection tends to be more challenging due to the lack of guidance from class labels. Along with the increasing variety of data sources, many datasets are also equipped with certain side information of heterogeneous structure. Such side information can be critical for feature selection when class labels are unavailable. In this paper, we propose a new feature selection method, SideFS, to exploit such rich side information. We model the complex side information as a heterogeneous network and derive instance correlations to guide subsequent feature selection. Representations are learned from the side information network and the feature selection is performed in a unified framework. Experimental results show that the proposed method can effectively enhance the quality of selected features by incorporating heterogeneous side information.
Xiaokai Wei, Bokai Cao, Philip S. Yu
CIKM2
2017 HitFraud: A Broad Learning Approach for Collective Fraud Detection in Heterogeneous Information Networks
abstract
On electronic game platforms, different payment transactions have different levels of risk. Risk is generally higher for digital goods in e-commerce. However, it differs based on product and its popularity, the offer type (packaged game, virtual currency to a game or subscription service), storefront and geography. Existing fraud policies and models make decisions independently for each transaction based on transaction attributes, payment velocities, user characteristics, and other relevant information. However, suspicious transactions may still evade detection and hence we propose a broad learning approach leveraging a graph based perspective to uncover relationships among suspicious transactions, i.e., inter-transaction dependency. Our focus is to detect suspicious transactions by capturing common fraudulent behaviors that would not be considered suspicious when being considered in isolation. In this paper, we present HitFraud that leverages heterogeneous information networks for collective fraud detection by exploring correlated and fast evolving fraudulent behaviors. First, a heterogeneous information network is designed to link entities of interest in the transaction database via different semantics. Then, graph based features are efficiently discovered from the network exploiting the concept of meta-paths, and decisions on frauds are made collectively on test instances. Experiments on real-world payment transaction data from Electronic Arts demonstrate that the prediction performance is effectively boosted by HitFraud with fast convergence.
Bokai Cao, Mia Mao, Siim Viidu, Philip S. Yu
ICDM1
2017 Multi-view unsupervised feature selection by cross-diffused matrix alignment
abstract
Multi-view high-dimensional data become increasingly popular in the big data era. Feature selection is a useful technique for alleviating the curse of dimensionality in multi-view learning. In this paper, we study unsupervised feature selection for multi-view data, as class labels are usually expensive to obtain. Traditional feature selection methods are mostly designed for single-view data and cannot fully exploit the rich information from multi-view data. Existing multi-view feature selection methods are usually based on noisy cluster labels which might not preserve sufficient information from multi-view data. To better utilize multi-view information, we propose a method, CDMA-FS, to select features for each view by performing alignment on a cross diffused matrix. We formulate it as a constrained optimization problem and solve it using Quasi-Newton based method. Experiments results on four real-world datasets show that the proposed method is more effective than the state-of-the-art methods in multi-view setting.
Xiaokai Wei, Bokai Cao, Philip S. Yu
IJCNN2
2017 DeepMood: Modeling Mobile Phone Typing Dynamics for Mood Detection
abstract
The increasing use of electronic forms of communication presents new opportunities in the study of mental health, including the ability to investigate the manifestations of psychiatric diseases unobtrusively and in the setting of patients' daily lives. A pilot study to explore the possible connections between bipolar affective disorder and mobile phone usage was conducted. In this study, participants were provided a mobile phone to use as their primary phone. This phone was loaded with a custom keyboard that collected metadata consisting of keypress entry time and accelerometer movement. Individual character data with the exceptions of the backspace key and space bar were not collected due to privacy concerns. We propose an end-to-end deep architecture based on late fusion, named DeepMood, to model the multi-view metadata for the prediction of mood scores. Experimental results show that 90.31% prediction accuracy on the depression score can be achieved based on session-level mobile phone typing dynamics which is typically less than one minute. It demonstrates the feasibility of using mobile phone metadata to infer mood disturbance and severity.
Bokai Cao, Lei Zheng 0001, Philip S. Yu, Andrea Piscitello, John Zulueta, Olusola Ajilore, Kelly Ryan, Alex D. Leow
KDD1
2017 Structural Deep Brain Network Mining
abstract
Mining from neuroimaging data is becoming increasingly popular in the field of healthcare and bioinformatics, due to its potential to discover clinically meaningful structure patterns that could facilitate the understanding and diagnosis of neurological and neuropsychiatric disorders. Most recent research concentrates on applying subgraph mining techniques to discover connected subgraph patterns in the brain network. However, the underlying brain network structure is complicated. As a shallow linear model, subgraph mining cannot capture the highly non-linear structures, resulting in sub-optimal patterns. Therefore, how to learn representations that can capture the highly non-linearity of brain networks and preserve the underlying structures is a critical problem.
Shen Wang 0005, Lifang He 0001, Bokai Cao, Chun-Ta Lu, Philip S. Yu, Ann B. Ragin
KDD3
2017 Sequential Keystroke Behavioral Biometrics for Mobile User Identification via Multi-view Deep Learning
Lichao Sun 0001, Bokai Cao, Philip S. Yu, Witawas Srisa-an, Alex D. Leow
ECML/PKDD (3)3
2017 Rethinking Unsupervised Feature Selection: From Pseudo Labels to Pseudo Must-Links
Xiaokai Wei, Sihong Xie, Bokai Cao, Philip S. Yu
ECML/PKDD (1)3
2017 t-BNE: Tensor-based Brain Network Embedding
abstract
Brain network embedding is the process of converting brain network data to discriminative representations of subjects, so that patients with brain disorders and normal controls can be easily separated. Computer-aided diagnosis based on such representations is potentially transformative for investigating disease mechanisms and for informing therapeutic interventions. However, existing methods either limit themselves to extracting graph-theoretical measures and subgraph patterns, or fail to incorporate brain network properties and domain knowledge in medical science. In this paper, we propose t-BNE, a novel Brain Network Embedding model based on constrained tensor factorization. t-BNE incorporates 1) symmetric property of brain networks, 2) side information guidance to obtain representations consistent with auxiliary measures, 3) orthogonal constraint to make the latent factors distinct with each other, and 4) classifier learning procedure to introduce supervision from labeled data. The Alternating Direction Method of Multipliers (ADMM) framework is utilized to solve the optimization objective. We evaluate t-BNE on three EEG brain network datasets. Experimental results illustrate the superior performance of the proposed model on graph classification tasks with significant improvement 20.51%, 6.38% and 12.85%, respectively. Furthermore, the derived factors are visualized which could be informative for investigating disease mechanisms under different emotion regulation tasks.
Bokai Cao, Lifang He 0001, Xiaokai Wei, Mengqi Xing, Philip S. Yu, Heide Klumpp, Alex D. Leow
SDM1
2017 Multilinear Factorization Machines for Multi-Task Multi-View Learning
abstract
Many real-world problems, such as web image analysis, document categorization and product recommendation, often exhibit dual-heterogeneity: heterogeneous features obtained in multiple views, and multiple tasks might be related to each other through one or more shared views. To address these Multi-Task Multi-View (MTMV) problems, we propose a tensor-based framework for learning the predictive multilinear structure from the full-order feature interactions within the heterogeneous data. The usage of tensor structure is to strengthen and capture the complex relationships between multiple tasks with multiple views. We further develop efficient multilinear factorization machines (MFMs) that can learn the task-specific feature map and the task-view shared multilinear structures, without physically building the tensor. In the proposed method, a joint factorization is applied to the full-order interactions such that the consensus representation can be learned. In this manner, it can deal with the partially incomplete data without difficulty as the learning procedure does not simply rely on any particular view. Furthermore, the complexity of MFMs is linear in the number of parameters, which makes MFMs suitable to large-scale real-world problems. Extensive experiments on four real-world datasets demonstrate that the proposed method significantly outperforms several state-of-the-art methods in a wide variety of MTMV problems.
Chun-Ta Lu, Lifang He 0001, Weixiang Shao, Bokai Cao, Philip S. Yu
WSDM4
2017 Cross View Link Prediction by Learning Noise-resilient Representation Consensus
abstract
Link Prediction has been an important task for social and information networks. Existing approaches usually assume the completeness of network structure. However, in many real-world networks, the links and node attributes can usually be partially observable. In this paper, we study the problem of Cross View Link Prediction (CVLP) on partially observable networks, where the focus is to recommend nodes with only links to nodes with only attributes (or vice versa). We aim to bridge the information gap by learning a robust consensus for link-based and attribute-based representations so that nodes become comparable in the latent space. Also, the link-based and attribute-based representations can lend strength to each other via this consensus learning. Moreover, attribute selection is performed jointly with the representation learning to alleviate the effect of noisy high-dimensional attributes. We present two instantiations of this framework with different loss functions and develop an alternating optimization framework to solve the problem. Experimental results on four real-world datasets show the proposed algorithm outperforms the baseline methods significantly for cross-view link prediction.
Xiaokai Wei, Linchuan Xu, Bokai Cao, Philip S. Yu
WWW3
2016 Unsupervised Feature Selection on Networks: A Generative View
abstract
In the past decade, social and information networks have become prevalent, and research on the network data has attracted much attention. Besides the link structure, network data are often equipped with the content information (i.e, node attributes) that is usually noisy and characterized by high dimensionality. As the curse of dimensionality could hamper the performance of many machine learning tasks on networks (e.g., community detection and link prediction), feature selection can be a useful technique for alleviating such issue. In this paper, we investigate the problem of unsupervised feature selection on networks. Most existing feature selection methods fail to incorporate the linkage information, and the state-of-the-art approaches usually rely on pseudo labels generated from clustering. Such cluster labels may be far from accurate and can mislead the feature selection process. To address these issues, we propose a generative point of view for unsupervised features selection on networks that can seamlessly exploit the linkage and content information in a more effective manner. We assume that the link structures and node content are generated from a succinct set of high-quality features, and we find these features through maximizing the likelihood of the generation process. Experimental results on three real-world datasets show that our approach can select more discriminative features than state-of-the-art methods.
Xiaokai Wei, Bokai Cao, Philip S. Yu
AAAI2
2016 Community detection with partially observable links and node attributes
abstract
Community detection has been an important task for social and information networks. Existing approaches usually assume the completeness of linkage and content information. However, the links and node attributes can usually be partially observable in many real-world networks. For example, users can specify their privacy settings to prevent non-friends from viewing their posts or connections. Such incompleteness poses additional challenges to community detection algorithms. In this paper, we aim to detect communities with partially observable link structure and node attributes. To fuse such incomplete information, we learn link-based and attribute-based representations via kernel alignment and a co-regularization approach is proposed to combine the information from both sources (i.e., links and attributes). The link-based and attribute-based representations can lend strength to each other via the partial consensus learning. We present two instantiations of this framework by enforcing hard and soft consensus constraint respectively. Experimental results on real-world datasets show the superiority of the proposed approaches over the baseline methods and its robustness under different observable levels.
Xiaokai Wei, Bokai Cao, Weixiang Shao, Chun-Ta Lu, Philip S. Yu
IEEE BigData2
2016 Semi-supervised Tensor Factorization for Brain Network Analysis
Bokai Cao, Chun-Ta Lu, Xiaokai Wei, Philip S. Yu, Alex D. Leow
ECML/PKDD (1)1
2016 Multi-graph Clustering Based on Interior-Node Topology with Applications to Brain Networks
Guixiang Ma, Lifang He 0001, Bokai Cao, Jiawei Zhang 0001, Philip S. Yu, Ann B. Ragin
ECML/PKDD (1)3
2016 Nonlinear Joint Unsupervised Feature Selection
abstract
In the era of big data, one is often confronted with the problem of high dimensional data for many machine learning or data mining tasks. Feature selection, as a dimension reduction technique, is useful for alleviating the curse of dimensionality while preserving interpretability. In this paper, we focus on unsupervised feature selection, as class labels are usually expensive to obtain. Unsupervised feature selection is typically more challenging than its supervised counterpart due to the lack of guidance from class labels. Recently, regression-based methods with L2,1 norms have gained much popularity as they are able to evaluate features jointly which, however, consider only linear correlations between features and pseudo-labels. In this paper, we propose a novel nonlinear joint unsupervised feature selection method based on kernel alignment. The aim is to find a succinct set of features that best aligns with the original features in the kernel space. It can evaluate features jointly in a nonlinear manner and provides a good ‘0/1’ approximation for the selection indicator vector. We formulate it as a constrained optimization problem and develop a Spectral Projected Gradient (SPG) method to solve the optimization problem. Experimental results on several real-world datasets demonstrate that our proposed method outperforms the state-of-the-art approaches significantly.
Xiaokai Wei, Bokai Cao, Philip S. Yu
SDM2
2016 Identifying Connectivity Patterns for Brain Diseases via Multi-side-view Guided Deep Architectures
abstract
There is considerable interest in mining neuroimage data to discover clinically meaningful connectivity patterns to inform an understanding of neurological and neuropsychiatric disorders. Subgraph mining models have been used to discover connected subgraph patterns. However, it is difficult to capture the complicated interplay among patterns. As a result, classification performance based on these results may not be satisfactory. To address this issue, we propose to learn non-linear representations of brain connectivity patterns from deep learning architectures. This is non-trivial, due to the limited subjects and the high costs of acquiring the data. Fortunately, auxiliary information from multiple side views such as clinical, serologic, immunologic, cognitive and other diagnostic testing also characterizes the states of subjects from different perspectives. In this paper, we present a novel Multi-side-View guided AutoEncoder (MVAE) that incorporates multiple side views into the process of deep learning to tackle the bias in the construction of connectivity patterns caused by the scarce clinical data. Extensive experiments show that MVAE not only captures discriminative connectivity patterns for classification, but also discovers meaningful information for clinical interpretation.
Bokai Cao, Sihong Xie, Chun-Ta Lu, Philip S. Yu, Ann B. Ragin
SDM2
2016 Multi-view Machines
abstract
With rapidly growing amount of data available on the web, it becomes increasingly likely to obtain data from different perspectives for multi-view learning. Some successive examples of web applications include recommendation and target advertising. Specifically, to predict whether a user will click an ad in a query context, there are available features extracted from user profile, ad information and query description, and each of them can only capture part of the task signals from a particular aspect/view. Different views provide complementary information to learn a practical model for these applications. Therefore, an effective integration of the multi-view information is critical to facilitate the learning performance.
Bokai Cao, Hucheng Zhou, Guoqiang Li 0008, Philip S. Yu
WSDM1
2015 Inferring crowd-sourced venues for tweets
abstract
Knowing the geo-located venue of a tweet can facilitate better understanding of a user's geographic context, allowing apps to more precisely present information, recommend services, and target advertisements. However, due to privacy concerns, few users choose to enable geotagging of their tweets, resulting in a small percentage of tweets being geotagged; furthermore, even if the geo-coordinates are available, the closest venue to the geolocation may be incorrect. In this paper, we present a method for providing a ranked list of geo-located venues for a non-geotagged tweet, which simultaneously indicates the venue name and the geo-location at a very fine-grained granularity. In our proposed method for Venue Inference for Tweets (VIT), we construct a heterogeneous social network in order to analyze the embedded social relations, and leverage available but limited geographic data to estimate the geo-located venue of tweets. A single classifier is trained to estimate the probability of a tweet and a geo-located venue being linked, rather than training a separate model for each venue. We examine the performance of four types of social relation features and three types of geographic features embedded in a social network when inferring whether a tweet and a venue are linked, with a best accuracy of over 88%. We use the classifier probability estimates to rank the candidate geo-located venues of a non-geotagged tweet from over 19k possibilities, and observed an average top-5 accuracy of 29%.
Bokai Cao, Francine Chen 0001, Dhiraj Joshi, Philip S. Yu
IEEE BigData1
2015 Mining Brain Networks Using Multiple Side Views for Neurological Disorder Identification
abstract
Mining discriminative subgraph patterns from graph data has attracted great interest in recent years. It has a wide variety of applications in disease diagnosis, neuroimaging, etc. Most research on subgraph mining focuses on the graph representation alone. However, in many real-world applications, the side information is available along with the graph data. For example, for neurological disorder identification, in addition to the brain networks derived from neuroimaging data, hundreds of clinical, immunologic, serologic and cognitive measures may also be documented for each subject. These measures compose multiple side views encoding a tremendous amount of supplemental information for diagnostic purposes, yet are often ignored. In this paper, we study the problem of discriminative subgraph selection using multiple side views and propose a novel solution to find an optimal set of subgraph features for graph classification by exploring a plurality of side views. We derive a feature evaluation criterion, named gSide, to estimate the usefulness of subgraph patterns based upon side views. Then we develop a branch-and-bound algorithm, called gMSV, to efficiently search for optimal subgraph features by integrating the subgraph mining process and the procedure of discriminative feature selection. Empirical studies on graph classification tasks for neurological disorders using brain networks demonstrate that subgraph patterns selected by the multi-side-view guided subgraph selection approach can effectively boost graph classification performances and are relevant to disease diagnosis.
Bokai Cao, Xiangnan Kong, Philip S. Yu, Ann B. Ragin
ICDM1
2014 Tensor-Based Multi-view Feature Selection with Applications to Brain Diseases
abstract
In the era of big data, we can easily access information from multiple views which may be obtained from different sources or feature subsets. Generally, different views provide complementary information for learning tasks. Thus, multi-view learning can facilitate the learning process and is prevalent in a wide range of application domains. For example, in medical science, measurements from a series of medical examinations are documented for each subject, including clinical, imaging, immunologic, serologic and cognitive measures which are obtained from multiple sources. Specifically, for brain diagnosis, we can have different quantitative analysis which can be seen as different feature subsets of a subject. It is desirable to combine all these features in an effective way for disease diagnosis. However, some measurements from less relevant medical examinations can introduce irrelevant information which can even be exaggerated after view combinations. Feature selection should therefore be incorporated in the process of multi-view learning. In this paper, we explore tensor product to bring different views together in a joint space, and present a dual method of tensor-based multi-view feature selection (dual-Tmfs) based on the idea of support vector machine recursive feature elimination. Experiments conducted on datasets derived from neurological disorder demonstrate the features selected by our proposed method yield better classification performance and are relevant to disease diagnosis.
Bokai Cao, Lifang He 0001, Xiangnan Kong, Philip S. Yu, Ann B. Ragin
ICDM1
2014 Collective Prediction of Multiple Types of Links in Heterogeneous Information Networks
abstract
Link prediction has become an important and active research topic in recent years, which is prevalent in many real-world applications. Current research on link prediction focuses on predicting one single type of links, such as friendship links in social networks, or predicting multiple types of links independently. However, many real-world networks involve more than one type of links, and different types of links are not independent, but related with complex dependencies among them. In such networks, the prediction tasks for different types of links are also correlated and the links of different types should be predicted collectively. In this paper, we study the problem of collective prediction of multiple types of links in heterogeneous information networks. To address this problem, we introduce the linkage homophily principle and design a relatedness measure, called RM, between different types of objects to compute the existence probability of a link. We also extend conventional proximity measures to heterogeneous links. Furthermore, we propose an iterative framework for heterogeneous collective link prediction, called HCLP, to predict multiple types of links collectively by exploiting diverse and complex linkage information in heterogeneous information networks. Empirical studies on real-world tasks demonstrate that the proposed collective link prediction approach can effectively boost link prediction performances in heterogeneous information networks.
Bokai Cao, Xiangnan Kong, Philip S. Yu
ICDM1
2013 Multi-label classification by mining label and instance correlations from heterogeneous information networks
abstract
Multi-label classification is prevalent in many real-world applications, where each example can be associated with a set of multiple labels simultaneously. The key challenge of multi-label classification comes from the large space of all possible label sets, which is exponential to the number of candidate labels. Most previous work focuses on exploiting correlations among different labels to facilitate the learning process. It is usually assumed that the label correlations are given beforehand or can be derived directly from data samples by counting their label co-occurrences. However, in many real-world multi-label classification tasks, the label correlations are not given and can be hard to learn directly from data samples within a moderate-sized training set. Heterogeneous information networks can provide abundant knowledge about relationships among different types of entities including data samples and class labels. In this paper, we propose to use heterogeneous information networks to facilitate the multi-label classification process. By mining the linkage structure of heterogeneous information networks, multiple types of relationships among different class labels and data samples can be extracted. Then we can use these relationships to effectively infer the correlations among different class labels in general, as well as the dependencies among the label sets of data examples inter-connected in the network. Empirical studies on real-world tasks demonstrate that the performance of multi-label classification can be effectively boosted using heterogeneous information net- works.
Xiangnan Kong, Bokai Cao, Philip S. Yu
KDD2