VLDB 2026 Research / reviewers in the wild / expert
Vincent Wenchen Zheng
dblp:47/1961 · also Vincent W. Zheng
· DBLP profile ↗
64ranked-venue papers
12as first author
9since 2021 · last 2025
0000-0002-0904-3184ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 9 first-author · 3 since 2021Databases, data management, data science and information retrieval · 34 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-authorHuman-computer interaction and ubiquitous computing · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Multi-Expert Tabular Language Model for BankingabstractPre-training large Tabular Language Models (TaLMs) on tabular data has shown effectiveness for table understanding tasks. However, training proprietary large TaLMs on a company's private databases requires substantial computational resources. This paper presents an efficient multi-expert TaLM architecture and training method tailored for multi-domain databases and modest infrastructure. This architecture leverages a divide-and-conquer pretraining approach and a sparsely activated fine-tuning paradigm to reduce computation. Using this architecture, we pre-train and fine-tune a TaLM with 10 billion parameters on a banking database under simple computational infrastructures. We apply our TaLM to support various important banking applications, including risk assessment, information prediction, and profit assessment. Compared with previous baselines, our model achieves +29.3% in [email protected]% on risk assessment and +16.5% in accuracy on information prediction, showing great effectiveness and profitability of our model. This model is successfully deployed in WeBank and now supports many real business scenarios. Yue Guo 0009, Vincent Wenchen Zheng, Yi Yang 0042 |
KDD (1) | 4 |
| 2025 | Meta-path based proximity learning in heterogeneous information networks
Wenyi Xiao, Huan Zhao 0002, Vincent Wenchen Zheng, Yangqiu Song |
Data Min. Knowl. Discov. | 3 |
| 2023 | Neighbor-Anchoring Adversarial Graph Neural NetworksabstractGraph neural networks (GNNs) have witnessed widespread adoption due to their ability to learn superior representations for graph data. While GNNs exhibit strong discriminative power, they often fall short of learning the underlying node distribution for increased robustness. To deal with this, inspired by generative adversarial networks (GANs), we investigate the problem of adversarial learning on graph neural networks, and propose a novel framework named NAGNN (i.e., Neighbor-anchoring Adversarial Graph Neural Networks) for graph representation learning, which trains not only a discriminator but also a generator that compete with each other. In particular, we propose a novel neighbor-anchoring strategy, where the generator produces samples with explicit features and neighborhood structures anchored on a reference real node, so that the discriminator can perform neighborhood aggregation on the fake samples to learn superior representation. The advantage of our neighbor-anchoring strategy can be demonstrated both theoretically and empirically. Furthermore, as a by-product, our generator can synthesize realistic-looking features, enabling potential applications such as automatic content summarization. Finally, we conduct extensive experiments on four public benchmark datasets, and achieve promising results under both quantitative and qualitative evaluations. Yuan Fang 0001, Yong Liu 0020, Vincent Wenchen Zheng |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Neighbor-Anchoring Adversarial Graph Neural Networks (Extended Abstract)abstractWhile graph neural networks (GNNs) exhibit strong discriminative power, they often fall short of learning the underlying node distribution for increased robustness. To deal with this, inspired by generative adversarial networks (GANs), we investigate the problem of adversarial learning on graph neural networks, and propose a novel framework named NAGNN (i.e., Neighbor-anchoring Adversarial Graph Neural Networks) for graph representation learning, which trains not only a discriminator but also a generator that compete with each other. In particular, we propose a novel neighbor-anchoring strategy, where the generator produces samples with explicit features and neighborhood structures anchored on a reference real node, so that the discriminator can perform neighborhood aggregation on the fake samples to learn superior representations. Yuan Fang 0001, Yong Liu 0020, Vincent Wenchen Zheng |
ICDE | 4 |
| 2021 | Differentially Private Federated Knowledge Graphs EmbeddingabstractKnowledge graph embedding plays an important role in knowledge representation, reasoning, and data mining applications. However, for multiple cross-domain knowledge graphs, state-of-the-art embedding models cannot make full use of the data from different knowledge domains while preserving the privacy of exchanged data. In addition, the centralized embedding model may not scale to the extensive real-world knowledge graphs. Therefore, we propose a novel decentralized scalable learning framework, Federated Knowledge Graphs Embedding (FKGE), where embeddings from different knowledge graphs can be learnt in an asynchronous and peer-to-peer manner while being privacy-preserving. FKGE exploits adversarial generation between pairs of knowledge graphs to translate identical entities and relations of different domains into near embedding spaces. In order to protect the privacy of the training data, FKGE further implements a privacy-preserving neural network structure to guarantee no raw data leakage. We conduct extensive experiments to evaluate FKGE on 11 knowledge graphs, demonstrating a significant and consistent improvement in model quality with at most 17.85% and 7.90% increases in performance on triple classification and link prediction tasks. Hao Peng 0001, Haoran Li 0003, Yangqiu Song, Vincent Wenchen Zheng, Jianxin Li 0002 |
CIKM | 4 |
| 2021 | Neural PathSim for Inductive Similarity Search in Heterogeneous Information NetworksabstractPathSim is a widely used meta-path-based similarity in heterogeneous information networks. Numerous applications rely on the computation of PathSim, including similarity search and clustering. Computing PathSim scores on large graphs is computationally challenging due to its high time and storage complexity. In this paper, we propose to transform the problem of approximating the ground truth PathSim scores into a learning problem. We design an encoder-decoder based framework, NeuPath, where the algorithmic structure of PathSim is considered. Specifically, the encoder module identifies Top T optimized path instances, which can approximate the ground truth PathSim, and maps each path instance to an embedding vector. The decoder transforms each embedding vector into a scalar respectively, which identifies the similarity score. We perform extensive experiments on two real-world datasets in different domains, ACM and IMDB. Our results demonstrate that NeuPath performs better than state-of-the-art baselines in the PathSim approximation task and similarity search task. Wenyi Xiao, Huan Zhao 0002, Vincent Wenchen Zheng, Yangqiu Song |
CIKM | 3 |
| 2021 | Social explorative attention based recommendation for content distribution platforms
Wenyi Xiao, Huan Zhao 0002, Haojie Pan, Yangqiu Song, Vincent Wenchen Zheng, Qiang Yang 0001 |
Data Min. Knowl. Discov. | 5 |
| 2021 | Accelerating Large-Scale Heterogeneous Interaction Graph Embedding Learning via Importance SamplingabstractIn real-world problems, heterogeneous entities are often related to each other through multiple interactions, forming a Heterogeneous Interaction Graph (HIG). While modeling HIGs to deal with fundamental tasks, graph neural networks present an attractive opportunity that can make full use of the heterogeneity and rich semantic information by aggregating and propagating information from different types of neighborhoods. However, learning on such complex graphs, often with millions or billions of nodes, edges, and various attributes, could suffer from expensive time cost and high memory consumption. In this article, we attempt to accelerate representation learning on large-scale HIGs by adopting the importance sampling of heterogeneous neighborhoods in a batch-wise manner, which naturally fits with most batch-based optimizations. Distinct from traditional homogeneous strategies neglecting semantic types of nodes and edges, to handle the rich heterogeneous semantics within HIGs, we devise both type-dependent and type-fusion samplers where the former respectively samples neighborhoods of each type and the latter jointly samples from candidates of all types. Furthermore, to overcome the imbalance between the down-sampled and the original information, we respectively propose heterogeneous estimators including the self-normalized and the adaptive estimators to improve the robustness of our sampling strategies. Finally, we evaluate the performance of our models for node classification and link prediction on five real-world datasets, respectively. The empirical results demonstrate that our approach performs significantly better than other state-of-the-art alternatives, and is able to reduce the number of edges in computation by up to 93%, the memory cost by up to 92% and the time cost by up to 86%. Yugang Ji, Mingyang Yin, Hongxia Yang, Jingren Zhou 0001, Vincent Wenchen Zheng, Chuan Shi 0001, Yuan Fang 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2021 | Metagraph-Based Learning on Heterogeneous GraphsabstractData in the form of graphs are prevalent, ranging from biological and social networks to citation graphs and the Web. In particular, most real-world graphs are heterogeneous, containing objects of multiple types, which present new opportunities for many problems on graphs. Consider a typical proximity search problem on graphs, which boils down to measuring the proximity between two given nodes. Most earlier studies on homogeneous or bipartite graphs only measure a generic form of proximity, without accounting for different “semantic classes”-for instance, on a social network two users can be close for different reasons, such as being classmates or family members, which represent two distinct semantic classes. Learning these semantic classes are made possible on heterogeneous graphs through the concept of metagraphs. In this study, we identify metagraphs as a novel and effective means to characterize the common structures for a desired class of proximity. Subsequently, we propose a family of metagraph-based proximity, and employ a learning-to-rank technique that automatically learns the right parameters to suit the desired semantic class. In terms of efficiency, we develop a symmetry-based matching algorithm to speed up the computation of metagraph instances. Empirically, extensive experiments reveal that our metagraph-based proximity substantially outperforms the best competitor by more than 10 percent, and our matching algorithm can reduce matching time by more than half. As a further generalization, we aim to derive a general node and edge representation for heterogeneous graphs, in order to support arbitrary machine learning tasks beyond proximity search. In particular, we propose the finer-grained anchored metagraph, which is capable of discriminating the roles of nodes within the same metagraph. Finally, further experiments on the general representation show that we can outperform the state of the art significantly and consistently across various machine learning tasks. Yuan Fang 0001, Wenqing Lin, Vincent Wenchen Zheng, Min Wu 0008, Kevin Chen-Chuan Chang, Xiaoli Li 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | A Federated Recommender System for Online ServicesabstractDue to privacy and security constraints, directly sharing user data between parties is undesired. Such decentralized data silo issues commonly exist in recommender systems. In general, recommender systems are data-driven. The more data it uses, the better performance it obtains. The data silo issues is a severe limitation of the recommender’s performance. Federated learning is an emerging technology, which bridges the data silos and builds machine learning models without compromising user privacy and data security. We design a recommender system based on federated learning. It is known as the federated recommender system. The system implements plenty of popular algorithms to support various online recommendation services. The algorithm implementation is open-sourced. We also deploy the system on a real-world content recommendation application, achieving significant performance improvement. In this demonstration, we present the architecture of the federated recommender system and give an online demo to show its detailed working procedures and results in content recommendations. Ben Tan, Bo Liu 0015, Vincent Wenchen Zheng, Qiang Yang 0001 |
RecSys | 3 |
| 2020 | Vertex-reinforced Random Walk for Network EmbeddingabstractIn this paper, we study the fundamental problem of random walk for network embedding. We propose to use non-Markovian random walk, variants of vertex-reinforced random walk (VRRW), to fully use the history of a random walk path. To solve the getting stuck problem of VRRW, we introduce an exploitation-exploration mechanism to help the random walk jump out of the stuck set. The new random walk algorithms share the same convergence property of VRRW and thus can be used to learn stable network embeddings. Experimental results on two link prediction benchmark datasets and three node classification benchmark datasets show that our proposed approach reinforce2vec can outperform state-of-the-art random walk based embedding methods by a large margin. Wenyi Xiao, Huan Zhao 0002, Vincent Wenchen Zheng, Yangqiu Song |
SDM | 3 |
| 2020 | Adam revisited: a weighted past gradients perspective
Zaiyi Chen, Chuan Qin 0002, Zai Huang, Vincent Wenchen Zheng, Tong Xu 0001, Enhong Chen |
Frontiers Comput. Sci. | 5 |
| 2019 | Alpha-Beta Sampling for Pairwise Ranking in One-Class Collaborative FilteringabstractThis paper introduces Alpha-Beta Sampling (ABS) strategy, which is particularly intended for the sampling problem of pairwise ranking in one-class collaborative filtering (PROCCF). Specifically, ABS strategy places more emphasis on such training examples, including positive item with a lower preference score and negative items with a higher preference score for each gradient step. Then, we provide the corresponding proofs for the ABS strategy from both gradient and ranking perspectives. First, we prove that sampled training examples by ABS strategy can update the model parameters with a large magnitude and analyze two instantiations by combining two specific pairwise algorithms. Second, it can be proved that ABS strategy is equivalent to optimizing for ranking-aware evaluation metrics like Normalized Discounted Cumulative Gain (NDCG). Furthermore, ABS strategy can be very general and applicable in a lot of pairwise structures of pairwise algorithms. Based on ABS strategy, we provide an effective sampling algorithm to dynamically draw items for each SGD update. Finally, we evaluate the ABS strategy by conducting sampling tasks in two representative pairwise algorithms. The experiment results show that the ABS strategy performs significantly better than the baseline strategies. Mingyue Cheng 0004, Runlong Yu, Qi Liu 0003, Vincent Wenchen Zheng, Hongke Zhao, Hefu Zhang, Enhong Chen |
ICDM | 4 |
| 2019 | Explainable Fashion Recommendation: A Semantic Attribute Region Guided ApproachabstractIn fashion recommender systems, each product usually consists of multiple semantic attributes (e.g., sleeves, collar, etc). When making cloth decisions, people usually show preferences for different semantic attributes (e.g., the clothes with v-neck collar). Nevertheless, most previous fashion recommendation models comprehend the clothing images with a global content representation and lack detailed understanding of users' semantic preferences, which usually leads to inferior recommendation performance. To bridge this gap, we propose a novel Semantic Attribute Explainable Recommender System (SAERS). Specifically, we first introduce a fine-grained interpretable semantic space. We then develop a Semantic Extraction Network (SEN) and Fine-grained Preferences Attention (FPA) module to project users and items into this space, respectively. With SAERS, we are capable of not only providing cloth recommendations for users, but also explaining the reason why we recommend the cloth through intuitive visual attribute semantic highlights in a personalized manner. Extensive experiments conducted on real-world datasets clearly demonstrate the effectiveness of our approach compared with the state-of-the-art methods. Min Hou 0004, Le Wu 0001, Enhong Chen, Zhi Li 0057, Vincent Wenchen Zheng, Qi Liu 0003 |
IJCAI | 5 |
| 2019 | An Attention-Based Model for Learning Dynamic Interaction NetworksabstractIn the physical world, complex systems are generally created as the composition of multiple primitive components that interact with each other rather than a single monolithic structure. Recently, spatio-temporal graphs received a reasonable amount of attention from the research community since they emerged as a natural representational tool able to capture the interactive and interrelated structure of a complex problem. To better understand the nature of complex systems, there is the need to define models that can easily explain the learned causal relationship. To this end, we propose an attentive model able to learn and project the relational structure into a fixed-size embedding. Such representation naturally captures the dynamic influence that each neighbors exert over a given vertex providing a valuable description of the problem setting. The proposed architecture has been extensively evaluated against strong baselines on toy as well as real-world tasks, such as prediction of household energy load and traffic congestion. Sandro Cavallari, Soujanya Poria, Erik Cambria, Vincent Wenchen Zheng, Hongyun Cai 0001 |
IJCNN | 4 |
| 2019 | Hidden POI Ranking with Spatial CrowdsourcingabstractExploring Hidden Points of Interest (H-POIs), which are rarely referred in online search and recommendation systems due to insufficient check-in records, benefits business and individuals. In this work, we investigate how to eliminate the hidden feature of H-POIs by enhancing conventional crowdsourced ranking aggregation framework with heterogeneous (i.e., H-POI and Popular Point of Interest (P-POI)) pairwise tasks. We propose a two-phase solution focusing on both effectiveness and efficiency. In offline phase, we substantially narrow down the search space by retrieving a set of geo-textual valid heterogeneous pairs as the initial candidates and develop two practical data-driven strategies to compute worker qualities. In the online phase, we minimize the cost of assessment by introducing an active learning algorithm to jointly select pairs and workers with worker quality, uncertainty of P-POI rankings and uncertainty of the model taken into account. In addition, a (Minimum Spanning) Tree-constrained Skip search strategy is proposed for the purpose of reducing search time cost. Empirical experiments based on real POI datasets verify that the ranking accuracy of H-POIs can be greatly improved with small number of query iterations. Yue Cui 0001, Liwei Deng 0001, Yan Zhao 0008, Bin Yao 0002, Vincent Wenchen Zheng, Kai Zheng 0001 |
KDD | 5 |
| 2019 | Beyond Personalization: Social Content Recommendation for Creator Equality and Consumer SatisfactionabstractAn effective content recommendation in modern social media platforms should benefit both creators to bring genuine benefits to them and consumers to help them get really interesting content. In this paper, we propose a model called Social Explorative Attention Network (SEAN) for content recommendation. SEAN uses a personalized content recommendation model to encourage personal interests driven recommendation. Moreover, SEAN allows the personalization factors to attend to users' higher-order friends on the social network to improve the accuracy and diversity of recommendation results. Constructing two datasets from a popular decentralized content distribution platform, Steemit, we compare SEAN with state-of-the-art CF and content based recommendation approaches. Experimental results demonstrate the effectiveness of SEAN in terms of both Gini coefficients for recommendation equality and F1 scores for recommendation performance. Wenyi Xiao, Huan Zhao 0002, Haojie Pan, Yangqiu Song, Vincent Wenchen Zheng, Qiang Yang 0001 |
KDD | 5 |
| 2019 | A Survey of Zero-Shot Learning: Settings, Methods, and ApplicationsabstractMost machine-learning methods focus on classifying instances whose classes have already been seen in training. In practice, many applications require classifying instances whose classes have not been seen previously. Zero-shot learning is a powerful and promising learning paradigm, in which the classes covered by training instances and the classes we aim to classify are disjoint. In this paper, we provide a comprehensive survey of zero-shot learning. First of all, we provide an overview of zero-shot learning. According to the data utilized in model optimization, we classify zero-shot learning into three learning settings. Second, we describe different semantic spaces adopted in existing zero-shot learning works. Third, we categorize existing zero-shot learning methods and introduce representative methods under each category. Fourth, we discuss different applications of zero-shot learning. Finally, we highlight promising future research directions of zero-shot learning. Wei Wang 0272, Vincent Wenchen Zheng, Han Yu 0001, Chunyan Miao |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2018 | Distance-Aware DAG Embedding for Proximity Search on Heterogeneous GraphsabstractProximity search on heterogeneous graphs aims to measure the proximity between two nodes on a graph w.r.t. some semantic relation for ranking. Pioneer work often tries to measure such proximity by paths connecting the two nodes. However, paths as linear sequences have limited expressiveness for the complex network connections. In this paper, we explore a more expressive DAG (directed acyclic graph) data structure for modeling the connections between two nodes. Particularly, we are interested in learning a representation for the DAGs to encode the proximity between two nodes. We face two challenges to use DAGs, including how to efficiently generate DAGs and how to effectively learn DAG embedding for proximity search. We find distance-awareness as important for proximity search and the key to solve the above challenges. Thus we develop a novel Distance-aware DAG Embedding (D2AGE) model. We evaluate D2AGE on three benchmark data sets with six semantic relations, and we show that D2AGE outperforms the state-of-the-art baselines. We release the code on https://github.com/shuaiOKshuai. Vincent Wenchen Zheng, Zhou Zhao 0001, Fanwei Zhu, Kevin Chen-Chuan Chang, Minghui Wu 0001, Jing Ying |
AAAI | 2 |
| 2018 | Node, Motif and Subgraph: Leveraging Network Functional Blocks Through Structural ConvolutionabstractNetworks or graphs provide a natural and generic way for modeling rich structured data. Recent research on graph analysis has been focused on representation learning, of which the goal is to encode the network structures into distributed embedding vectors, so as to enable various downstream applications through off-the-shelf machine learning. However, existing methods mostly focus on node-level embedding, which is insufficient for subgraph analysis. Moreover, their leverage of network structures through path sampling or neighborhood preserving is implicit and coarse. Network motifs allow graph analysis in a finer granularity, but existing methods based on motif matching are limited to enumerated simple motifs and do not leverage node labels and supervision. In this paper, we develop NEST, a novel hierarchical network embedding method combining motif filtering and convolutional neural networks. Motif-based filtering enables NEST to capture exact small structures within networks, and convolution over the filtered embedding allows it to fully explore complex substructures and their combinations. NEST can be trivially applied to any domain and provide insight into particular network functional blocks. Extensive experiments on protein function prediction, drug toxicity prediction and social network community identification have demonstrated its effectiveness and efficiency. Carl Yang 0001, Mengxiong Liu, Vincent Wenchen Zheng, Jiawei Han 0001 |
ASONAM | 3 |
| 2018 | Prerequisite-Driven Deep Knowledge TracingabstractKnowledge tracing serves as the key technique in the computer supported education environment (e.g., intelligent tutoring systems) to model student's knowledge states. While the Bayesian knowledge tracing and deep knowledge tracing models have been developed, the sparseness of student's exercise data still limits knowledge tracing's performance and applications. In order to address this issue, we advocate for and propose to incorporate the knowledge structure information, especially the prerequisite relations between pedagogical concepts, into the knowledge tracing model. Specifically, by considering how students master pedagogical concepts and their prerequisites, we model prerequisite concept pairs as ordering pairs. With a proper mathematical formulation, this property can be utilized as constraints in designing knowledge tracing model. As a result, the obtained model can have a better performance on student concept mastery prediction. In order to evaluate this model, we test it on five different real world datasets, and the experimental results show that the proposed model achieves a significant performance improvement by comparing with three knowledge tracing models. Penghe Chen, Yu Lu 0003, Vincent Wenchen Zheng, Yang Pian |
ICDM | 3 |
| 2018 | Heterogeneous Embedding Propagation for Large-Scale E-Commerce User AlignmentabstractWe study the important problem of user alignment in e-commerce: to predict whether two online user identities that access an e-commerce site from different devices belong to one real-world person. As input, we have a set of user activity logs from Taobao and some labeled user identity linkages. User activity logs can be modeled using a heterogeneous interaction graph (HIG), and subsequently the user alignment task can be formulated as a semi-supervised HIG embedding problem. HIG embedding is challenging for two reasons: its heterogeneous nature and the presence of edge features. To address the challenges, we propose a novel Heterogeneous Embedding Propagation (HEP) model. The core idea is to iteratively reconstruct a node's embedding from its heterogeneous neighbors in a weighted manner, and meanwhile propagate its embedding updates from reconstruction loss and/or classification loss to its neighbors. We conduct extensive experiments on large-scale datasets from Taobao, demonstrating that HEP significantly outperforms state-of-the-art baselines often by more than 10% in F-scores. Vincent Wenchen Zheng, Mo Sha 0002, Yuchen Li 0001, Hongxia Yang, Yuan Fang 0001, Kian-Lee Tan, Kevin Chen-Chuan Chang |
ICDM | 1 |
| 2018 | Engineering Graph Features via Network Functional BlocksabstractGraph is a prevalent data structure that enables many predictive tasks. How to engineer graph features is a fundamental question. Our concept is to go beyond nodes and edges, and explore richer structures (e.g., paths, subgraphs) for graph feature engineering. We call such richer structures as network functional blocks, because each structure serves as a network building block but with some different functionality. We use semantic proximity search as an example application to share our recent work on exploiting different granularities of network functional blocks. We show that network functional blocks are effective, and they can be useful for a wide range of applications. Vincent Wenchen Zheng |
IJCAI | 1 |
| 2018 | High-order Proximity Preserving Information Network HashingabstractInformation network embedding is an effective way for efficient graph analytics. However, it still faces with computational challenges in problems such as link prediction and node recommendation, particularly with increasing scale of networks. Hashing is a promising approach for accelerating these problems by orders of magnitude. However, no prior studies have been focused on seeking binary codes for information networks to preserve high-order proximity. Since matrix factorization (MF) unifies and outperforms several well-known embedding methods with high-order proximity preserved, we propose a MF-based \underlineI nformation \underlineN etwork \underlineH ashing (INH-MF) algorithm, to learn binary codes which can preserve high-order proximity. We also suggest Hamming subspace learning, which only updates partial binary codes each time, to scale up INH-MF. We finally evaluate INH-MF on four real-world information network datasets with respect to the tasks of node classification and node recommendation. The results demonstrate that INH-MF can perform significantly better than competing learning to hash baselines in both tasks, and surprisingly outperforms network embedding methods, including DeepWalk, LINE and NetMF, in the task of node recommendation. The source code of INH-MF is available online\footnote\urlhttps://github.com/DefuLian/network . Defu Lian, Kai Zheng 0001, Vincent Wenchen Zheng, Yong Ge 0001, Longbing Cao, Ivor W. Tsang, Xing Xie 0001 |
KDD | 3 |
| 2018 | Interactive Paths Embedding for Semantic Proximity Search on Heterogeneous GraphsabstractSemantic proximity search on heterogeneous graph is an important task, and is useful for many applications. It aims to measure the proximity between two nodes on a heterogeneous graph w.r.t. some given semantic relation. Prior work often tries to measure the semantic proximity by paths connecting a query object and a target object. Despite the success of such path-based approaches, they often modeled the paths in a weakly coupled manner, which overlooked the rich interactions among paths. In this paper, we introduce a novel concept of interactive paths to model the inter-dependency among multiple paths between a query object and a target object. We then propose an Interactive Paths Embedding (IPE) model, which learns low-dimensional representations for the resulting interactive-paths structures for proximity estimation. We conduct experiments on seven relations with four different types of heterogeneous graphs, and show that our model outperforms the state-of-the-art baselines. Vincent Wenchen Zheng, Zhou Zhao 0001, Zhao Li 0007, Hongxia Yang, Minghui Wu 0001, Jing Ying |
KDD | 2 |
| 2018 | An automatic knowledge graph construction system for K-12 educationabstractMotivated by the pressing need of educational applications with knowledge graph, we develop a system, called K12EduKG, to automatically construct knowledge graphs for K-12 educational subjects. Leveraging on heterogeneous domain-specific educational data, K12EduKG extracts educational concepts and identifies implicit relations with high educational significance. More specifically, it adopts named entity recognition (NER) techniques on educational data like curriculum standards to extract educational concepts, and employs data mining techniques to identify the cognitive prerequisite relations between educational concepts. In this paper, we present details of K12EduKG and demonstrate it with a knowledge graph constructed for the subject of mathematics. Penghe Chen, Yu Lu 0003, Vincent Wenchen Zheng, Xiyang Chen |
L@S | 3 |
| 2018 | Adaptive Attention Network for Review Sentiment Classification
Chuantao Zong, Wenfeng Feng 0001, Vincent Wenchen Zheng, Hankui Zhuo |
PAKDD (1) | 3 |
| 2018 | Robust Asymmetric Recommendation via Min-Max OptimizationabstractRecommender systems with implicit feedback (e.g. clicks and purchases) suffer from two critical limitations: 1) imbalanced labels may mislead the learning process of the conventional models that assign balanced weights to the classes; and 2) outliers with large reconstruction errors may dominate the objective function by the conventional $L_2$-norm loss. To address these issues, we propose a robust asymmetric recommendation model. It integrates cost-sensitive learning with capped unilateral loss into a joint objective function, which can be optimized by an iteratively weighted approach. To reduce the computational cost of low-rank approximation, we exploit the dual characterization of the nuclear norm to derive a min-max optimization problem and design a subgradient algorithm without performing full SVD. Finally, promising empirical results demonstrate the effectiveness of our algorithm on benchmark recommendation datasets. Peng Yang 0010, Peilin Zhao, Vincent Wenchen Zheng, Lizhong Ding 0001, Xin Gao 0001 |
SIGIR | 3 |
| 2018 | Unsupervised Multi-view Nonlinear Graph Embedding
Zhao Li 0007, Vincent Wenchen Zheng, Yifan Yang 0001, Yuanmi Chen |
UAI | 3 |
| 2018 | Subgraph-augmented Path Embedding for Semantic User Search on Heterogeneous Social NetworkabstractSemantic user search is an important task on heterogeneous social networks. Its core problem is to measure the proximity between two user objects in the network w.r.t. certain semantic user relation. State-of-the-art solutions often take a path-based approach, which uses the sequences of objects connecting a query user and a target user to measure their proximity. Despite their success, we assert that path as a low-order structure is insufficient to capture the rich semantics between two users. Therefore, in this paper we introduce a new concept of subgraph-augmented path for semantic user search. Specifically, we consider sampling a set of object paths from a query user to a target user; then in each object path, we replace the linear object sequence between its every two neighboring users with their shared subgraph instances. Such subgraph-augmented paths are expected to leverage both path»s distance awareness and subgraph»s high-order structure. As it is non-trivial to model such subgraph-augmented paths, we develop a Subgraph-augmented Path Embedding (SPE) framework to accomplish the task. We evaluate our solution on six semantic user relations in three real-world public data sets, and show that it outperforms the baselines. Vincent Wenchen Zheng, Zhou Zhao 0001, Hongxia Yang, Kevin Chen-Chuan Chang, Minghui Wu 0001, Jing Ying |
WWW | 2 |
| 2018 | A Comprehensive Survey of Graph Embedding: Problems, Techniques, and ApplicationsabstractGraph is an important data representation which appears in a wide diversity of real-world scenarios. Effective graph analytics provides users a deeper understanding of what is behind the data, and thus can benefit a lot of useful applications such as node classification, node recommendation, link prediction, etc. However, most graph analytics methods suffer the high computation and space cost. Graph embedding is an effective yet efficient way to solve the graph analytics problem. It converts the graph data into a low dimensional space in which the graph structural information and graph properties are maximumly preserved. In this survey, we conduct a comprehensive review of the literature in graph embedding. We first introduce the formal definition of graph embedding as well as the related concepts. After that, we propose two taxonomies of graph embedding which correspond to what challenges exist in different graph embedding problem settings and how the existing work addresses these challenges in their solutions. Finally, we summarize the applications that graph embedding enables and suggest four promising future research directions in terms of computation efficiency, problem settings, techniques, and application scenarios. Hongyun Cai 0001, Vincent Wenchen Zheng, Kevin Chen-Chuan Chang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | Semantic Proximity Search on Heterogeneous Graph by Proximity EmbeddingabstractMany real-world networks have a rich collection of objects. The semantics of these objects allows us to capture different classes of proximities, thus enabling an important task of semantic proximity search. As the core of semantic proximity search, we have to measure the proximity on a heterogeneous graph, whose nodes are various types of objects. Most of the existing methods rely on engineering features about the graph structure between two nodes to measure their proximity. With recent development on graph embedding, we see a good chance to avoid feature engineering for semantic proximity search. There is very little work on using graph embedding for semantic proximity search. We also observe that graph embedding methods typically focus on embedding nodes, which is an "indirect'' approach to learn the proximity. Thus, we introduce a new concept of proximity embedding, which directly embeds the network structure between two possibly distant nodes. We also design our proximity embedding, so as to flexibly support both symmetric and asymmetric proximities. Based on the proximity embedding, we can easily estimate the proximity score between two nodes and enable search on the graph. We evaluate our proximity embedding method on three real-world public data sets, and show it outperforms the state-of-the-art baselines. Vincent Wenchen Zheng, Zhou Zhao 0001, Fanwei Zhu, Kevin Chen-Chuan Chang, Minghui Wu 0001, Jing Ying |
AAAI | 2 |
| 2017 | Community-Based Question Answering via Asymmetric Multi-Faceted Ranking Network LearningabstractNowadays the community-based question answering (CQA) sites become the popular Internet-based web service, which have accumulated millions of questions and their posted answers over time. Thus, question answering becomes an essential problem in CQA sites, which ranks the high-quality answers to the given question. Currently, most of the existing works study the problem of question answering based on the deep semantic matching model to rank the answers based on their semantic relevance, while ignoring the authority of answerers to the given question. In this paper, we consider the problem of community-based question answering from the viewpoint of asymmetric multi-faceted ranking network embedding. We propose a novel asymmetric multi-faceted ranking network learning framework for community-based question answering by jointly exploiting the deep semantic relevance between question-answer pairs and the answerers' authority to the given question. We then develop an asymmetric ranking network learning method with deep recurrent neural networks by integrating both answers' relative quality rank to the given question and the answerers' following relations in CQA sites. The extensive experiments on a large-scale dataset from a real world CQA site show that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Hanqing Lu, Vincent Wenchen Zheng, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
AAAI | 3 |
| 2017 | Learning Community Embedding with Community Detection and Node Embedding on GraphsabstractIn this paper, we study an important yet largely under-explored setting of graph embedding, i.e., embedding communities instead of each individual nodes. We find that community embedding is not only useful for community-level applications such as graph visualization, but also beneficial to both community detection and node classification. To learn such embedding, our insight hinges upon a closed loop among community embedding, community detection and node embedding. On the one hand, node embedding can help improve community detection, which outputs good communities for fitting better community embedding. On the other hand, community embedding can be used to optimize the node embedding by introducing a community-aware high-order proximity. Guided by this insight, we propose a novel community embedding framework that jointly solves the three tasks together. We evaluate such a framework on multiple real-world datasets, and show that it improves graph visualization and outperforms state-of-the-art baselines in various application tasks, e.g., community detection and node classification. Sandro Cavallari, Vincent Wenchen Zheng, Hongyun Cai 0001, Kevin Chen-Chuan Chang, Erik Cambria |
CIKM | 2 |
| 2017 | A Simple Regularization-based Algorithm for Learning Cross-Domain Word EmbeddingsabstractLearning word embeddings has received a significant amount of attention recently.Often, word embeddings are learned in an unsupervised manner from a large collection of text.The genre of the text typically plays an important role in the effectiveness of the resulting embeddings.How to effectively train word embedding models using data from different domains remains a problem that is underexplored.In this paper, we present a simple yet effective method for learning word embeddings based on text from different domains.We demonstrate the effectiveness of our approach through extensive experiments on various down-stream NLP tasks. Wei Yang 0017, Wei Lu 0011, Vincent Wenchen Zheng |
EMNLP | 3 |
| 2017 | SocialLens: Searching and Browsing Communities by Content and InteractionabstractCommunity analysis is an important task in graph mining. Most of the existing community studies are community detection, which aim to find the community membership for each user based on the user friendship links. However, membership alone, without a complete profile of what a community is and how it interacts with other communities, has limited applications. This motivates us to consider systematically profiling the communities and thereby developing useful community-level applications. In this paper, we introduce a novel concept of community profiling, upon which we build a SocialLens system1 to enable searching and browsing communities by content and interaction. We deploy SocialLens on two social graphs: Twitter and DBLP. We demonstrate two useful applications of SocialLens, including interactive community visualization and profile-aware community ranking. Hongyun Cai 0001, Vincent Wenchen Zheng, Penghe Chen, Fanwei Zhu, Kevin Chen-Chuan Chang, Zi Huang |
ICDE | 2 |
| 2017 | Topological Recurrent Neural Network for Diffusion PredictionabstractIn this paper, we study the problem of using representation learning to assist information diffusion prediction on graphs. In particular, we aim at estimating the probability of an inactive node to be activated next in a cascade. Despite the success of recent deep learning methods for diffusion, we find that they often underexplore the cascade structure. We consider a cascade as not merely a sequence of nodes ordered by their activation time stamps; instead, it has a richer structure indicating the diffusion process over the data graph. As a result, we introduce a new data model, namely diffusion topologies, to fully describe the cascade structure. We find it challenging to model diffusion topologies, which are dynamic directed acyclic graphs (DAGs), with the existing neural networks. Therefore, we propose a novel topological recurrent neural network, namely Topo-LSTM, for modeling dynamic DAGs. We customize Topo-LSTM for the diffusion prediction task, and show it improves the state-of-the-art baselines, by 20.1%-56.6% (MAP) relatively, across multiple real-world data sets. Vincent Wenchen Zheng, Kevin Chen-Chuan Chang |
ICDM | 2 |
| 2017 | Link Prediction via Ranking Metric Dual-Level Attention Network LearningabstractLink prediction is a challenging problem for complex network analysis, arising in many disciplines such as social networks and telecommunication networks. Currently, many existing approaches estimate the proximity of the link endpoints for link prediction from their feature or the local neighborhood around them, which suffer from the localized view of network connections and insufficiency of discriminative feature representation. In this paper, we consider the problem of link prediction from the viewpoint of learning discriminative path-based proximity ranking metric embedding. We propose a novel ranking metric network learning framework by jointly exploiting both node-level and path-level attentional proximity of the endpoints for link prediction. We then develop the path-based dual-level reasoning attentional learning method with recurrent neural network for proximity ranking metric embedding. The extensive experiments on two large-scale datasets show that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Ben Gao, Vincent Wenchen Zheng, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 3 |
| 2017 | From Community Detection to Community ProfilingabstractMost existing community-related studies focus on detection, which aim to find the community membership for each user from user friendship links. However, membership alone, without a complete profile of what a community is and how it interacts with other communities, has limited applications. This motivates us to consider systematically profiling the communities and thereby developing useful community-level applications. In this paper, we for the first time formalize the concept of community profiling. With rich user information on the network, such as user published content and user diffusion links, we characterize a community in terms of both its internal content profile and external diffusion profile. The difficulty of community profiling is often underestimated. We novelly identify three unique challenges and propose a joint Community Profiling and Detection (CPD) model to address them accordingly. We also contribute a scalable inference algorithm, which scales linearly with the data size and it is easily parallelizable. We evaluate CPD on large-scale real-world data sets, and show that it is significantly better than the state-of-the-art baselines in various tasks. Hongyun Cai 0001, Vincent Wenchen Zheng, Fanwei Zhu, Kevin Chen-Chuan Chang, Zi Huang |
Proc. VLDB Endow. | 2 |
| 2016 | Cold-Start Heterogeneous-Device Wireless LocalizationabstractIn this paper, we study a cold-start heterogeneous-devicelocalization problem. This problem is challenging, becauseit results in an extreme inductive transfer learning setting,where there is only source domain data but no target do-main data. This problem is also underexplored. As there is notarget domain data for calibration, we aim to learn a robustfeature representation only from the source domain. There islittle previous work on such a robust feature learning task; besides, the existing robust feature representation propos-als are both heuristic and inexpressive. As our contribution,we for the first time provide a principled and expressive robust feature representation to solve the challenging cold-startheterogeneous-device localization problem. We evaluate ourmodel on two public real-world data sets, and show that itsignificantly outperforms the best baseline by 23.1%–91.3%across four pairs of heterogeneous devices. Vincent Wenchen Zheng, Shenghua Gao, Aditi Adhikari, Kevin Chen-Chuan Chang |
AAAI | 1 |
| 2016 | Regularizing Structured Classifier with Conditional Probabilistic Constraints for Semi-supervised LearningabstractConstraints have been shown as an effective way to incorporate unlabeled data for semi-supervised structured classification. We recognize that, constraints are often conditional and probabilistic; moreover, a constraint can have its condition depend on either just observations (which we call x-type constraint) or even hidden variables (which we call y-type constraint). We wish to design a constraint formulation that can flexibly model the constraint probability for both x-type and y-type constraints, and later use it to regularize general structured classifiers for semi-supervision. Surprisingly, none of the existing models have such a constraint formulation. Thus in this paper, we propose a new conditional probabilistic formulation for modeling both x-type and y-type constraints. We also recognize the inference complication for y-type constraint, and propose a systematic selective evaluation approach to efficiently realize the constraints. Finally, we evaluate our model in three applications, including named entity recognition, part-of-speech tagging and entity information extraction, with totally nine data sets. We show that our model is generally more accurate and efficient than the state-of-the-art baselines. Our code and data are available at https://bitbucket.org/vwz/cikm2016-cpf/. Vincent Wenchen Zheng, Kevin Chen-Chuan Chang |
CIKM | 1 |
| 2016 | Semantic proximity search on graphs with metagraph-based learningabstractGiven ubiquitous graph data such as the Web and social networks, proximity search on graphs has been an active research topic. The task boils down to measuring the proximity between two nodes on a graph. Although most earlier studies deal with homogeneous or bipartite graphs only, many real-world graphs are heterogeneous with objects of various types, giving rise to different semantic classes of proximity. For instance, on a social network two users can be close for different reasons, such as being classmates or family members, which represent two distinct classes of proximity. Thus, it becomes inadequate to only measure a “generic” form of proximity as previous works have focused on. In this paper, we identify metagraphs as a novel and effective means to characterize the common structures for a desired class of proximity. Subsequently, we propose a family of metagraph-based proximity, and employ a supervised technique to automatically learn the right form of proximity within its family to suit the desired class. As it is expensive to match (i.e., find the instances of) a metagraph, we propose the novel approaches of dual-stage training and symmetry-based matching to speed up. Finally, our experiments reveal that our approach is significantly more accurate and efficient. For accuracy, we outperform the baselines by 11% and 16% in NDCG and MAP, respectively. For efficiency, dual-stage training reduces the overall matching cost by 83%, and symmetry-based matching further decreases the cost of individual metagraphs by 52%. Yuan Fang 0001, Wenqing Lin, Vincent Wenchen Zheng, Min Wu 0008, Kevin Chen-Chuan Chang, Xiaoli Li 0001 |
ICDE | 3 |
| 2016 | Learning to query: Focused web page harvesting for entity aspectsabstractAs the Web hosts rich information about real-world entities, our information quests become increasingly entity centric. In this paper, we study the problem of focused harvesting of Web pages for entity aspects, to support downstream applications such as business analytics and building a vertical portal. Given that search engines are the de facto gateways to assess information on the Web, we recognize the essence of our problem as Learning to Query (L2Q) - to intelligently select queries so that we can harvest pages, via a search engine, focused on an entity aspect of interest. Thus, it is crucial to quantify the utilities of the candidate queries w.r.t. some entity aspect. In order to better estimate the utilities, we identify two opportunities and address their challenges. First, a target entity in a given domain has many peers. We leverage these peer entities to become domain aware. Second, a candidate query may “overlap” with the past queries that have already been fired. We account for these past queries to become context aware. Empirical results show that our approach significantly outperforms both algorithmic and manual baselines by 16% and 10% in F-scores, respectively. Yuan Fang 0001, Vincent Wenchen Zheng, Kevin Chen-Chuan Chang |
ICDE | 2 |
| 2016 | Detecting signals of detrimental prescribing cascades from social media
Tao Hoang, Jixue Liu, Nicole Pratt, Vincent Wenchen Zheng, Kevin Chen-Chuan Chang, Elizabeth Roughead, Jiuyong Li |
Artif. Intell. Medicine | 4 |
| 2015 | An Aggressive Graph-Based Selective Sampling Algorithm for ClassificationabstractTraditional online learning algorithms are designed for vector data only, which assume that the labels of all the training examples are provided. In this paper, we study graph classification where only limited nodes are chosen for labelling by selective sampling. Particularly, we first adapt a spectral-based graph regularization technique to derive a novel online learning linear algorithm which can handle graph data, although it still queries the labels of all nodes and thus is not preferred, as labelling is typically time-consuming. To address this issue, we then propose a new confidence-based query method for selective sampling. The theoretical result shows that our online learning algorithm with a fraction of queried labels can achieve a mistake bound comparable with the one learning on all labels of the nodes. In addition, the algorithm based on our proposed query strategy can achieve a mistake bound better than the one based on other query methods. However, our algorithm is conservative to update the model whenever error happens, which obviously wastes training labels that are valuable for the model. To take advantage of these labels, we further propose a novel aggressive algorithm, which can update the model aggressively even if no error occurs. The theoretical analysis shows that our aggressive approach can achieve a mistake bound better than its conservative and fully-supervised counterpart, with substantially fewer queried times. We empirically evaluate our algorithm on several real-world graph datasets and the experimental results demonstrate that our method is highly effective. Peng Yang 0010, Peilin Zhao, Vincent Wenchen Zheng, Xiaoli Li 0001 |
ICDM | 3 |
| 2015 | Mobility Profiling for User Verification with Anonymized Location Data
Vincent Wenchen Zheng, Kevin Chen-Chuan Chang, Shonali Krishnaswamy |
IJCAI | 3 |
| 2015 | Modeling User Activity Preference by Leveraging User Spatial Temporal Characteristics in LBSNsabstractWith the recent surge of location based social networks (LBSNs), activity data of millions of users has become attainable. This data contains not only spatial and temporal stamps of user activity, but also its semantic information. LBSNs can help to understand mobile users' spatial temporal activity preference (STAP), which can enable a wide range of ubiquitous applications, such as personalized context-aware location recommendation and group-oriented advertisement. However, modeling such user-specific STAP needs to tackle high-dimensional data, i.e., user-location-time-activity quadruples, which is complicated and usually suffers from a data sparsity problem. In order to address this problem, we propose a STAP model. It first models the spatial and temporal activity preference separately, and then uses a principle way to combine them for preference inference. In order to characterize the impact of spatial features on user activity preference, we propose the notion of personal functional region and related parameters to model and infer user spatial activity preference. In order to model the user temporal activity preference with sparse user activity data in LBSNs, we propose to exploit the temporal activity similarity among different users and apply nonnegative tensor factorization to collaboratively infer temporal activity preference. Finally, we put forward a context-aware fusion framework to combine the spatial and temporal activity preference models for preference inference. We evaluate our proposed approach on three real-world datasets collected from New York and Tokyo, and show that our STAP model consistently outperforms the baseline approaches in various settings. Dingqi Yang, Daqing Zhang 0001, Vincent Wenchen Zheng, Zhiyong Yu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2013 | Covariate Shift in Hilbert Space: A Solution via Sorrogate KernelsabstractCovariate shift is a unconventional learning scenario in which training and testing data have different distributions. A general principle to solve the problem is to make the training data distribution similar to the test one, such that classifiers computed on the former generalizes well to the latter. Current approaches typically target on the sample distribution in the input space, however, for kernel-based learning methods, the algorithm performance depends directly on the geometry of the kernel-induced feature space. Motivated by this, we propose to match data distributions in the Hilbert space, which, given a pre-defined empirical kernel map, can be formulated as aligning kernel matrices across domains. In particular, to evaluate similarity of kernel matrices defined on arbitrarily different samples, the novel concept of surrogate kernel is introduced based on the Mercer's theorem. Our approach caters the model adaptation specifically to kernel-based learning mechanism, and demonstrates promising results on several real-world applications. Kai Zhang 0001, Vincent Wenchen Zheng, Qiaojun Wang, James T. Kwok, Qiang Yang 0001, Ivan Marsic |
ICML (3) | 2 |
| 2012 | Towards mobile intelligence: Learning from GPS history data for collaborative recommendation
Vincent Wenchen Zheng, Yu Zheng 0004, Xing Xie 0001, Qiang Yang 0001 |
Artif. Intell. | 1 |
| 2012 | Integrated QSAR study for inhibitors of hedgehog signal pathway against multiple cell lines: a collaborative filtering methodabstractBACKGROUND: The Hedgehog Signaling Pathway is one of signaling pathways that are very important to embryonic development. The participation of inhibitors in the Hedgehog Signal Pathway can control cell growth and death, and searching novel inhibitors to the functioning of the pathway are in a great demand. As the matter of fact, effective inhibitors could provide efficient therapies for a wide range of malignancies, and targeting such pathway in cells represents a promising new paradigm for cell growth and death control. Current research mainly focuses on the syntheses of the inhibitors of cyclopamine derivatives, which bind specifically to the Smo protein, and can be used for cancer therapy. While quantitatively structure-activity relationship (QSAR) studies have been performed for these compounds among different cell lines, none of them have achieved acceptable results in the prediction of activity values of new compounds. In this study, we proposed a novel collaborative QSAR model for inhibitors of the Hedgehog Signaling Pathway by integration the information from multiple cell lines. Such a model is expected to substantially improve the QSAR ability from single cell lines, and provide useful clues in developing clinically effective inhibitors and modifications of parent lead compounds for target on the Hedgehog Signaling Pathway. RESULTS: In this study, we have presented: (1) a collaborative QSAR model, which is used to integrate information among multiple cell lines to boost the QSAR results, rather than only a single cell line QSAR modeling. Our experiments have shown that the performance of our model is significantly better than single cell line QSAR methods; and (2) an efficient feature selection strategy under such collaborative environment, which can derive the commonly important features related to the entire given cell lines, while simultaneously showing their specific contributions to a specific cell-line. Based on feature selection results, we have proposed several possible chemical modifications to improve the inhibitor affinity towards multiple targets in the Hedgehog Signaling Pathway. CONCLUSIONS: Our model with the feature selection strategy presented here is efficient, robust, and flexible, and can be easily extended to model large-scale multiple cell line/QSAR data. The data and scripts for collaborative QSAR modeling are available in the Additional file 1. Dongsheng Che, Vincent Wenchen Zheng, Ruixin Zhu, Qi Liu 0019 |
BMC Bioinform. | 3 |
| 2011 | User-Dependent Aspect Model for Collaborative Activity RecognitionabstractActivity recognition aims to discover one or more users' actions and goals based on sensor readings. In the real world, a single user's data are often insufficient for training an activity recognition model due to the data sparsity problem. This is especially true when we are interested in obtaining a personalized model. In this paper, we study how to collaboratively use different users' sensor data to train a model that can provide personalized activity recognition for each user. We propose a user-dependent aspect model for this collaborative activity recognition task. Our model introduces user aspect variables to capture the user grouping information, so that a target user can also benefit from her similar users in the same group to train the recognition model. In this way, we can greatly reduce the need for much valuable and expensive labeled data required in training the recognition model for each user. Our model is also capable of incorporating time information and handling new user in activity recognition. We evaluate our model on a real-world WiFi data set obtained from an indoor environment, and show that the proposed model can outperform several state-of-art baseline algorithms. Vincent Wenchen Zheng, Qiang Yang 0001 |
IJCAI | 1 |
| 2011 | Cross-domain activity recognition via transfer learning
Derek Hao Hu, Vincent Wenchen Zheng, Qiang Yang 0001 |
Pervasive Mob. Comput. | 2 |
| 2010 | Collaborative Filtering Meets Mobile Recommendation: A User-Centered ApproachabstractWith the increasing popularity of location tracking services such as GPS, more and more mobile data are being accumulated. Based on such data, a potentially useful service is to make timely and targeted recommendations for users on places where they might be interested to go and activities that they are likely to conduct. For example, a user arriving in Beijing might wonder where to visit and what she can do around the Forbidden City. A key challenge for such recommendation problems is that the data we have on each individual user might be very limited, while to make useful and accurate recommendations, we need extensive annotated location and activity information from user trace data. In this paper, we present a new approach, known as user-centered collaborative location and activity filtering (UCLAF), to pull many users’ data together and apply collaborative filtering to find like-minded users and like-patterned activities at different locations. We model the userlocation- activity relations with a tensor representation, and propose a regularized tensor and matrix decomposition solution which can better address the sparse data problem in mobile information retrieval. We empirically evaluate UCLAF using a real-world GPS dataset collected from 164 users over 2.5 years, and showed that our system can outperform several state-of-the-art solutions to the problem. Vincent Wenchen Zheng, Bin Cao 0001, Yu Zheng 0004, Xing Xie 0001, Qiang Yang 0001 |
AAAI | 1 |
| 2010 | Indoor localization in multi-floor environments with reduced effortabstractIn pervasive computing, localizing a user in wireless indoor environments is an important yet challenging task. Among the state-of-art localization methods, fingerprinting is shown to be quite successful by statistically learning the signal to location relations. However, a major drawback for fingerprinting is that, it usually requires a lot of labeled data to train an accurate localization model. To establish a fingerprinting-based localization model in a building with many floors, we have to collect sufficient labeled data on each floor. This effort can be very burdensome. In this paper, we study how to reduce this calibration effort by only collecting the labeled data on one floor, while collecting unlabeled data on other floors. Our idea is inspired by the observation that, although the wireless signals can be quite different, the floor-plans in a building are similar. Therefore, if we co-embed these different floors' data in some common low-dimensional manifold, we are able to align the unlabeled data with the labeled data well so that we can then propagate the labels to the unlabeled data. We conduct empirical evaluations on real-world multi-floor data sets to validate our proposed method. Hua-Yan Wang, Vincent Wenchen Zheng, Qiang Yang 0001 |
PerCom | 2 |
| 2010 | Collaborative location and activity recommendations with GPS history dataabstractWith the increasing popularity of location-based services, such as tour guide and location-based social network, we now have accumulated many location data on the Web. In this paper, we show that, by using the location data based on GPS and users' comments at various locations, we can discover interesting locations and possible activities that can be performed there for recommendations. Our research is highlighted in the following location-related queries in our daily life: 1) if we want to do something such as sightseeing or food-hunting in a large city such as Beijing, where should we go? 2) If we have already visited some places such as the Bird's Nest building in Beijing's Olympic park, what else can we do there? By using our system, for the first question, we can recommend her to visit a list of interesting locations such as Tiananmen Square, Bird's Nest, etc. For the second question, if the user visits Bird's Nest, we can recommend her to not only do sightseeing but also to experience its outdoor exercise facilities or try some nice food nearby. To achieve this goal, we first model the users' location and activity histories that we take as input. We then mine knowledge, such as the location features and activity-activity correlations from the geographical databases and the Web, to gather additional inputs. Finally, we apply a collective matrix factorization method to mine interesting locations and activities, and use them to recommend to the users where they can visit if they want to perform some specific activities and what they can do if they visit some specific places. We empirically evaluated our system using a large GPS dataset collected by 162 users over a period of 2.5 years in the real-world. We extensively evaluated our system and showed that our system can outperform several state-of-the-art baselines. Vincent Wenchen Zheng, Yu Zheng 0004, Xing Xie 0001, Qiang Yang 0001 |
WWW | 1 |
| 2010 | Multi-task learning for cross-platform siRNA efficacy prediction: an in-silico studyabstractBACKGROUND: Gene silencing using exogenous small interfering RNAs (siRNAs) is now a widespread molecular tool for gene functional study and new-drug target identification. The key mechanism in this technique is to design efficient siRNAs that incorporated into the RNA-induced silencing complexes (RISC) to bind and interact with the mRNA targets to repress their translations to proteins. Although considerable progress has been made in the computational analysis of siRNA binding efficacy, few joint analysis of different RNAi experiments conducted under different experimental scenarios has been done in research so far, while the joint analysis is an important issue in cross-platform siRNA efficacy prediction. A collective analysis of RNAi mechanisms for different datasets and experimental conditions can often provide new clues on the design of potent siRNAs. RESULTS: An elegant multi-task learning paradigm for cross-platform siRNA efficacy prediction is proposed. Experimental studies were performed on a large dataset of siRNA sequences which encompass several RNAi experiments recently conducted by different research groups. By using our multi-task learning method, the synergy among different experiments is exploited and an efficient multi-task predictor for siRNA efficacy prediction is obtained. The 19 most popular biological features for siRNA according to their jointly importance in multi-task learning were ranked. Furthermore, the hypothesis is validated out that the siRNA binding efficacy on different messenger RNAs(mRNAs) have different conditional distribution, thus the multi-task learning can be conducted by viewing tasks at an "mRNA"-level rather than at the "experiment"-level. Such distribution diversity derived from siRNAs bound to different mRNAs help indicate that the properties of target mRNA have important implications on the siRNA binding efficacy. CONCLUSIONS: The knowledge gained from our study provides useful insights on how to analyze various cross-platform RNAi data for uncovering of their complex mechanism. Qi Liu 0019, Qian Xu 0005, Vincent Wenchen Zheng, Hong Xue 0001, Qiang Yang 0001 |
BMC Bioinform. | 3 |
| 2009 | Cross-domain activity recognitionabstractIn activity recognition, one major challenge is huge manual effort in labeling when a new domain of activities is to be tested. In this paper, we ask an interesting question: can we transfer the available labeled data from a set of existing activities in one domain to help recognize the activities in another different but related domain? Our answer is "yes", provided that the sensor data from the two domains are related in some way. We develop a bridge between the activities in two domains by learning a similarity function via Web search, under the condition that the sensor data are from the same feature space. Based on the learned similarity measures, our algorithm interprets the data from the source domain as the data in the domain with different confidence levels, thus accomplishing the cross-domain knowledge transfer task. Our algorithm is evaluated on several real-world datasets to demonstrate its effectiveness. Vincent Wenchen Zheng, Derek Hao Hu, Qiang Yang 0001 |
UbiComp | 1 |
| 2009 | Abnormal Activity Recognition Based on HDP-HMM Models
Derek Hao Hu, Xian-Xing Zhang, Jie Yin 0001, Vincent Wenchen Zheng, Qiang Yang 0001 |
IJCAI | 4 |
| 2008 | Transferring Multi-device Localization Models using Latent Multi-task Learning
Vincent Wenchen Zheng, Sinno Jialin Pan, Qiang Yang 0001, Jeffrey Junfeng Pan |
AAAI | 1 |
| 2008 | Transferring Localization Models over Time
Vincent Wenchen Zheng, Evan Wei Xiang, Qiang Yang 0001, Dou Shen |
AAAI | 1 |
| 2008 | Real world activity recognition with multiple goalsabstractRecognizing and understanding the activities of people from sensor readings is an important task in ubiquitous computing. Activity recognition is also a particularly difficult task because of the inherent uncertainty and complexity of the data collected by the sensors. Many researchers have tackled this problem in an overly simplistic setting by assuming that users often carry out single activities one at a time or multiple activities consecutively, one after another. However, so far there has been no formal exploration on the degree in which humans perform concurrent or interleaving activities, and no thorough study on how to detect multiple goals in a real world scenario. In this article, we ask the fundamental questions of whether users often carry out multiple concurrent and interleaving activities or single activities in their daily life, and if so, whether such complex behavior can be detected accurately using sensors. We define several classes of complexity levels under a goal taxonomy that describe different granularities of activities, and relate the recognition accuracy with different complexity levels or granularities. We present a theoretical framework for recognizing multiple concurrent and interleaving activities, and evaluate the framework in several real-world ubiquitous computing environments. Derek Hao Hu, Sinno Jialin Pan, Vincent Wenchen Zheng, Nathan Nan Liu, Qiang Yang 0001 |
UbiComp | 3 |
| 2008 | Digital Wall: A Power-efficient Solution for Location-based Data SharingabstractWith the proliferation of wireless and sensor techniques, data can be shared conveniently through the air. However, wireless communication is vulnerable since unauthorized machine may try to intrude a server without being physically connected. In this paper, we wish to control the communication between a wireless client and the infrastructure based on the client location. Our idea is to implement a digital wall, which is a user-defined boundary so that access is allowed within the boundary and denied outside the boundary. To do this, we need to do accurate location estimation since the decision around the boundary line is critical. Furthermore, computational efficiency is also important since we need to reduce computation cost so as to save power energy. In this paper, we propose k-nearest-neighbor (KNN) based method to determine the location of a mobile client based on received signal strength (RSS) values. We further use information gain as a feature selection criterion to reduce the estimation time. We study how the performance may be affected when the controlled area is confined by physical or virtual walls. Experimental results show that we can well distinguish whether a client device is located within an expected digital wall. Jeffrey Junfeng Pan, Sinno Jialin Pan, Vincent Wenchen Zheng, Qiang Yang 0001 |
PerCom | 3 |
| 2007 | Domain-constrained semi-supervised mining of tracking models in sensor networksabstractAccurate localization of mobile objects is a major research problem in sensor networks and an important data mining application. Specifically, the localization problem is to determine the location of a client device accurately given the radio signal strength values received at the client device from multiple beacon sensors or access points. Conventional data mining and machine learning methods can be applied to solve this problem. However, all of them require large amounts of labeled training data, which can be quite expensive. In this paper, we propose a probabilistic semi supervised learning approach to reduce the calibration effort and increase the tracking accuracy. Our method is based on semi-supervised conditional random fields which can enhance the learned model from a small set of training data with abundant unlabeled data effectively. To make our method more efficient, we exploit a Generalized EM algorithm coupled with domain constraints. We validate our method through extensive experiments in a real sensor network using Crossbow MICA2 sensors. The results demonstrate the advantages of methods compared to other state-of-the-art object-tracking algorithms. Vincent Wenchen Zheng, Jeffrey Junfeng Pan, Dou Shen, Sinno Jialin Pan, Qiang Yang 0001 |
KDD | 3 |
| 2007 | A Joint Power Control, Link Scheduling and Rate Control Algorithm for Wireless Ad Hoc NetworksabstractIn this paper, we present a joint power control, link scheduling and rate control (PSR) algorithm for wireless ad hoc networks by using the convex optimization theory. This algorithm practically considers the power control problem in the interference-based link scheduling process, and provides a congestion control on the transport layer. Both our theoretical analysis and simulation results prove that the PSR algorithm can converge quickly and is possible for distributed implementation in the wireless ad hoc networks. Vincent Wenchen Zheng, Xinming Zhang 0001, Daoke Liu, Dan Keun Sung |
WCNC | 1 |