Jie Zhou 0001

dblp:00/5012-1 · DBLP profile ↗
← Back
17ranked-venue papers in the field
0as first author
3since 2021 · last 2024
0000-0001-7701-234XORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 6Data Mining & Knowledge Discovery · 5Other / Interdisciplinary · 3Database Systems & Data Management · 2Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2024 A Secure and Fair Client Selection Based on DDPG for Federated Learning
abstract
Federated learning (FL) is a machine learning technique in which a large number of clients collaborate to train models without sharing private data. However, FL’s integrity is vulnerable to unreliable models; for instance, data poisoning attacks can compromise the system. In addition, system preferences and resource disparities preclude fair participation by reliable clients. To address this challenge, we propose a novel client selection strategy that introduces a security‐fairness value to measure client performance in FL. The value in question is a composite metric that combines a security score and a fairness score. The former is dynamically calculated from a beta distribution reflecting past performance, while the latter considers the client’s participation frequency in the aggregation process. The weighting strategy based on the deep deterministic policy gradient (DDPG) determines these scores. Experimental results confirm that our method fairly effectively selects reliable clients and maintains the security and fairness of the FL system.
Tao Wan 0003, Shun Feng, Weichuan Liao, Nan Jiang 0013, Jie Zhou 0001
Int. J. Intell. Syst.5
2023 Regional digital infrastructure, enterprise digital transformation and entrepreneurial orientation: Empirical evidence based on the broadband china strategy
Wenqing Wu 0003, Jie Zhou 0001
Inf. Process. Manag.4
2022 Incorporating multi-interest into recommendation with graph convolution networks
abstract
In recent years, the appearance of graph convolutional networks (GCNs) provides a new idea for graph structure data processing. Because of that, they can learn excellent user and item embedding by using cooperative signals of high-order neighbors, and the GCNs technique shows great potential in the recommendation. The common problem with the bulk of GCN-based models is that it appears the situation of performance degradation during the stacking of network layers. The recently proposed IMP-GCN alleviates this problem to some extent. It aims to avoid the influence of downside information from high-order propagation on embedding learning. However, we consider that it ignores the multi-interest factor, in which users may have different interests. In this paper, we present a multi-interest GCN(MI-GCN) model for a recommendation, and it conducts high-order graph convolution operations in three sets of subgraphs. Users with similar interests and the corresponding interaction items belong to the identical subgraph. As for the formation of the subgraph, we adopt two varied clustering methods and the user feature to form a subgraph generation mechanism. This mechanism can generate three groups of differential subgraphs to divide users into multi-interest groups and make subgraph division more reasonable. We carry out massive experiments on three real-world datasets, demonstrating the effectiveness of our model. Experimental results confirm that our presented MI-GCN outperforms the state-of-the-art GCN-based recommendation models.
Nan Jiang 0013, Zilin Zeng, Jie Zhou 0001, Tao Wan 0003, Ximeng Liu, Honglong Chen
Int. J. Intell. Syst.4
2020 Discovering attractive segments in the user-generated video streams
Zheng Wang 0044, Jie Zhou 0001, Jing Ma 0004, Jingjing Li 0001, Jiangbo Ai, Yang Yang 0002
Inf. Process. Manag.2
2020 Toward optimal participant decisions with voting-based incentive model for crowd sensing
Nan Jiang 0013, Dong Xu 0020, Jie Zhou 0001, Hongyang Yan, Tao Wan 0003, Jiaqi Zheng 0001
Inf. Sci.3
2015 Optimal answerer ranking for new questions in community question answering
Zhenlei Yan, Jie Zhou 0001
Inf. Process. Manag.2
2013 Dynamic Personalized Recommendation on Sparse Data
abstract
Recommendation techniques are very important in the fields of E-commerce and other web-based services. One of the main difficulties is dynamically providing high-quality recommendation on sparse data. In this paper, a novel dynamic personalized recommendation algorithm is proposed, in which information contained in both ratings and profile contents are utilized by exploring latent relations between ratings, a set of dynamic features are designed to describe user preferences in multiple phases, and finally, a recommendation is made by adaptively weighting the features. Experimental results on public data sets show that the proposed algorithm has satisfying performance.
Xiangyu Tang, Jie Zhou 0001
IEEE Trans. Knowl. Data Eng.2
2012 A New Approach to Answerer Recommendation in Community Question Answering Services
Zhenlei Yan, Jie Zhou 0001
ECIR2
2010 Collaborative Filtering: Weighted Nonnegative Matrix Factorization Incorporating User and Item Graphs
abstract
Collaborative filtering is an important topic in data mining and has been widely used in recommendation system.In this paper, we proposed a unified model for collaborative filtering based on graph regularized weighted nonnegative matrix factorization.In our model, two graphs are constructed on users and items, which exploit the internal information (e.g.neighborhood information in the user-item rating matrix) and external information (e.g.content information such as user's occupation and item's genre, or other kind of knowledge such as social trust network).The proposed method not only inherits the advantages of model-based method, but also owns the merits of memory-based method which considers the neighborhood information.Moreover, it has the ability to make use of content information and any additional information regarding user-user such as social trust network.Due to the use of these internal and external information, the proposed method is able to find more interpretable lowdimensional representations for users and items, which is helpful for improving the recommendation accuracy.Experimental results on benchmark collaborative filtering data sets demonstrate that the proposed methods outperform the state of the art collaborative filtering methods a lot.
Quanquan Gu, Jie Zhou 0001, Chris Ding
SDM2
2010 Closing the Loop in Webpage Understanding
abstract
The two most important tasks in information extraction from the Web are webpage structure understanding and natural language sentences processing. However, little work has been done toward an integrated statistical model for understanding webpage structures and processing natural language sentences within the HTML elements. Our recent work on webpage understanding introduces a joint model of Hierarchical Conditional Random Fields (HCRFs) and extended Semi-Markov Conditional Random Fields (Semi-CRFs) to leverage the page structure understanding results in free text segmentation and labeling. In this top-down integration model, the decision of the HCRF model could guide the decision making of the Semi-CRF model. However, the drawback of the top-down integration strategy is also apparent, i.e., the decision of the Semi-CRF model could not be used by the HCRF model to guide its decision making. This paper proposed a novel framework called WebNLP, which enables bidirectional integration of page structure understanding and text understanding in an iterative manner. We have applied the proposed framework to local business entity extraction and Chinese person and organization name extraction. Experiments show that the WebNLP framework achieved significantly better performance than existing methods.
Chunyu Yang 0005, Zaiqing Nie, Jie Zhou 0001, Ji-Rong Wen
IEEE Trans. Knowl. Data Eng.4
2009 Subspace maximum margin clustering
abstract
In text mining, we are often confronted with very high dimensional data. Clustering with high dimensional data is a challenging problem due to the curse of dimensionality. In this paper, to address this problem, we propose an subspace maximum margin clustering (SMMC) method, which performs dimensionality reduction and maximum margin clustering simultaneously within a unified framework. We aim to learn a subspace, in which we try to find a cluster assignment of the data points, together with a hyperplane classifier, such that the resultant margin is maximized among all possible cluster assignments and all possible subspaces. The original problem is transformed from learning the subspace to learning a positive semi-definite matrix, in order to avoid tuning the dimensionality of the subspace. The transformed problem can be solved efficiently via cutting plane technique and constrained concave-convex procedure (CCCP). Since the sub-problem in each iteration of CCCP is joint convex, alternating minimization is adopted to obtain the global optimum. Experiments on benchmark data sets illustrate that the proposed method outperforms the state of the art clustering methods as well as many dimensionality reduction based clustering approaches.
Quanquan Gu, Jie Zhou 0001
CIKM2
2009 Learning the Shared Subspace for Multi-task Clustering and Transductive Transfer Classification
abstract
There are many clustering tasks which are closely related in the real world, e.g. clustering the Web pages of different universities. However, existing clustering approaches neglect the underlying relation and treat these clustering tasks either individually or simply together. In this paper, we will study a novel clustering paradigm, namely multi-task clustering, which performs multiple related clustering tasks together and utilizes the relation of these tasks to enhance the clustering performance. We aim to learn a subspace shared by all the tasks, through which the knowledge of the tasks can be transferred to each other. The objective of our approach consists of two parts: (1) Within-task clustering: clustering the data of each task in its input space individually; and (2) Cross-task clustering: simultaneous learning the shared subspace and clustering the data of all the tasks together. We will show that it can be solved by alternating minimization, and its convergence is theoretically guaranteed. Furthermore, we will show that given the labels of one task, our multi-task clustering method can be extended to transductive transfer classification (a.k.a. cross-domain classification, domain adaption). Experiments on several cross-domain text data sets demonstrate that the proposed multi-task clustering outperforms traditional single-task clustering methods greatly. And the transductive transfer classification method is comparable to or even better than several existing transductive transfer classification approaches.
Quanquan Gu, Jie Zhou 0001
ICDM2
2009 Co-clustering on manifolds
abstract
Co-clustering is based on the duality between data points (e.g. documents) and features (e.g. words), i.e. data points can be grouped based on their distribution on features, while features can be grouped based on their distribution on the data points. In the past decade, several co-clustering algorithms have been proposed and shown to be superior to traditional one-side clustering. However, existing co-clustering algorithms fail to consider the geometric structure in the data, which is essential for clustering data on manifold. To address this problem, in this paper, we propose a Dual Regularized Co-Clustering (DRCC) method based on semi-nonnegative matrix tri-factorization. We deem that not only the data points, but also the features are sampled from some manifolds, namely data manifold and feature manifold respectively. As a result, we construct two graphs, i.e. data graph and feature graph, to explore the geometric structure of data manifold and feature manifold. Then our co-clustering method is formulated as semi-nonnegative matrix tri-factorization with two graph regularizers, requiring that the cluster labels of data points are smooth with respect to the data manifold, while the cluster labels of features are smooth with respect to the feature manifold. We will show that DRCC can be solved via alternating minimization, and its convergence is theoretically guaranteed. Experiments of clustering on many benchmark data sets demonstrate that the proposed method outperforms many state of the art clustering methods.
Quanquan Gu, Jie Zhou 0001
KDD2
2009 Transductive Classification via Dual Regularization
Quanquan Gu, Jie Zhou 0001
ECML/PKDD (1)2
2009 Local Relevance Weighted Maximum Margin Criterion for Text Classification
abstract
Text classification is a very important task in information retrieval and data mining. In vector space model (VSM), document is represented as a high dimensional vector, and a feature extraction phase is usually needed to reduce the dimensionality of the document. In this paper, we propose a feature extraction method, named Local Relevance Weighted Maximum Margin Criterion (LRWMMC). It aims to learn a subspace in which the documents in the same class are as near as possible while the documents in the different classes are as far as possible in the local region of each document. Furthermore, the relevance is taken into account as a weight to determine the extent to which the documents will be projected. LRWMMC is able to find the low dimensional manifold embedded in the high dimensional ambient space. In addition, We generalize LRWMMC to Reproducing Kernel Hilbert Space (RKHS), which can resolve the nonlinearity of the input space. We also generalize LRWMMC to tensor space which is suitable for a new document representation, named tensor space model (TSM). On the other hand, in order to utilize the large amount of unlabeled documents, we also present a Semi-Supervised LRWMMC, which aims to find a projection inferred from the labeled samples, as well as the unlabeled samples. Finally, we present a fast algorithm based on QR-decomposition to make the methods proposed in this paper apply for large scale data set. Encouraging experimental results on benchmark text classification data sets indicate that the proposed methods outperform many existing feature extraction methods for text classification.
Quanquan Gu, Jie Zhou 0001
SDM2
2009 Stock Price Forecasting by Combining News Mining and Time Series Analysis
abstract
Stock price forecasting has aroused great concern in research of economy, machine learning and other fields. Time series analysis methods are usually utilized to deal with this task. In this paper, we propose to combine news mining and time series analysis to forecast inter-day stock prices. News reports are automatically analyzed with text mining techniques, and then the mining results are used to improve the accuracy of time series analysis algorithms. The experimental result on a half year Chinese stock market data indicates that the proposed algorithm can help to improve the performance of normal time series analysis in stock price forecasting significantly. Moreover, the proposed algorithm also performs well in stock price trend forecasting.
Xiangyu Tang, Chunyu Yang 0005, Jie Zhou 0001
Web Intelligence3
2008 Closing the loop in webpage understanding
abstract
Little work has been done towards an integrated statistical model for understanding webpage structures and processing natural language sentences within the HTML elements. This paper proposed a novel framework called WebNLP which enables bidirectional integration of page structure understanding and text understanding in an iterative manner. Experiments show that the WebNLP framework achieved significantly better performance.
Chunyu Yang 0005, Zaiqing Nie, Jie Zhou 0001, Ji-Rong Wen
CIKM4