Zhiang Wu 0001

dblp:00/4248 · DBLP profile ↗
← Back
72ranked-venue papers
15as first author
10since 2021 · last 2026
0000-0002-0591-1861ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 27 · 6 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 1 since 2021Systems, architecture and hardware · 6 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4Software engineering, systems software and programming languages · 3Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing text classification with neural label embedding and weakly-supervised learning
Xiao Jing, Zhe Li 0039, Zhiang Wu 0001
Expert Syst. Appl.3
2026 Data-centric cross-city knowledge transfer for regional socioeconomic prediction
Hongru Lu, Dong Kan, Xiao Jing, Zhiang Wu 0001
Knowl. Based Syst.4
2026 Few-shot urban strategic group discovery with LLM annotation and hypergraph learning
Hongru Lu, Xiao Jing, Zhiang Wu 0001
World Wide Web (WWW)6
2025 Revisiting intelligent audit from a data science perspective
Hongru Lu, Zhiang Wu 0001
Neurocomputing2
2025 Efficient and Flexible Long-Tail Recommendation Using Cosine Patterns
abstract
With the increasing use of recommender systems in various application domains, many algorithms have been proposed for improving the accuracy of recommendations. Among various dimensions of recommender systems performance, long-tail (niche) recommendation performance remains an important challenge in large part because of the popularity bias of many existing recommendation techniques. In this study, we propose CORE, a cosine pattern–based technique, for effective long-tail recommendation. Comprehensive experimental results compare the proposed approach with a wide variety of classic, widely used recommendation algorithms and demonstrate its practical benefits in accuracy, flexibility, and scalability in addition to the superior long-tail recommendation performance. 1 History: Accepted by Ramaswamy Ramesh, Area Editor for Data Science & Machine Learning. Funding: This work was supported by the National Natural Science Foundation of China [Grants 72031001, 72072091, 72242101]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.0194 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2022.0194 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .
Yaqiong Wang, Junjie Wu 0002, Zhiang Wu 0001, Gediminas Adomavicius
INFORMS J. Comput.3
2023 Electrical Fault Diagnosis via Text Mining: A Weakly-Supervised Learning Model
abstract
To automatically and accurately identify electrical fault categories by means of data-driven approaches has attracted much attention in recent years. However, the research on fault diagnosis by mining unstructured data such as text data is far from being perfect. One problem lies in that a large number of human-labeled samples is often costly and difficult to obtain in real applications. In this study, we present a novel Label Embedding joint with Weakly-supervised Classification model (LEMWEC) for predicting fault category given the fault-descriptive corpus. Our model leverages a few labeled documents and a good number of pseudo-labeled documents to learn both sentence embeddings and the multi-label classifier. Then, the model will be iteratively refined by the newly generated pseudo documents. Experiments on a real-life dataset demonstrate the LEMWEC can improve the accuracy of electrical fault classification remarkably compared with a comprehensive set of baselines.
Xiao Jing, Zhiang Wu 0001
CSCWD2
2023 PipePar: Enabling fast DNN pipeline parallel training in heterogeneous GPU clusters
Jinghui Zhang 0001, Geng Niu, Qiangsheng Dai, Fang Dong 0001, Zhiang Wu 0001
Neurocomputing7
2022 CrimeTensor: Fine-Scale Crime Prediction via Tensor Learning with Spatiotemporal Consistency
abstract
Crime poses a major threat to human life and property, which has been recognized as one of the most crucial problems in our society. Predicting the number of crime incidents in each region of a city before they happen is of great importance to fight against crime. There has been a great deal of research focused on crime prediction, ranging from introducing diversified data sources to exploring various prediction models. However, most of the existing approaches fail to offer fine-scale prediction results and take little notice of the intricate spatial-temporal-categorical correlations contained in crime incidents. In this article, we propose a tailor-made framework called CrimeTensor to predict the number of crime incidents belonging to different categories within each target region via tensor learning with spatiotemporal consistency. In particular, we model the crime data as a tensor and present an objective function which tries to take full advantage of the spatial, temporal, and categorical correlations contained in crime incidents. Moreover, a well-designed optimization algorithm which transforms the objective into a compact form and then applies CP decomposition to find the optimal solution is elaborated to solve the objective function. Furthermore, we develop an enhanced framework which takes a set of pre-selected regions to conduct prediction so as to further improve the computational efficiency of the optimization algorithm. Finally, extensive experiments are performed on both proprietary and public datasets and our framework significantly outperforms all the baselines in terms of each evaluation metric.
Weichao Liang, Zhiang Wu 0001, Zhe Li 0039, Yong Ge 0001
ACM Trans. Intell. Syst. Technol.2
2022 Community Detection via Local Learning Based on Generalized Metric With Neighboring Regularization
abstract
Community detection has long been a fundamental problem in network analysis. A great deal of previous research has regarded community detection as an optimization process, where a variety of internal quality metrics are typically treated as objective functions, such as modularity (${Q}$) and weighted community clustering (WCC). However, purely optimizing a predefined quality metric probably results in an extreme unbalance in the scale of the detected communities, e.g., few giant communities with many very small communities. To reveal the true mesoscopic structure inside a big network under the suitable number of communities, we propose a novel community detection framework called LL-GMR, which is a local learning framework based on generalized metric with neighboring regularization. LL-GMR is qualified for both nonoverlapping and overlapping detection tasks. In LL-GMR, we propose a generalized representation and illustrate that it can be instantiated into 12 well-known internal quality metrics. When the generalized metric is used as an objective function, we encode node-level and community-level neighborhood information into two regularization terms to alleviate the dilemma of unbalanced communities. The experimental results show that our LL-GMR consistently outperforms other state-of-the-art community detection approaches in terms of discovering ground-truth communities in six real-life networks.
Guangliang Gao, Zhiang Wu 0001, Lu Zhang 0030, Jie Cao 0001, Xinzhe Qi
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Leveraging multi-level dependency of relational sequences for social spammer detection
Jun Yin 0005, Qian Li 0003, Shaowu Liu, Zhiang Wu 0001, Guandong Xu
Neurocomputing4
2020 Personalized itinerary recommendation: Deep and collaborative learning with textual information
Lei Chen 0079, Lu Zhang 0030, Shanshan Cao, Zhiang Wu 0001, Jie Cao 0001
Expert Syst. Appl.4
2020 Intelligent online catastrophe assessment and preventive control via a stacked denoising autoencoder
Mingyu Zhai, Jiahui Jin 0001, Aibo Song, Jikeng Lin, Zhiang Wu 0001, Yixin Zhao
Neurocomputing6
2020 Efficiently Translating Complex SQL Query to MapReduce Jobflow on Cloud
abstract
MapReduce is a widely-used programming model in cloud environment for parallel processing large-scale data sets. The combination of the high-level language with a SQL-to-MapReduce translator allows programmers to code using SQL-like declarative language, so that each program can afterwards be complied into a MapReduce jobflow automatically. This way is helpful to narrow the gap between non-professional users and cloud platforms, and thus significantly improve the usability of the cloud. Although a number of translators have been developed, the auto-generated MapReduce programs still suffered from extremely inefficiency. In this paper, we present an efficient Cost-Aware SQL-to-MapReduce Translator (CAT). CAT has two notable features. First, it defines two intra-SQL correlations: Generalized Job Flow Correlation (GJFC) and Input Correlation (IC), based on which a set of looser merging rules are introduced. Thus, both Top-Down (TD) and Bottom-Up (BU) merging strategies are proposed and integrated into CAT simultaneously. Second, it adopts a cost estimation model for MapReduce jobflows to guide the selection of a more efficient MapReduce jobflows auto-generated by TD and BU merging strategies. Finally, comparative experiments on TPC-H benchmark demonstrate the effectiveness and scalability of CAT.
Zhiang Wu 0001, Aibo Song, Jie Cao 0001, Junzhou Luo, Lu Zhang 0030
IEEE Trans. Cloud Comput.1
2020 hPSD: A Hybrid PU-Learning-Based Spammer Detection Model for Product Reviews
abstract
Spammers, who manipulate online reviews to promote or suppress products, are flooding in online commerce. To combat this trend, there has been a great deal of research focused on detecting review spammers, most of which design diversified features and thus develop various classifiers. The widespread growth of crowdsourcing platforms has created large-scale deceptive review writers who behave more like normal users, that the way they can more easily evade detection by the classifiers that are purely based on fixed characteristics. In this paper, we propose a hybrid semisupervised learning model titled hybrid PU-learning-based spammer detection (hPSD) for spammer detection to leverage both the users' characteristics and the user-product relations. Specifically, the hPSD model can iteratively detect multitype spammers by injecting different positive samples, and allows the construction of classifiers in a semisupervised hybrid learning framework. Comprehensive experiments on movie dataset with shilling injection confirm the superior performance of hPSD over existing baseline methods. The hPSD is then utilized to detect the hidden spammers from real-life Amazon data. A set of spammers and their underlying employers (e.g., book publishers) are successfully discovered and validated. These demonstrate that hPSD meets the real-world application scenarios and can thus effectively detect the potentially deceptive review writers.
Zhiang Wu 0001, Jie Cao 0001, Yaqiong Wang, Youquan Wang, Lu Zhang 0030, Junjie Wu 0002
IEEE Trans. Cybern.1
2020 Travel Recommendation via Fusing Multi-Auxiliary Information into Matrix Factorization
abstract
As an e-commerce feature, the personalized recommendation is invariably highly-valued by both consumers and merchants. The e-tourism has become one of the hottest industries with the adoption of recommendation systems. Several lines of evidence have confirmed the travel-product recommendation is quite different from traditional recommendations. Travel products are usually browsed and purchased relatively infrequently compared with other traditional products (e.g., books and food), which gives rise to the extreme sparsity of travel data. Meanwhile, the choice of a suitable travel product is affected by an army of factors such as departure, destination, and financial and time budgets. To address these challenging problems, in this article, we propose a Probabilistic Matrix Factorization with Multi-Auxiliary Information (PMF-MAI) model in the context of the travel-product recommendation. In particular, PMF-MAI is able to fuse the probabilistic matrix factorization on the user-item interaction matrix with the linear regression on a suite of features constructed by the multiple auxiliary information. In order to fit the sparse data, PMF-MAI is built by a whole-data based learning approach that utilizes unobserved data to increase the coupling between probabilistic matrix factorization and linear regression. Extensive experiments are conducted on a real-world dataset provided by a large tourism e-commerce company. PMF-MAI shows an overwhelming superiority over all competitive baselines on the recommendation performance. Also, the importance of features is examined to reveal the crucial auxiliary information having a great impact on the adoption of travel products.
Lei Chen 0079, Zhiang Wu 0001, Jie Cao 0001, Guixiang Zhu, Yong Ge 0001
ACM Trans. Intell. Syst. Technol.2
2020 On Scalability of Association-rule-based Recommendation: A Unified Distributed-computing Framework
abstract
The association-rule-based approach is one of the most common technologies for building recommender systems and it has been extensively adopted for commercial use. A variety of techniques, mainly including eligible rule selection and multiple rules combination, have been developed to create effective recommendation. Unfortunately, little attention has been paid to the scalability concern of rule-based recommendation methods. However, the computational complexity of rule-based methods shall increase drastically with the growth of both online customers and rules, which are usually several millions in typical e-commerce platforms. Moreover, the dynamic change of users’ actions requires rule-based methods make recommendations in nearly real-time, which further highlights the scalability issue of rule-based recommender systems. In this article, we present a distributed framework that can scale different association-rule-based recommendation methods in a unified way. Specifically, based on the summarization of existing rule-based approaches, a generic tree-type structure is defined to store separate kinds of patterns, and an efficient algorithm is designed for mining eligible patterns along with computing recommendation scores. To handle the ever-increasing number of online customers, a distributed framework is proposed, where two load-balanced strategies for partitioning tree are put forward to fit sparse and dense data, respectively. Extensive experiments on five real-life data sets demonstrate that the efficiency of association-rule-based recommender systems can be significantly improved by the proposed framework.
Zhiang Wu 0001, Jie Cao 0001, Yong Ge 0001
ACM Trans. Web1
2020 A Pseudo-document-based Topical N-grams model for short texts
Hao Lin 0002, Yuan Zuo, Guannan Liu 0004, Junjie Wu 0002, Zhiang Wu 0001
World Wide Web6
2019 MBECN: Enabling ECN with Micro-burst Traffic in Multi-queue Data Center
abstract
Modern multi-queue data centers often use the standard Explicit Congestion Notification (ECN) scheme to achieve high network performance. However, one substantial drawback of this approach is that micro-burst traffic can cause the instantaneous queue length to exceed the ECN's threshold, resulting in numerous mismarkings. After enduring too many mismarkings, senders may overreact, leading to severe throughput loss. As a solution to this dilemma, we propose our own adaptationthe Micro-burst ECN (MBECN) scheme-to mitigate mismarking. MBECN finds a more appropriate threshold baseline for each queue to absorb micro-bursts, based on steady-state analysis and an ideal generalized processor sharing (GPS) model. By adopting a queue-occupation-based dynamically adjusting algorithm, MBECN effectively handles packet backlog without hurting latency. Through testbed experiments, we find that MBECN improves throughput by ~20% and reduces flow completion time (FCT) by ~40%. Using large scale simulations, we find that throughput can be improved by 1.5~2.4× with DCTCP and 1.26~1.35× with ECN*. We also measure network delay and find that latency only increases by 7.36%.
Kexi Kang, Jinghui Zhang 0001, Jiahui Jin 0001, Dian Shen, Junzhou Luo, Zhiang Wu 0001
CLUSTER7
2019 Verifying the claimed sale-ranking trustworthy: A maximum marginal relevance-based ranking method
abstract
Summary Various online contents on Internet platforms or search engines are related to the corporate reputation. Facing the huge amount of online contents, we need a mining method that can automatically extract and analyze a large number of network‐related information and obtain the real reliability of aspect for the content claimed by companies. In this paper, we propose to generate a ranking model to verify whether the sales‐rankings claimed by companies are trustworthy. The key idea is that the company that has higher confidence score should be supported by the online media. We use a unique data set of public opinion data related with a specific company, which we supplement with data from various online news platform and retrieval webpages using a distributed and generic Web crawler. Meanwhile, basic information and open financial data of companies are also collected for auxiliary analysis. We present a Maximal Marginal Relevance‐based ranking model to compute the confidence score of each company, taking into consideration the two technologies of word embedding and KL‐Divergence to filter the irrelevant documents. Extensive experiments show that the proposed method outperforms the state‐of‐the‐art MMR‐based method, and we showcase three representative cases about the corporate reputation built by us that gives positive, neutral, and negative support respectively to the sales‐ranking claim of companies.
Youquan Wang, Changjian Fang, Dongqin Shen, Zhiang Wu 0001, Jie Cao 0001
Concurr. Comput. Pract. Exp.4
2019 A local random walk model for complex networks based on discriminative feature combinations
Aibo Song, Zhiang Wu 0001, Mingyu Zhai, Junzhou Luo
Expert Syst. Appl.3
2019 Link communities detection: an embedding method on the line hypergraph
Haicheng Tao, Zhe Li 0039, Zhiang Wu 0001, Jie Cao 0001
Neurocomputing3
2019 Link prediction in temporal networks: Integrating survival analysis and game theory
Zhan Bu, Hui-Jia Li, Jiuchuan Jiang, Zhiang Wu 0001, Jie Cao 0001
Inf. Sci.5
2018 Understanding Customer Behavior in Shopping Mall from Indoor Tracking Data
abstract
The prosperity of various indoor positioning technologies makes possible the large collection of tracking data in indoor spaces. Much of the focus has been on several fundamental problems such as indoor localization, indoor space modeling, indoor data cleansing, indexing and querying. However, this paper attempts to analyze customer behavior from a unique indoor tracking data, which will promote the convergence between various applications and the underlying data. In particular, we introduce a real-life indoor tracking data collected at an electrical mall in China. Then, we cluster users into several groups and summarize the most characteristic behaviors of each cluster. Last but not least, we analyze customer's individual behaviors through two aspects: 1) the regression model is used to reveal hot regions where customers are likely to stay in long time; and 2) a transition matrix combined with the connectivity is presented to demonstrate hot paths.
Weichao Liang, Zhiang Wu 0001, Jie Cao 0001
CSCWD2
2018 An Efficient Distributed-Computing Framework for Association-Rule-Based Recommendation
abstract
The association-rule-based recommendation model is one of the most widely used commercial recommendation engines in e-commerce websites. Existing studies mostly focus on how to select eligible rules to enhance the recommendation performance, but the efficiency of recommendation has been paid few attentions. To remedy this, this paper develops a distributed-computing framework for improving the computational efficiency of rule-based recommendation. Specifically, a tree-typed structure called Ordered-Patterns Forest (OPF) is designed to compress and store frequent patterns. Then, we transform eligible rules mining to a path-searching problem on OPF, and present a path-searching algorithm running on single machine. Finally, a load-balanced strategy for data partitioning is clarified. Experimental results demonstrate that the efficiency improved remarkably by the proposed OPF, compared with the traditional Brute-Force method.
Weichao Liang, Zhiang Wu 0001, Jie Cao 0001
ICWS3
2018 A Sequential Recommendation for Mobile Apps: What Will User Click Next App?
abstract
With the advances in smartphones users install abundant apps to facilitate their daily lives. Both users and related developers have increasing requirements to understand the mobile App usage pattern, for individual and commercial use. Respectively, personalized App recommendation methods and systems have emerged as a novel attractive topic that can demonstrate the human App usage behavior. The mobile Apps recommendation can serve as a cornerstone for a variety of intelligent services, such as fast-launching UIs, intelligent user-phone interactions, and battery management of cellphones. In this paper, we develop a novel App recommendation framework combining the historical App usage data with the sequence of recently-used Apps. Specifically, our framework is an extension of the user-based collaborative filtering technique, where the set of nearest neighbors is employed for training the prediction model. However, our prediction scheme is constructed on the temporal sequential data and is modeled by using the chain-augmented Naive Bayes model. Experiments with a real mobile Internet record dataset demonstrate that the accuracy of our framework outperforms several baseline App recommendation approaches.
Chaoyi Pu, Zhiang Wu 0001, Jie Cao 0001
ICWS2
2018 Social Spammer Detection: A Multi-Relational Embedding Approach
Jun Yin 0005, Shaowu Liu, Zhiang Wu 0001, Guandong Xu
PAKDD (1)4
2018 A generalized game theoretic framework for mining communities in complex networks
Guangliang Gao, Jie Cao 0001, Zhan Bu, Hui-Jia Li, Zhiang Wu 0001
Expert Syst. Appl.5
2017 Word-sentence co-ranking for automatic extractive text summarization
Changjian Fang, Zhenghong Deng 0001, Zhiang Wu 0001
Expert Syst. Appl.4
2017 A recommendation engine for travel products based on topic sequential patterns
Guixiang Zhu, Jie Cao 0001, Zhiang Wu 0001
Multim. Tools Appl.4
2017 Fuzzy Consensus Clustering With Applications on Big Data
abstract
Consensus clustering aims to find a single partition of data that agrees as much as possible with existing basic partitions. Given its robustness and generalizability, consensus clustering has emerged as a promising solution to find cluster structures inside heterogeneous big data rising from various application domains. In the area of fuzzy systems, however, research along this line is still in its initial stage with some unsystematic algorithmic studies. Finding a fuzzy consensus partition from multiple fuzzy basic partitions in an efficient, flexible, and robust way is still an exciting open problem calling for further investigation. In light of this, this paper provides a systematic study of fuzzy consensus clustering (FCC) from a utility perspective. Specifically, we first define the objective function of FCC clearly using the novel fuzzified contingency matrix. We then derive a family of FCC Utility functions termed as FCCU that can transform FCC to a weighted piecewise fuzzy $c$ -means clustering (piFCM) problem. This helps us to establish an algorithmic framework for FCC with flexible choice of utility functions, and speeds FCC significantly with a FCM-like iterative process of piFCM. To meet the big data challenge, we further parallelize FCC on the Spark platform with both vertical and horizontal segmentation schemes. Extensive experiments on various real-world datasets demonstrate the excellent performance of FCC, even with a majority of poor basic partitions. In particular, our method exhibits interesting potential for big data clustering in two real-life applications concerned with online event detection and overlapping community detection, respectively.
Junjie Wu 0002, Zhiang Wu 0001, Jie Cao 0001, Hongfu Liu 0001, Yanchun Zhang
IEEE Trans. Fuzzy Syst.2
2016 A Relaxed Ranking-Based Factor Model for Recommender System from Implicit Feedback
Richang Hong, Defu Lian, Zhiang Wu 0001, Meng Wang 0001, Yong Ge 0001
IJCAI4
2016 A Spatial-Temporal Probabilistic Matrix Factorization Model for Point-of-Interest Recommendation
abstract
With the rapid development of Location-based Social Network (LBSN) services, a large number of Point-of-Interests (POIs) have been available, which consequently raises a great demand of building personalized POI recommender systems. A personalized POI recommender system can significantly help users to find their preferred POIs and assist POI owners to attract more customers. However, due to the complexity of users' checkin decision making process that is influenced by many different factors such as POI distance and region's prosperity, and the dynamics of user's preference, POI recommender systems usually suffer from many challenges. Although different latent factor based methods (e.g., probabilistic matrix factorization) have been proposed, most of them do not successfully incorporate both geographical influence and temporal effect together into latent factor models. To this end, in this paper, we propose a new Spatial-Temporal Probabilistic Matrix Factorization (STPMF) model that models a user's preference for POI as the combination of his geographical preference and other general interest in POI. Furthermore, in addition to static general interest of user, we capture the temporal dynamics of user's interest as well by modeling checkin data in a unique way. To evaluate the proposed STPMF model, we conduct extensive experiments with many state-of-the-art baseline methods and evaluation metrics on two real-world data sets. The experimental results clearly demonstrate the effectiveness of our proposed STPMF model.
Richang Hong, Zhiang Wu 0001, Yong Ge 0001
SDM3
2016 Predicting Replacement of Smartphones with Mobile App Usage
Dun Yang, Zhiang Wu 0001, Jie Cao 0001, Guandong Xu
WISE (1)2
2016 CAT: A Cost-Aware Translator for SQL-query workflow to MapReduce jobflow
Aibo Song, Zhiang Wu 0001, Junzhou Luo
Data Knowl. Eng.2
2016 Game theory based emotional evolution analysis for chinese online reviews
Zhan Bu, Hui-Jia Li, Jie Cao 0001, Zhiang Wu 0001, Lu Zhang 0030
Knowl. Based Syst.4
2016 Local Community Mining on Distributed and Dynamic Networks From a Multiagent Perspective
abstract
Distributed and dynamic networks are ubiquitous in many real-world applications. Due to the huge-scale, decentralized, and dynamic characteristics, the global topological view is either too hard to obtain or even not available. So, most existing community detection methods working on the global view fail to handle such decentralized and dynamic large networks. In this paper, we propose a novel autonomy-oriented computing-based method for community mining (AOCCM) from the multiagent perspective in the distributed environment. In particular, AOCCM utilizes reactive agents to pick the neighborhood node with the largest structural similarity as the candidate node, and thus determine whether it should be added into local community based on the modularity gain. We further improve AOCCM to a more efficient incremental version named AOCCM-i for mining communities from dynamic networks. AOCCM and AOCCM-i can be easily expanded to detect both nonoverlapping and overlapping global community structures. Experimental results on real-life networks demonstrate that the proposed methods can reduce the computational cost by avoiding repeated structural similarity calculation and can still obtain the high-quality communities.
Zhan Bu, Zhiang Wu 0001, Jie Cao 0001, Yichuan Jiang
IEEE Trans. Cybern.2
2016 Efficient batch similarity join processing of social images based on arbitrary features
Yi Zhuang 0001, Nan Jiang 0002, Zhiang Wu 0001, Jie Cao 0001, Chunhua Ju
World Wide Web3
2015 Spammers Detection from Product Reviews: A Hybrid Model
abstract
Driven by profits, spam reviews for product promotion or suppression become increasingly rampant in online shopping platforms. This paper focuses on detecting hidden spam users based on product reviews. In the literature, there have been tremendous studies suggesting diversified methods for spammer detection, but whether these methods can be combined effectively for higher performance remains unclear. Along this line, a hybrid PU-learning-based Spammer Detection (hPSD) model is proposed in this paper. On one hand, hPSD can detect multi-type spammers by injecting or recognizing only a small portion of positive samples, which meets particularly real-world application scenarios. More importantly, hPSD can leverage both user features and user relations to build a spammer classifier via a semi-supervised hybrid learning framework. Experimental results on movie data sets with shilling injection show that hPSD outperforms several state-of-the-art baseline methods. In particular, hPSD shows great potential in detecting hidden spammers as well as their underlying employers from a real-life Amazon data set. These demonstrate the effectiveness and practical value of hPSD for real-life applications.
Zhiang Wu 0001, Youquan Wang, Yaqiong Wang, Junjie Wu 0002, Jie Cao 0001, Lu Zhang 0030
ICDM1
2015 Regularizing Flat Latent Variables with Hierarchical Structures
Rongcheng Lin, Xiaojun Quan, Richang Hong, Zhiang Wu 0001, Yong Ge 0001
IJCAI5
2015 Coupling Multiple Views of Relations for Recommendation
Guandong Xu, Longbing Cao, Zhiang Wu 0001
PAKDD (2)5
2015 Detecting overlapping communities in poly-relational networks
Zhiang Wu 0001, Jie Cao 0001, Guixiang Zhu, Wenpeng Yin 0001, Alfredo Cuzzocrea
World Wide Web1
2014 Local Community Extraction for Non-overlapping and Overlapping Community Detection
Zhan Bu, Guangliang Gao, Zhiang Wu 0001, Jie Cao 0001
ADMA3
2014 A backbone extraction method with Local Search for complex weighted networks
abstract
The backbone is the natural abstraction of a complex network, which can help people to understand it in a more simplified form. Backbone extraction becomes more challenging as many networks are evolving into large scale and the weight distributions are spanning several orders of magnitude. Traditional filter-based methods tend to include many outliers into the backbone. What is more, they often suffer from the computational inefficiency-the exhaustive search of all nodes or edges is often prohibitively expensive. In this work, we propose a Local Search based Backbone Extraction Heuristic (LS-BEH) to find the backbone in a complex weighted network. First, a strict filtering rule is carefully designed to determine edges to be preserved or discarded. Second, we present a local search model to examine part of edges in an iterative way. Experimental results on two real-life networks demonstrate the advantage of LS-BEH over the classic disparity filter method by either effectiveness or efficiency validity.
Zhan Bu, Zhiang Wu 0001, Liqiang Qian, Jie Cao 0001, Guandong Xu
ASONAM2
2014 Detecting Genuine Communities from Large-Scale Social Networks: A Pattern-Based Method
abstract
Community detection is a long-standing yet very difficult task in social network analysis. It becomes more challenging as many online social networking sites are evolving into super-large scales. Numerous methods have been proposed for community detection from massive networks, but how to reconcile the partitioning efficiency and the community quality remains an open problem. In this paper, we attempt to address this challenge by introducing a COSine-pattern-based COMmunity extraction framework: COSCOM. The COSCOM adopts an extracting view of community detection. It first extracts the so-called asymptotically equivalent structures (AESs) from networks, from which the nodes are further partitioned into crisp communities using any of the existing methods. Specifically, we prove that an AES is a very tight group of nodes, and is actually a cosine pattern defined by the extended cosine similarity. A novel cosine-pattern mining algorithm based on the ordered anti-monotone of cosine similarity is thus proposed for the efficient extraction of AESs. Experiments on various real-world social networks demonstrate the advantage of the extracting view of community detection. In particular, COSCOM shows merits in detecting genuine communities by either internal or external validity.
Zhiang Wu 0001, Jie Cao 0001, Junjie Wu 0002, Youquan Wang
Comput. J.1
2014 Scaling up cosine interesting pattern discovery: A depth-first method
Jie Cao 0001, Zhiang Wu 0001, Junjie Wu 0002
Inf. Sci.2
2014 Efficient and robust large medical image retrieval in mobile cloud computing environment
Yi Zhuang 0001, Nan Jiang 0002, Zhiang Wu 0001, Qing Li 0001, Dickson K. W. Chiu, Hua Hu 0001
Inf. Sci.3
2014 Efficient Personalized Probabilistic Retrieval of Chinese Calligraphic Manuscript Images in Mobile Cloud Environment
abstract
Ancient language manuscripts constitute a key part of the cultural heritage of mankind. As one of the most important languages, Chinese historical calligraphy work has contributed to not only the Chinese cultural heritage but also the world civilization at large, especially for Asia. To support deeper and more convenient appreciation of Chinese calligraphy works, based on our previous work on the probabilistic retrieval of historical Chinese calligraphic character manuscripts repositories, we propose a system framework of the multi-feature-based Chinese calligraphic character images probabilistic retrieval in the mobile cloud network environment, which is called the DPRC . To ensure retrieval efficiency, we further propose four enabling techniques: (1) DRL-based probability propagation, (2) optimal data placement scheme, (3) adaptive data robust transmission algorithm, and (4) index support filtering scheme. Comprehensive experiments are conducted to testify the effectiveness and efficiency of our proposed DPRC method.
Yi Zhuang 0001, Qing Li 0001, Dickson K. W. Chiu, Zhiang Wu 0001
ACM Trans. Asian Lang. Inf. Process.4
2013 How Many Zombies Around You?
abstract
Recent years have witnessed the explosive growth of online social media. Weibo, a famous "Chinese Twitter", has attracted over half billion users in less than four years. Among them are zombie users or bogus users, who are seemingly active common users but actually marionettes manipulated by intelligent software for economic interests. To probe such users thus becomes critically important for a healthy Weibo, but the existing studies along this line are still in initial stage due to the serious lack of labeled zombies and the limited attributes for user profiling. In light of this, in this paper, we figure out a commercial way for training set labeling, and propose a two-stage cascading model called ProZombie for zombie user recognition. ProZombie decomposes the training/predicting process into fast and refined phases in cascade, which greatly improves the modeling efficiency without sacrificing the accuracy. Moreover, 35 attributes including 16 newly proposed ones are employed for a panoramic description of Weibo users. Experiments on real-world labeled Weibo users demonstrate the effectiveness and efficiency of ProZombie. More interestingly, two case studies based on ProZombie successfully unveil the zombies hidden around common users, and their impact to information propagation on Weibo. To our best knowledge, this study is among the first to quantify these interesting observations on Weibo.
Hongfu Liu 0001, Hao Lin 0002, Junjie Wu 0002, Zhiang Wu 0001
ICDM5
2013 SEA: a system for event analysis on chinese tweets
abstract
Recent years have witnessed the explosive growth of online social media. Weibo, a famous "Chinese Twitter", has attracted over 0.5 billion users in less than four years, with more than 1000 tweets generated in every second. These tweets are informative but very fragmented, and thus would be better archived from an event perspective, as done by Weibo itself in the "Micro-Topic" program. This effort, however, is yet far from satisfaction for not providing enough analytical power to events. In light of this, in this demo paper, we propose SEA, a System for Event Analysis on Chinese tweets. In general, SEA is an event-centric, multi-functional platform that conducts panoramic analysis on Weibo events from various aspects, including the semantic information of the events, the temporal and spatial trends, the public sentiments, the hidden sub-events, the key users in the event diffusion and their preferences, etc. These functions are enabled by the integration of various analytical models and by the noSQL techniques adopted purposefully for massive tweets management. Finally, a case study on the "Spring Festival" event demonstrates the effectiveness of SEA. To our best knowledge, SEA is the first third-party system that provides panoramic analysis to Weibo events.
Yaqiong Wang, Hongfu Liu 0001, Hao Lin 0002, Junjie Wu 0002, Zhiang Wu 0001, Jie Cao 0001
KDD5
2013 A Cloud System for Community Extraction from Super-Large Scale Social Networks
Zhiang Wu 0001, Haicheng Tao, Youquan Wang, Changjian Fang, Jie Cao 0001
WISE (2)1
2013 Community Detection in Multi-relational Social Networks
Zhiang Wu 0001, Wenpeng Yin 0001, Jie Cao 0001, Guandong Xu, Alfredo Cuzzocrea
WISE (2)1
2013 A novel noise filter based on interesting pattern mining for bag-of-features images
Zhiang Wu 0001, Jie Cao 0001, Haicheng Tao, Yi Zhuang 0001
Expert Syst. Appl.1
2013 Hybrid Collaborative Filtering algorithm for bidirectional Web service recommendation
Jie Cao 0001, Zhiang Wu 0001, Youquan Wang, Yi Zhuang 0001
Knowl. Inf. Syst.2
2013 Towards information-theoretic K-means clustering for image indexing
Jie Cao 0001, Zhiang Wu 0001, Junjie Wu 0002
Signal Process.2
2013 SAIL: Summation-bAsed Incremental Learning for Information-Theoretic Text Clustering
abstract
Information-theoretic clustering aims to exploit information-theoretic measures as the clustering criteria. A common practice on this topic is the so-called Info-Kmeans, which performs K-means clustering with KL-divergence as the proximity function. While expert efforts on Info-Kmeans have shown promising results, a remaining challenge is to deal with high-dimensional sparse data such as text corpora. Indeed, it is possible that the centroids contain many zero-value features for high-dimensional text vectors, which leads to infinite KL-divergence values and creates a dilemma in assigning objects to centroids during the iteration process of Info-Kmeans. To meet this challenge, in this paper, we propose a Summation-bAsed Incremental Learning (SAIL) algorithm for Info-Kmeans clustering. Specifically, by using an equivalent objective function, SAIL replaces the computation of KL-divergence by the incremental computation of Shannon entropy. This can avoid the zero-feature dilemma caused by the use of KL-divergence. To improve the clustering quality, we further introduce the variable neighborhood search scheme and propose the V-SAIL algorithm, which is then accelerated by a multithreaded scheme in PV-SAIL. Our experimental results on various real-world text collections have shown that, with SAIL as a booster, the clustering performance of Info-Kmeans can be significantly improved. Also, V-SAIL and PV-SAIL indeed help improve the clustering quality at a lower cost of computation.
Jie Cao 0001, Zhiang Wu 0001, Junjie Wu 0002, Hui Xiong 0001
IEEE Trans. Cybern.2
2013 Shilling attack detection utilizing semi-supervised learning method for collaborative recommender system
Jie Cao 0001, Zhiang Wu 0001, Yanchun Zhang
World Wide Web2
2012 Predicting Driving Direction with Weighted Markov Model
Jie Cao 0001, Zhiang Wu 0001, Guangyan Huang, Jingjun Li
ADMA3
2012 Towards a Tricksy Group Shilling Attack Model against Recommender Systems
Youquan Wang, Zhiang Wu 0001, Jie Cao 0001, Changjian Fang
ADMA2
2012 Personalized Clustering for Social Image Search Results Based on Integration of Multiple Features
Yi Zhuang 0001, Dickson K. W. Chiu, Nan Jiang 0002, Guochang Jiang, Zhiang Wu 0001
ADMA5
2012 Efficient Probabilistic Image Retrieval Based on a Mixed Feature Model
Yi Zhuang 0001, Zhiang Wu 0001, Nan Jiang 0002, Guochang Jiang, Dickson K. W. Chiu, Hua Hu 0001
APWeb2
2012 HySAD: a semi-supervised hybrid shilling attack detector for trustworthy product recommendation
abstract
Shilling attackers apply biased rating profiles to recommender systems for manipulating online product recommendations. Although many studies have been devoted to shilling attack detection, few of them can handle the hybrid shilling attacks that usually happen in practice, and the studies for real-life applications are rarely seen. Moreover, little attention has yet been paid to modeling both labeled and unlabeled user profiles, although there are often a few labeled but numerous unlabeled users available in practice. This paper presents a Hybrid Shilling Attack Detector, or HySAD for short, to tackle these problems. In particular, HySAD introduces MC-Relief to select effective detection metrics, and Semi-supervised Naive Bayes (SNB_lambda) to precisely separate Random-Filler model attackers and Average-Filler model attackers from normal users. Thorough experiments on MovieLens and Netflix datasets demonstrate the effectiveness of HySAD in detecting hybrid shilling attacks, and its robustness for various obfuscated strategies. A real-life case study on product reviews of Amazon.cn is also provided, which further demonstrates that HySAD can effectively improve the accuracy of a collaborative-filtering based recommender system, and provide interesting opportunities for in-depth analysis of attacker behaviors. These, in turn, justify the value of HySAD for real-world applications.
Zhiang Wu 0001, Junjie Wu 0002, Jie Cao 0001, Dacheng Tao
KDD1
2012 Pick-Up Tree Based Route Recommendation from Taxi Trajectories
Zhiang Wu 0001, Yi Zhuang 0001, Jie Cao 0001, Jingui Pan
WAIM2
2012 Bandwidth-Aware Medical Image Retrieval in Mobile Cloud Computing Network
Yi Zhuang 0001, Nan Jiang 0002, Zhiang Wu 0001, Dickson K. W. Chiu, Guochang Jiang, Hua Hu 0001
WAIM3
2012 Dynamic multi-resource advance reservation in grid environment
Junzhou Luo, Zhiang Wu 0001, Jiuxin Cao, Tian Tian 0009
J. Supercomput.2
2011 Dynamic Advance Reservation for Grid System Using Resource Pools
Zhiang Wu 0001, Jie Cao 0001, Youquan Wang
NPC1
2011 Semi-SAD: applying semi-supervised learning to shilling attack detection
abstract
Collaborative filtering (CF) based recommender systems are vulnerable to shilling attacks. In some leading e-commerce sites, there exists a large number of unlabeled users, and it is expensive to obtain their identities. Existing research efforts on shilling attack detection fail to exploit these unlabeled users. In this article, Semi-SAD, a new semi-supervised learning based shilling attack detection algorithm is proposed. Semi-SAD is trained with the labeled and unlabeled user profiles using the combination of naïve Bayes classifier and EM-», augmented Expectation Maximization (EM). Experiments on MovieLens datasets show that our proposed Semi-SAD is efficient and effective.
Zhiang Wu 0001, Jie Cao 0001, Youquan Wang
RecSys1
2010 An Improved Protocol for Deadlock and Livelock Avoidance Resource Co-allocation in Network Computing
Jie Cao 0001, Zhiang Wu 0001
World Wide Web2
2009 An adaptive algorithm for QoS-aware service composition in grid environments
Junzhou Luo, Jingya Zhou, Zhiang Wu 0001
Serv. Oriented Comput. Appl.3
2008 QoS adaption aware algorithm for grid service selection
abstract
Grid service composition has been recognized as a flexible way for resource sharing and application integration since appearance of service-oriented architecture. Approaches are needed to select service candidates with various Quality of Service (QoS) levels according to use’s performance requirements. In this paper We model this problem as the Multi-Constrained Optimal Path Selection Problem (MCOP), due to the dynamic property of grid service, adaptive mechanism is introduced to ensure the whole QoS when some service candidates fail. An algorithm QAGSS is proposed. Simulation results show that QAGSS has greater success rate and lower cost than previous algorithms.
Jingya Zhou, Junzhou Luo, Zhiang Wu 0001
CSCWD3
2007 QoS Deviation Distance Based Negotiation Algorithm in Grid Resource Advance Reservation
abstract
How to guarantee user's QoS (Quality of Service) demands becomes increasingly important in service-oriented grid environment. Current research on grid resource advance reservation, a well-known and effective mechanism to guarantee QoS, forces on proposing theoretic architecture to support advance reservation. Detailed research on advance reservation is scarce. For that, SNAP (Service Negotiation and Acquisition Protocol) is extended to support QoS description in fine granularity and new calculation method of QoS deviation distance is proposed. Then, advance reservation state switching is analyzed and QoS deviation distance based negotiation algorithm is addressed. Preliminary results show that proposed negotiation algorithm that considers QoS deviation distance can produce remarkable improvement in user satisfaction and resource utilization.
Zhiang Wu 0001, Junzhou Luo, Fang Dong 0001, Xudong Ni
CSCWD1
2007 Dynamic Multi-resource Advance Reservation in Grid Environment
Zhiang Wu 0001, Junzhou Luo
NPC1
2006 The Measurement Model of Grid QoS
abstract
As QoS-based resource management is becoming increasingly hot, the research on grid QoS plays a fundamental role in service-oriented grid environment. Grid QoS is classified into five categories in virtual organization layer in our early research. In this paper, we propose a kind of scheme to measure grid QoS based on this classification. In the scheme, the grid QoS parameters are quantified respectively. And QoS of grid application is put forward and the measurement model is applied to scheduling heuristics, so that grid QoS can be used as a critical part of a scheduling heuristic's objective function and to implement, analyze and compare resource management systems. Finally, we make a simple simulation and results indicate that applying the measuring model to scheduling heuristic as objective function can enhancing the performance of grid environments drastically
Zhiang Wu 0001, Junzhou Luo
CSCWD1