Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Kevin Chiew

dblp:34/7145 · DBLP profile ↗
← Back
28ranked-venue papers
2as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 16Artificial intelligence and machine learning · 9Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSecurity and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 86% Information retrieval · 7% Web and social media mining · 7%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
interactive data mining
0.412020
Interactive Rare-Category-of-Interest Mining from Large Datasets · AAAI 2020
Data mining
pattern mining
0.412020
Interactive Rare-Category-of-Interest Mining from Large Datasets · AAAI 2020
Data mining › predictive modeling › classification › class imbalance
rare category analysis
0.412020
Interactive Rare-Category-of-Interest Mining from Large Datasets · AAAI 2020
Data mining › clustering
arbitrary shape clustering
0.212014
Towards effective and efficient mining of arbitrary shaped clusters · ICDE 2014
Data mining
clustering
0.212014
Towards effective and efficient mining of arbitrary shaped clusters · ICDE 2014
Data mining › text mining
information extraction and text analysis
0.212013
Commodity query by snapping · SIGIR 2013
Information retrieval
multimedia analysis and retrieval
0.212013
Commodity query by snapping · SIGIR 2013
Web and social media mining › social tagging
social bookmarking
0.212013
Browse with a social web directory · SIGIR 2013
Data mining › clustering › hierarchical clustering
agglomerative clustering
0.112014
Towards effective and efficient mining of arbitrary shaped clusters · ICDE 2014
Data mining › text mining › text classification
hierarchical classification
0.012013
Browse with a social web directory · SIGIR 2013

Methods — techniques the papers use, named apart from their topics

pk similarity metric · 0.2mutual k-nearest neighbors · 0.2data set shrinking · 0.2statistical deduction · 0.2social bookmarking integration · 0.2semantic comprehension · 0.2ranking · 0.2pattern recognition · 0.2
YearPublicationVenuePosition
2020 Interactive Rare-Category-of-Interest Mining from Large Datasets
Zhenguang Liu, Sihao Hu, Yifang Yin, Jianhai Chen, Kevin Chiew, Zetian Wu
AAAI5
2018 Rare category exploration with noisy labels
Haiqin Weng, Kevin Chiew, Zhenguang Liu, Qinming He, Roger Zimmermann
Expert Syst. Appl.2
2018 Evaluation of local community metrics: from an experimental perspective
Lianhang Ma, Kevin Chiew, Hao Huang 0001, Qinming He
J. Intell. Inf. Syst.2
2016 Rare category exploration via wavelet analysis: Theory and applications
Zhenguang Liu, Kevin Chiew, Beibei Zhang 0007, Qinming He, Roger Zimmermann
Expert Syst. Appl.2
2016 A formalized framework for incorporating expert labels in crowdsourcing environment
Qingyang Hu, Qinming He, Hao Huang 0001, Kevin Chiew, Zhenguang Liu
J. Intell. Inf. Syst.4
2015 Rare Category Exploration on Linear Time Complexity
Zhenguang Liu, Hao Huang 0001, Qinming He, Kevin Chiew, Yunjun Gao
DASFAA (2)4
2015 Rare Category Detection Forest
abstract
Rare category detecion (RCD) aims to discover rare categories in a massive unlabeled data set with the help of a labeling oracle. A challenging task in RCD is to discover rare categories which are concealed by numerous data examples from major categories. Only a few algorithms have been proposed for this issue, most of which are on quadratic or cubic time complexity. In this paper, we propose a novel tree-based algorithm known as RCD-Forest with $$O(\varphi n \log {(n/s)})$$ time complexity and high query efficiency where n is the size of the unlabeled data set. Experimental results on both synthetic and real data sets verify the effectiveness and efficiency of our method.
Haiqin Weng, Zhenguang Liu, Kevin Chiew, Qinming He
KSEM3
2014 Towards effective and efficient mining of arbitrary shaped clusters
abstract
Mining arbitrary shaped clusters in large data sets is an open challenge in data mining. Various approaches to this problem have been proposed with high time complexity. To save computational cost, some algorithms try to shrink a data set size to a smaller amount of representative data examples. However, their user-defined shrinking ratios may significantly affect the clustering performance. In this paper, we present CLASP an effective and efficient algorithm for mining arbitrary shaped clusters. It automatically shrinks the size of a data set while effectively preserving the shape information of clusters in the data set with representative data examples. Then, it adjusts the positions of these representative data examples to enhance their intrinsic relationship and make the cluster structures more clear and distinct for clustering. Finally, it performs agglomerative clustering to identify the cluster structures with the help of a mutual k-nearest neighbors-based similarity metric called Pk. Extensive experiments on both synthetic and real data sets are conducted, and the results verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Yunjun Gao, Kevin Chiew, Lei Chen 0002, Qinming He
ICDE3
2014 Learning from Crowds under Experts' Supervision
Qingyang Hu, Qinming He, Hao Huang 0001, Kevin Chiew, Zhenguang Liu
PAKDD (1)4
2014 Rare Category Detection on O(dN) Time Complexity
Zhenguang Liu, Hao Huang 0001, Qinming He, Kevin Chiew, Lianhang Ma
PAKDD (2)4
2014 Recovering Missing Labels of Crowdsourcing Workers
abstract
Data sets collected from crowdsourcing platforms are well known for their cheap costs. But cheap costs may lead to low quality, i.e., labels may be incorrect or missing. Most of the existing work focuses on modeling the labeling errors of crowd workers, but missing labels can also cause problems when modeling the data. In this paper, we present an algorithm to predict the missing labels of crowd workers, in which we adopt thoughts from semi-supervised learning and utilize the particular consistency between crowd workers. We also define the consistency between workers by crowd labels and develop an algorithm to learn them from the data automatically. Experiments on both benchmark and real data show that our algorithm outperforms traditional semi-supervised learning algorithms in predicting missing labels, and the recovered crowd labels are capable of predicting the ground truth and reflecting real properties of crowd workers.
Qingyang Hu, Kevin Chiew, Hao Huang 0001, Qinming He
SDM2
2014 Rare category exploration
Hao Huang 0001, Kevin Chiew, Yunjun Gao, Qinming He, Qing Li 0001
Expert Syst. Appl.2
2014 Unsupervised analysis of top-k core members in poly-relational networks
Hao Huang 0001, Yunjun Gao, Kevin Chiew, Qinming He, Baihua Zheng
Expert Syst. Appl.3
2014 Prior-free rare category detection: More effective and efficient solutions
Zhenguang Liu, Kevin Chiew, Qinming He, Hao Huang 0001, Butian Huang
Expert Syst. Appl.2
2014 Toward seed-insensitive solutions to local community detection
Lianhang Ma, Hao Huang 0001, Qinming He, Kevin Chiew, Zhenguang Liu
J. Intell. Inf. Syst.4
2014 Mining regional co-location patterns with kNNG
Feng Qian 0006, Kevin Chiew, Qinming He, Hao Huang 0001
J. Intell. Inf. Syst.2
2013 AAGA: Affinity-Aware Grouping for Allocation of Virtual Machines
abstract
Virtualization technology enables various application services to be distributed and encapsulated within virtual machines (VMs), which are dynamically allocated to physical machines (PMs) in cloud computing environments. However, in many existing virtualized systems, the limited network bandwidth often becomes a bottleneck resource, leading to the intensification of network competition and the performance degradation for communication or data intensive applications. Aiming at reducing communication overheads and improving the application performance, in this paper, we propose an Affinity-Aware Grouping method for Allocation of VMs (AAGA). Firstly, we identity and model the problem of affinity-aware grouping-based allocation for virtual machines, and propose a detailed grouping method based on which a heuristic bin packing algorithm is used to deploy VM groups into PMs. In order to demonstrate the effectiveness of AAGA, we create multiple real virtual clusters (multi-VCs) with 56 VMs running multi-VM applications and compare application performance with Non-Affinity-aware Grouping-based Allocation methods (NAGA). Experimental results show that AAGA achieves better performance than NAGA.
Jianhai Chen, Kevin Chiew, Deshi Ye, Liangwei Zhu, Wenzhi Chen
AINA2
2013 GMAC: A Seed-Insensitive Approach to Local Community Detection
Lianhang Ma, Hao Huang 0001, Qinming He, Kevin Chiew, Jianan Wu, Yanzhe Che
DaWaK4
2013 Discovery of Regional Co-location Patterns with k-Nearest Neighbor Graph
Feng Qian 0006, Kevin Chiew, Qinming He, Hao Huang 0001, Lianhang Ma
PAKDD (1)2
2013 Commodity query by snapping
abstract
Commodity information such as prices and public reviews is always the concern of consumers. Helping them conveniently acquire these information as an instant reference is often of practical significance for their purchase activities. Nowadays, Web 2.0, linked data clouds, and the pervasiveness of smart hand held devices have created opportunities for this demand, i.e., users could just snap a photo of any commodity that is of interest at anytime and anywhere, and retrieve the relevant information via their Internet-linked mobile devices. Nonetheless, compared with the traditional keyword-based information retrieval, extracting the hidden information related to the commodities in photos is a much more complicated and challenging task, involving techniques such as pattern recognition, knowledge base construction, semantic comprehension, and statistic deduction. In this paper, we propose a framework to address this issue by leveraging on various techniques, and evaluate the effectiveness and efficiency of this framework with experiments on a prototype.
Hao Huang 0001, Yunjun Gao, Kevin Chiew, Qinming He, Lu Chen 0001
SIGIR3
2013 Browse with a social web directory
abstract
Browse with either web directories or social bookmarks is an important complementation to search by keywords in web information retrieval. To improve users' browse experiences and facilitate the web directory construction, in this paper, we propose a novel browse system called Social Web Directory (SWD for short) by integrating web directories and social bookmarks. In SWD, (1) web pages are automatically categorized to a hierarchical structure to be retrieved efficiently, and (2) the popular web pages, hottest tags, and expert users in each category are ranked to help users find information more conveniently. Extensive experimental results demonstrate the effectiveness of our SWD system.
Hao Huang 0001, Yunjun Gao, Lu Chen 0001, Kevin Chiew, Qinming He
SIGIR5
2013 EDA: an enhanced dual-active algorithm for location privacy preservation inmobile P2P networks
abstract
Various solutions have been proposed to enable mobile users to access location-based services while preserving their location privacy. Some of these solutions are based on a centralized architecture with the participation of a trustworthy third party, whereas some other approaches are based on a mobile peer-to-peer (P2P) architecture. The former approaches suffer from the scalability problem when networks grow large, while the latter have to endure either low anonymization success rates or high communication overheads. To address these issues, this paper deals with an enhanced dual-active spatial cloaking algorithm (EDA) for preserving location privacy in mobile P2P networks. The proposed EDA allows mobile users to collect and actively disseminate their location information to other users. Moreover, to deal with the challenging characteristics of mobile P2P networks, e.g., constrained network resources and user mobility, EDA enables users (1) to perform a negotiation process to minimize the number of duplicate locations to be shared so as to significantly reduce the communication overhead among users, (2) to predict user locations based on the latest available information so as to eliminate the inaccuracy problem introduced by using some out-of-date locations, and (3) to use a latest-record-highest-priority (LRHP) strategy to reduce the probability of broadcasting fewer useful locations. Extensive simulations are conducted for a range of P2P network scenarios to evaluate the performance of EDA in comparison with the existing solutions. Experimental results demonstrate that the proposed EDA can improve the performance in terms of anonymity and service time with minimized communication overhead.
Yanzhe Che, Kevin Chiew, Xiaoyan Hong, Qiang Yang 0004, Qinming He
J. Zhejiang Univ. Sci. C2
2013 CLOVER: a faster prior-free approach to rare-category detection
Hao Huang 0001, Qinming He, Kevin Chiew, Feng Qian 0006, Lianhang Ma
Knowl. Inf. Syst.3
2012 Spatial co-location pattern discovery without thresholds
Feng Qian 0006, Qinming He, Kevin Chiew, Jiangfeng He
Knowl. Inf. Syst.3
2010 An Approach to Image Spam Filtering Based on Base64 Encoding and N-Gram Feature Extraction
abstract
As compared with text spam, the image spam is a variant which is invented to escape from traditional text-based spam classification and filtering. Various approaches to image spam filtering have been proposed with respective advantages and drawbacks in terms of time cost and efficiency. In this paper, we propose a new approach based on Base64 encoding of image files and n-gram technique for feature extraction. By transforming normal images into Base64 presentation, we try to extract features of an image with n-gram technique. With these features we train an SVM (support vector machine) which shows effectiveness and efficiency in detecting spam images from legitimate images. With an online shared personal corpus of images as the input, experimental results show that our approach, in comparison with some of the existing methods of feature extraction, can achieve very high performance for image spam classification in terms of some basic measures such as accuracy, precision, and recall. Moreover, our approach shows its practicability by taking less running time for image spam classification in comparison to other methods.
Congfu Xu, Yafang Chen, Kevin Chiew
ICTAI (1)3
2010 Privacy Disclosure Analysis and Control for 2D Contingency Tables Containing Inaccurate Data
Kevin Chiew, Yingjiu Li, Yanjiang Yang
Privacy in Statistical Databases2
2009 Multistage Off-Line Permutation Packet Routing on a Mesh: An Approach with Elementary Mathematics
Kevin Chiew, Yingjiu Li
J. Comput. Sci. Technol.1
2009 Scheduling and Routing of AMOs in an Intelligent Transport System
abstract
Autonomous moving objects (AMOs), such as automated guided vehicles (AGVs) and autonomous robots, have widely been used in the industry for decades. In an intelligent transport system with a great number of AMOs involved, it is important to eliminate potential congestion and deadlocks among AMOs to maintain a well-organized traffic flow. In this paper, we propose an algorithm that adapts bitonic merge sort algorithm for concurrent scheduling and routing of a great number (i.e., 4n2) of AMOs on an ntimesn mesh topology of path network without congestion or deadlocks among AMOs during their moves. The results are tested by experiments with randomly generated data and the comparison of a related model.
Kevin Chiew, Shaowen Qin
IEEE Trans. Intell. Transp. Syst.1