VLDB 2026 Research / reviewers in the wild / expert
Chi-Yin Chow
dblp:77/6228
· DBLP profile ↗
77ranked-venue papers
14as first author
10since 2021 · last 2026
0000-0002-9566-0743ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 50 · 9 first-author · 2 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 7 · 2 since 2021Computer networks · 6 · 3 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Companion learning Networks: A deep reinforcement learning algorithm with partner networks
Jinfeng Bu, Jia-Dong Zhang, Chi-Yin Chow |
Expert Syst. Appl. | 5 |
| 2024 | MiST: Enhancing Traffic Predictions with a Mixing Spatio-temporal Neural NetworkabstractAccurately predicting traffic conditions is vital for smart city development, yet it remains challenging due to the intricate spatio-temporal dependencies in road networks. Existing works often propose intra-mixing deep learning-based prediction models for individual nodes and share parameters among them or spatial intermixing deep learning-based models for traffic predictions. However, these approaches may neglect essential principles of information exchange in traffic flow or capture useless or even erroneous spatio-temporal dependencies. To address these limitations, we propose a Mixing Spatio-Temporal neural network (MiST) for enhancing traffic predictions. In MiST, we propose (i) a temporal encoder that embeds the traffic data along with periodic features, (ii) a spatial encoder that embeds the positional information in graph and hypergraph spectral domains, as well as spatial node identities, and (iii) a mixing spatio-temporal encoder that merges the diverse features provided by the temporal and spatial encoders. Our empirical evaluations on real-world traffic prediction tasks, including flow and speed predictions, validate the superiority of MiST, underscoring its innovative contribution to traffic prediction methodologies. Zhixiang He, Mengzan Gong, Jia-Dong Zhang, Xiliang Liu, Chi-Yin Chow, Ning Li 0041 |
SIGSPATIAL/GIS | 5 |
| 2023 | Pairwise and Hyper-correlations Based Spatiotemporal Neural Networks for Traffic Speed PredictionsabstractThe problem of traffic speed predictions is still very challenging due to the complex and dynamic urban traffic conditions. Many existing works have implied the importance of integrating spatial correlations into models to explore nonlinear spatio-temporal dependencies and make traffic predictions in near future. However, some of the works only consider pairwise correlations and cannot model the hidden information among multiple nodes well, and the others only consider hyper-correlations (that can be shared by more than two nodes) and discount the role of the pairwise ones for propagating spatial dependencies. Therefore, we propose a Spatio-Temporal neural nEtwork based on both Pairwise and Hyper-correlations (STEPH) for traffic speed predictions. It is distinguished primarily by incorporating both types of spatial correlations into temporal information and designing new hybrid spatio-temporal blocks in neural networks to effectively overcome the challenge. Experiments on two real-world traffic datasets demonstrate the effectiveness of the proposed model, and show its superiority performance compared to other state-of-the-art baselines. Zhixiang He, Jia-Dong Zhang, Chi-Yin Chow, Ning Li 0041, Xiliang Liu, Pengfei Lin 0001 |
MDM | 3 |
| 2023 | Feature decomposition and enhancement for unsupervised medical ultrasound image denoising and instance segmentation
Kam-yiu Lam, Chi-Yin Chow |
Appl. Intell. | 4 |
| 2023 | DeepMAG: Deep reinforcement learning with multi-agent graphs for flexible job shop scheduling
Jia-Dong Zhang, Zhixiang He, Wing-Ho Chan, Chi-Yin Chow |
Knowl. Based Syst. | 4 |
| 2023 | H3Rec: Higher-Order Heterogeneous and Homogeneous Interaction Modeling for Group Recommendations of Web ServicesabstractRecommendations are important web services in the era of information explosion. Particularly, group recommendations aim to suggest new items to groups such that the members of groups are likely interested in. However, existing works still suffer from sparsity and cold-start issues (e.g., cold-start groups or items) for groups with few interactions on items. Most of them model the preferences or features of entities (i.e., users, items and groups) from heterogeneous interactions (i.e., user-item, group-item and user-group interactions) between two distinct types of entities, while ignoring the homogeneous interactions (i.e., user-user, item-item and group-group interactions) between entities of one type. To this end, we propose a new model, called H3Rec, which learns the representations of entities by developing two graph embedding layers based on an interaction graph of all entities. Specifically, the two graph embedding layers make full use of the hidden information in theHigher-orderHeterogeneous andHomogeneous interactionsof the graph. Therefore, H3Rec can alleviate the sparsity and cold-start issues and improve the performance of group recommendations. The experimental results on two real world datasets in different domains show the superiority of H3Rec in group recommendations, especially for cold-start groups and items. Zhixiang He, Chi-Yin Chow, Jia-Dong Zhang, Kam-yiu Lam |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | MVLevelDB: Using Log-Structured Tree to Support Temporal Queries in IoTabstractAlthough log-structured merge trees (LSM-trees) are commonly adopted in many NoSQLs as they can significantly improve the write performance in updating a database, most of the proposed LSM-trees are concentrated on storing a single version of data. On the other hand, in many Internet of Things (IoT) applications, it is important to maintain the old versions of data in addition to the latest version. In this article, we introduce our design and implementation of an enhancement of LevelDB to multiversion LevelDB (called MVLevelDB) with the purpose to efficiently support temporal queries on multiversion data in IoT applications. Based on the temporal consistency, we formulated the log-structured multiversion tree (LSMV-tree) to be implemented into MVLevelDB. In LSMV-tree, each data version is associated with two time-stamps to define its validity interval, and both the data versions and the components are time-sorted to improve the efficiency in searching data in processing temporal queries. To handle the problem of multicomponents data versions, we designed the data version duplication (DvD) method in which a data version will be duplicated in the next component if it is valid while its component is being flushed from the main memory to disk storage. Extensive experiments using a benchmark program have been performed to investigate the performance of MVLevelDB as compared with LevelDB both in writing and reading data. Xiaofei Zhao 0002, Kam-yiu Lam, Chun Jiang Zhu, Chi-Yin Chow, Tei-Wei Kuo |
IEEE Internet Things J. | 4 |
| 2021 | DA-BERT: Enhancing Knowledge Selection in Dialog via Domain Adapted BERT with Dynamic Masking ProbabilityabstractOne of the most challenging tasks in Knowledge-grounded Task-oriented Dialog Systems (KTDS) is the knowledge selection task, which aims to find the proper knowledge snippets to handle user requests. This paper proposes DA-BERT to employ pre-trained BERT with domain adaptive training and newly proposed dynamic masking probability to deal with knowledge selection in KTDS. Domain adaptive training minimizes the domain gap between the general text data BERT is pre-trained on and the dialog-knowledge joint data; and dynamic masking probability enhances the training in an easy-to-hard mode. Experimental results on the benchmark dataset show that our proposed training method outperforms the state-of-the-art models with large margins across all the evaluation metrics. Moreover, we analyze the bad case of our method and recognize several typical errors in the bad case set to facilitate further research in this direction. Chi-Yin Chow, Ning Li 0041 |
SMARTCOMP | 2 |
| 2021 | STNN: A Spatio-Temporal Neural Network for Traffic PredictionsabstractTraffic is very important to route planning and people’s daily lives. Traffic prediction is still very challenging as it is affected by many complex factors including dynamic spatio-temporal dependencies and external factors (e.g., road types and nearby points of interest) in the road network. Dynamic spatio-temporal dependencies simultaneously contain spatial and temporal dependencies. Existing models for predicting traffic of links only consider the spatial dependencies from the perspective of links or the whole road network by ignoring the spatial dependencies among regions. To this end, this paper proposes a new Spatio-Temporal Neural Network (STNN) with the encoder-decoder architecture to improve the accuracy of traffic predictions by additionally taking into account the region-based spatial dependencies and external factors. Specifically, STNN learns dynamic spatio-temporal dependencies from historical traffic time series via an encoder in the perspective of the road network, with two spatial models, i.e.,region-based spatial modelandlink-based spatial attention modelin the perspectives of regions and links, respectively. Further, STNN decodes the output from the encoder via a decoder with a temporal attention model for recording long-term dependencies and fuses external factors in the road network, to improve network-wide traffic predictions. We conduct extensive experiments to evaluate the performance of STNN on three real-world traffic datasets, which shows that STNN is significantly better than the state-of-the-art models. Zhixiang He, Chi-Yin Chow, Jia-Dong Zhang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Enabling Probabilistic Differential Privacy Protection for Location RecommendationsabstractThe sequential pattern in the human movement is one of the most important aspects for location recommendations in geosocial networks. Existing location recommenders have to access users' raw check-in data to mine their sequential patterns that raises serious location privacy breaches. In this paper, we propose a new Privacy-preserving LOcation REcommendation framework (PLORE) to address this privacy challenge. First, we employ the nnth-order additive Markov chain to exploit users' sequential patterns for location recommendations. Further, we contrive the probabilistic differential privacy mechanism to reach a good trade-off between high recommendation accuracy and strict location privacy protection. Finally, we conduct extensive experiments to evaluate the performance of PLORE using three large-scale real-world data sets. Extensive experimental results show that PLORE provides efficient and highly accurate location recommendations, and guarantees strict privacy protection for user check-in data in geosocial networks. Jia-Dong Zhang, Chi-Yin Chow |
IEEE Trans. Serv. Comput. | 2 |
| 2020 | GAMIT: A New Encoder-Decoder Framework with Graphical Space and Multi-grained Time for Traffic PredictionsabstractNowadays, many researchers study on characterizing complex and dynamic traffic environments by modeling the spatio-temporal dependencies in a road network for traffic predictions. However, existing works fail to investigate comprehensive spatio-temporal dependencies, because most of them ignore the spatial dependencies from the topological graph structure information in the road network, or only consider the temporal dependencies between fine-grained time slots whereas ignore those among coarse-grained time periods, in which a time period is composed of a number of time slots. To this end, we propose a new encoder-decoder framework called GAMIT with graphical space and multi-grained time by developing a spatiotemporal recurrent neural network (STRNN) for traffic predictions, where the graphical space consists of spatial networks which represent road networks. STRNN first devises a spatiotemporal convolution block to capture the fine-grained spatiotemporal dependencies between time slots in the road network. Then, STRNN uses the recurrent architecture to catch the coarsegrained spatio-temporal dependencies among time periods. The encoder finally applies STRNN to learn the multi-grained spatiotemporal dependencies which are fed into the decoder for computing traffic predictions based on STRNN as well. To evaluate the performance of GAMIT, we conduct extensive experiments on two real traffic flow datasets. Experimental results show that GAMIT outperforms the state-of-the-art traffic prediction models. Zhixiang He, Chi-Yin Chow, Jia-Dong Zhang |
IEEE BigData | 2 |
| 2020 | EMOVA: A Semi-supervised End-to-End Moving-Window Attentive Framework for Aspect Mining
Ning Li 0041, Chi-Yin Chow, Jia-Dong Zhang |
PAKDD (2) | 2 |
| 2020 | GAME: Learning Graphical and Attentive Multi-view Embeddings for Occasional Group RecommendationabstractGroup recommendation aims to suggest preferred items to a group of users rather than to an individual user. Most existing methods on group recommendation directly learn theinherent interests of groups and users orinherent features of items, i.e., independently modeling the inherent embeddings of groups, users or items. However, the independent view severely suffers from the cold-start problem when making recommendations for occasional groups that are temporally formed by a set of users and have few interactions on items. Actually, the groups, users and items are interdependent because they interact with one another. The interdependencies constitute an interaction graph that provides multiple views to model the embeddings of groups, users and items from their interacting counterparts to improve recommendation for occasional groups. To this end, we propose a model, named GAME to learn the Graphical and Attentive Multi-view Embeddings (i.e., representations) for the groups, users and items from the independent view and counterpart views based on the interaction graph. In the counterpart views, the embedding of a group, user or item is aggregated from the interacting counterparts based on an attention mechanism that derives the adaptive weight for each counterpart. For instance, a user's embedding may be aggregated from her interacting items or groups. Further, GAME applies neural collaborative filtering to investigate the interactions between the multi-view embeddings of groups (or users) and items for group recommendation. Finally, we conduct extensive experiments on two real datasets. The experimental results show that GAME outperforms other state-of-the-art models, especially on both cold-start groups (i.e., occasional groups) and cold-start items. Zhixiang He, Chi-Yin Chow, Jia-Dong Zhang |
SIGIR | 2 |
| 2019 | GRADI: Towards Group Recommendation Using Attentive Dual Top-Down and Bottom-Up InfluencesabstractMost of current group recommenders only consider the bottom-up influences, i.e., the preference of a group is greatly affected by the members in the group. For example, children usually dominate the preference of a family while senior experts often lead the preference of a professional group. However, in reality there also exist the top-down influences, i.e., a group inherently affects its every member, because a group often has some distinct themes which limit the preferences of all the members in the group. For instance, the members in a group for sports may prefer hiking and rock climbing, whereas the members in a group for entertainment would like to watch movies and play games. In other words, the influences between a group and its members are dual. To this end, this paper proposes a new model for Group Recommendation using Attentive Dual Influences (GRADI) that simultaneously explores both the bottom-up and top-down influences between a group and its members. The preference of a member in a group is represented as the group-specific member embedding by modeling the top-down influences from the group to the member. In addition, the preference of a group on a target item is represented as the item-specific group representation by considering the bottom-up influences from all the members to the group, where an attentive mechanism is developed to aggregate the preferences of all the members on a target item. Furthermore, GRADI investigates the interactions between groups and items with neural collaborative filtering. Results of extensive experiments conducted on two real-world datasets show that GRADI outperforms other state-of-the-art models. Zhixiang He, Chi-Yin Chow, Jia-Dong Zhang, Ning Li 0041 |
IEEE BigData | 2 |
| 2019 | STCNN: A Spatio-Temporal Convolutional Neural Network for Long-Term Traffic PredictionabstractAs many location-based applications provide services for users based on traffic conditions, an accurate traffic prediction model is very significant, particularly for long-term traffic predictions (e.g., one week in advance). As far, long-term traffic predictions are still very challenging due to the dynamic nature of traffic. In this paper, we propose a model, called Spatio-Temporal Convolutional Neural Network (STCNN) based on convolutional long short-term memory units to address this challenge. STCNN aims to learn the spatio-temporal correlations from historical traffic data for long-term traffic predictions. Specifically, STCNN captures the general spatio-temporal traffic dependencies and the periodic traffic pattern. Further, STCNN integrates both traffic dependencies and traffic patterns to predict the long-term traffic. Finally, we conduct extensive experiments to evaluate STCNN on two real-world traffic datasets. Experimental results show that STCNN is significantly better than other state-of-the-art models. Zhixiang He, Chi-Yin Chow, Jia-Dong Zhang |
MDM | 2 |
| 2019 | iMCRec: A multi-criteria framework for personalized point-of-interest recommendations
Chi-Yin Chow, Ran Wang 0001, Victor C. S. Lee |
Inf. Sci. | 2 |
| 2018 | Efficient evaluation of shortest travel-time path queries through spatial mashups
Detian Zhang, Chi-Yin Chow, An Liu 0002, Xiangliang Zhang 0001, Qingzhu Ding, Qing Li 0001 |
GeoInformatica | 2 |
| 2018 | TaxiRec: Recommending Road Clusters to Taxi Drivers Using Ranking-Based Extreme Learning MachinesabstractUtilizing large-scale GPS data to improve taxi services has become a popular research problem in the areas of data mining, intelligent transportation, geographical information systems, and the Internet of Things. In this paper, we utilize a large-scale GPS data set generated by over 7,000 taxis in a period of one month in Nanjing, China, and propose TaxiRec: a framework for evaluating and discovering the passenger-finding potentials of road clusters, which is incorporated into a recommender system for taxi drivers to seek passengers. In TaxiRec, the underlying road network is first segmented into a number of road clusters, a set of features for each road cluster is extracted from real-life data sets, and then a ranking-based extreme learning machine (ELM) model is proposed to evaluate the passenger-finding potential of each road cluster. In addition, TaxiRec can use this model with a training cluster selection algorithm to provide road cluster recommendations when taxi trajectory data is incomplete or unavailable. Experimental results demonstrate the feasibility and effectiveness of TaxiRec. Ran Wang 0001, Chi-Yin Chow, Victor C. S. Lee, Sam Kwong |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2018 | Entropy-Based Scheduling Policy for Cross Aggregate Ranking WorkloadsabstractMany data exploration applications require the ability to identify the top-k results according to a scoring function. We study a class of top-k ranking problems where top-k candidates in a dataset are scored with the assistance of another set. We call this class of workloads cross aggregate ranking. Example computation problems include evaluating the Hausdorff distance between two datasets, finding the medoid or radius within one dataset, and finding the closest or farthest pair between two datasets. In this paper, we propose a parallel and distributed solution to process cross aggregate ranking workloads. Our solution subdivides the aggregate score computation of each candidate into tasks while constantly maintains the tentative top-k results as an uncertain top-k result set. The crux of our proposed approach lies in our entropy-based scheduling technique to determine result-yielding tasks based on their abilities to reduce the uncertainty of the tentative result set. Experimental results show that our proposed approach consistently outperforms the best existing one in two different types of cross aggregate rank workloads using real datasets. Chengcheng Dai, Sarana Nutanong, Chi-Yin Chow, Reynold Cheng |
IEEE Trans. Serv. Comput. | 3 |
| 2017 | EventRec: Personalized Event Recommendations for Smart Event-Based Social NetworksabstractIn recent years, there has been a tremendous increase in the popularity of event-based social networks which allow social and physical interactions among their members. One major challenge for their members is the difficulty of searching events that meet their preferences from a large number of upcoming events. To tackle this challenge, we propose a personalized event recommendation framework called EventRec that exploits the geographical, social and temporal influences of events on users to generate personalized event recommendations. In EventRec, we model the influence of two-dimensional geographical location of an event using the Kernel Density Estimation method along with the popularity of the event location. Furthermore, the social influence in EventRec does not rely only on the relevance of a group to a user, but it also considers the relevance of the group to her friends. The geographical and social influences are integrated with the temporal influence that considers the preferences of the user and her friends on the days of the week and time of events. Our performance evaluation is conducted using two large Meetup.com data sets, and experimental results show that the quality of recommendations of EventRec outperforms the state-of-the-art event recommendation techniques. Tunde Joseph Ogundele, Chi-Yin Chow, Jia-Dong Zhang |
SMARTCOMP | 2 |
| 2017 | Coding-based cooperative caching in on-demand data broadcast environments
Houling Ji, Victor C. S. Lee, Chi-Yin Chow, Kai Liu 0001, Guoqing Wu 0004 |
Inf. Sci. | 3 |
| 2017 | Enabling Kernel-Based Attribute-Aware Matrix Factorization for Rating PredictionabstractIn recommender systems, one key task is to predict the personalized rating of a user to a new item and then return the new items having the top predicted ratings to the user. Recommender systems usually apply collaborative filtering techniques (e.g., matrix factorization) over a sparse user-item rating matrix to make rating prediction. However, the collaborative filtering techniques are severely affected by the data sparsity of the underlying user-item rating matrix and often confront the cold-start problems for new items and users. Since the attributes of items and social links between users become increasingly accessible in the Internet, this paper exploits the rich attributes of items and social links of users to alleviate the rating sparsity effect and tackle the cold-start problems. Specifically, we first propose a Kernel-based Attribute-aware Matrix Factorization model called KAMF to integrate the attribute information of items into matrix factorization. KAMF can discover the nonlinear interactions among attributes, users, and items, which mitigate the rating sparsity effect and deal with the cold-start problem for new items by nature. Further, we extend KAMF to address the cold-start problem for new users by utilizing the social links between users. Finally, we conduct a comprehensive performance evaluation for KAMF using two large-scale real-world data sets recently released in Yelp and MovieLens. Experimental results show that KAMF achieves significantly superior performance against other state-of-the-art rating prediction techniques. Jia-Dong Zhang, Chi-Yin Chow |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | Privacy-Preserving Location Sharing Services for Social NetworksabstractA common functionality of many location-based social networking applications is a location sharing service that allows a group of friends to share their locations. With a potentially untrusted server, such a location sharing service may threaten the privacy of users. Existing solutions for Privacy-Preserving Location Sharing Services (PPLSS) require a trusted third party that has access to the exact location of all users in the system or rely on expensive algorithms or protocols in terms of computational or communication overhead. Other solutions can only provide approximate query answers. To overcome these limitations, we propose a new encryption notion, called Order-Retrievable Encryption (ORE), for PPLSS for social networking applications. The distinguishing characteristics of our PPLSS are that it (1) allows a group of friends to share their exact locations without the need of any third party or leaking any location information to any server or users outside the group, (2) achieves low computational and communication cost by allowing users to receive the exact location of their friends without requiring any direct communication between users or multiple rounds of communication between a user and a server, (3) provides efficient query processing by designing an index structure for our ORE scheme, (4) supports dynamic location updates, and (5) provides personalized privacy protection within a group of friends by specifying a maximum distance where a user is willing to be located by his/her friends. Experimental results show that the computational and communication cost of our PPLSS is much better than the state-of-the-art solution. Roman Schlegel, Chi-Yin Chow, Qiong Huang 0001, Duncan S. Wong |
IEEE Trans. Serv. Comput. | 2 |
| 2016 | Efficient Evaluation of Shortest Travel-Time Path Queries in Road Networks by Optimizing Waypoints in Route Requests Through Spatial Mashups
Detian Zhang, Chi-Yin Chow, Qing Li 0001, An Liu 0002 |
APWeb (1) | 2 |
| 2016 | Exploring cell tower data dumps for supervised learning-based point-of-interest prediction (industrial paper)
Ran Wang 0001, Chi-Yin Chow, Victor C. S. Lee, Sarana Nutanong, Mingxuan Yuan |
GeoInformatica | 2 |
| 2016 | Location Data Management: A Tale of Two Systems and the "Next Destination"!abstractIn early 2000, we had the vision of ubiquitous location services, where each object is aware of its location, and continuously sends its location to a designated database server. This flood of location data opened the door for a myriad of location-based services that were considered visionary at that time, yet today they are a reality and have become ubiquitous. To realize our early vision, we identified two main challenges that needed to be addressed, namely, scalability and privacy. We have addressed these challenges through two main systems, PLACE and Casper. PLACE, developed at Purdue University from 2000 to 2005, set up the environment for built-in database support of scalable and continuous location-based services. The Casper system, developed at University of Minnesota from 2005 to 2010, was built inside the PLACE server allowing it to provide its high quality scalable service, while maintaining the privacy of its users' locations. This talk will take you through a time journey of location services from 2000 until today, and beyond, highlighting the development efforts of the PLACE and Casper systems, along with their impact on current and future research initiatives in both academia and industry. Mohamed F. Mokbel, Chi-Yin Chow, Walid G. Aref |
Proc. VLDB Endow. | 2 |
| 2016 | A Spatial Mashup Service for Efficient Evaluation of Concurrent k-NN QueriesabstractAlthough the travel time is the most important information in road networks, many spatial queries, e.g.,$k$-nearest-neighbor ($k$-NN) and range queries, for location-based services (LBS) are only based on the network distance. This is because it is costly for an LBS provider to collect real-time traffic data from vehicles or roadside sensors to compute the travel time between two locations. With the advance of web mapping services, e.g., Google Maps, Microsoft Bing Maps, and MapQuest Maps, there is an invaluable opportunity for using such services for processing spatial queries based on the travel time. In this paper, we propose a server-sideSpatialMashupService (SMS) that enables the LBS provider to efficiently evaluate$k$-NN queries in road networks using the route information and travel time retrieved from an external web mapping service. Due to the high cost of retrieving such external information, the usage limits of web mapping services, and the large number of spatial queries, we optimize the SMS for a large number of$k$-NN queries. We first discuss how the SMS processes a single$k$-NN query using two optimizations, namely,direction sharingandparallel requesting. Then, we extend them to process multiple concurrent$k$-NN queries and design a performance tuning tool to provide a trade-off between the query response time and the number of external requests and more importantly, to prevent a starvation problem in the parallel requesting optimization for concurrent queries. We evaluate the performance of the proposed SMS using MapQuest Maps, a real road network, real and synthetic data sets. Experimental results show the efficiency and scalability of our optimizations designed for the SMS. Detian Zhang, Chi-Yin Chow, Qing Li 0001, Xinming Zhang 0001, Yinlong Xu 0001 |
IEEE Trans. Computers | 2 |
| 2016 | Ambiguity-Based Multiclass Active LearningabstractMost existing works on active learning (AL) focus on binary classification problems, which limit their applications in various real-world scenarios. One solution to multiclass AL (MAL) is evaluating the informativeness of unlabeled samples by an uncertainty model and selecting the most uncertain one for query. In this paper, an ambiguity-based strategy is proposed to tackle this problem by applying a possibility approach. First, the possibilistic memberships of unlabeled samples in the multiple classes are calculated from the one-against-all-based support vector machine model. Then, by employing fuzzy logic operators, these memberships are aggregated into a new concept named k-order ambiguity, which estimates the risk of labeling a sample among k classes. Afterward, the k-order ambiguities are used to form an overall ambiguity measure to evaluate the uncertainty of the unlabeled samples. Finally, the sample with the maximum ambiguity is selected for query, and a new MAL strategy is developed. Experiments demonstrate the feasibility and effectiveness of the proposed method. Ran Wang 0001, Chi-Yin Chow, Sam Kwong |
IEEE Trans. Fuzzy Syst. | 2 |
| 2016 | Explaining Missing Answers to Top-k SQL QueriesabstractDue to the fact that existing database systems are increasingly more difficult to use, improving the quality and the usability of database systems has gained tremendous momentum over the last few years. In particular, the feature of explaining why some expected tuples are missing in the result of a query has received more attention. In this paper, we study the problem of explaining missing answers to top-k queries in the context of SQL (i.e., with selection, projection, join, and aggregation). To approach this problem, we use the query-refinement method. That is, given as inputs the original top-k SQL query and a set of missing tuples, our algorithms return to the user a refined query that includes both the missing tuples and the original query results. Case studies and experimental results show that our algorithms are able to return high quality explanations efficiently. Wenjian Xu, Zhian He, Eric Lo 0001, Chi-Yin Chow |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2016 | CRATS: An LDA-Based Model for Jointly Mining Latent Communities, Regions, Activities, Topics, and Sentiments from Geosocial Network DataabstractGeosocial networks like Yelp and Foursquare have been rapidly growing and accumulating plenty of data such as social links between users, user check-ins to venues, venue geographical locations, venue categories, and user textual comments on venues. These data contain rich knowledge on the user's social interactions in communities, geographical mobility patterns between regions, categorical preferences on activities, aspect interests in topics, and opinion expressions for sentiments. Such knowledge is essential for two key applications, namely, text sentiment classification and venue recommendations, which will be developed in this paper. To extract the knowledge from the data, the key task is to discover the latent communities, regions, activities, topics, and sentiments of users. However, these latent variables are interdependent, e.g., users in the same community usually travel on nearby regions and share common activities and topics, which renders a big challenge for modeling these latent variables. To tackle this challenge, in this study, we propose an LDA-based model called CRATS that jointly mines the latent Communities, Regions, Activities, Topics, and Sentiments based on the important dependencies among these latent variables. To the best of our knowledge, this is the first study to jointly model these five latent variables. Finally, we conduct a comprehensive performance evaluation for CRATS in different applications, including text sentiment classification and venue recommendations, using three large-scale real-world geosocial network data sets collected from Yelp and Foursquare. Experimental results show that CRATS achieves significantly superior performance against other state-of-the-art techniques. Jia-Dong Zhang, Chi-Yin Chow |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | A Location- and Diversity-Aware News Feed System for Mobile UsersabstractA location-aware news feed (LANF) system generates news feeds for a mobile user based on her spatial preference (i.e., her current location and future locations) and non-spatial preference (i.e., her interest). Existing LANF systems simply send the most relevant geo-tagged messages to their users. Unfortunately, the major limitation of such an existing approach is that, a news feed may contain messages related to the same location (i.e., point-of-interest) or the same category of locations (e.g., food, entertainment or sport). We argue that diversity is a very important feature for location-aware news feeds because it helps users discover new places and activities. In this paper, we propose D-MobiFeed; a new LANF system enables a user to specify the minimum number of message categories (h) for the messages in a news feed. In D-MobiFeed, our objective is to efficiently schedule news feeds for a mobile user at her current and predicted locations, such that (i) each news feed contains messages belonging to at least h different categories, and (ii) their total relevance to the user is maximized. To achieve this objective, we formulate the problem into two parts, namely, a decision problem and an optimization problem. For the decision problem, we provide an exact solution by modeling it as a maximum flow problem and proving its correctness. The optimization problem is solved by our proposed three-stage heuristic algorithm. We conduct a user study and experiments to evaluate the performance of D-MobiFeed using a real data set crawled from Foursquare. Experimental results show that our proposed three-stage heuristic scheduling algorithm outperforms the brute-force optimal algorithm by at least an order of magnitude in terms of running time and the relative error incurred by the heuristic algorithm is below 1 percent. D-MobiFeed with the location prediction method effectively improves the relevance, diversity, and efficiency of news feeds. Wenjian Xu, Chi-Yin Chow |
IEEE Trans. Serv. Comput. | 2 |
| 2016 | TICRec: A Probabilistic Framework to Utilize Temporal Influence Correlations for Time-Aware Location RecommendationsabstractIn location-based social networks (LBSNs), time significantly affects users' check-in behaviors, for example, people usually visit different places at different times of weekdays and weekends, e.g., restaurants at noon on weekdays and bars at midnight on weekends. Current studies use the temporal influence to recommend locations through dividing users' check-in locations into time slots based on their check-in time and learning their preferences to locations in each time slot separately. Unfortunately, these studies generally suffer from two major limitations: (1) the loss of time information because of dividing a day into time slots and (2) the lack of temporal influence correlations due to modeling users' preferences to locations for each time slot separately. In this paper, we propose a probabilistic framework called TICRec that utilizes temporal influence correlations (TIC) of both weekdays and weekends for time-aware location recommendations. TICRec not only recommends locations to users, but it also suggests when a user should visit a recommended location. In TICRec, we estimate a time probability density of a user visiting a new location without splitting the continuous time into discrete time slots to avoid the time information loss. To leverage the TIC, TICRec considers both user-based TIC (i.e., different users' check-in behaviors to the same location at different times) and location-based TIC (i.e., the same user's check-in behaviors to different locations at different times). Finally, we conduct a comprehensive performance evaluation for TICRec using two real data sets collected from Foursquare and Gowalla. Experimental results show that TICRec achieves significantly superior location recommendations compared to other state-of-the-art recommendation techniques with temporal influence. Jia-Dong Zhang, Chi-Yin Chow |
IEEE Trans. Serv. Comput. | 2 |
| 2015 | Sampling Big Trajectory DataabstractThe increasing prevalence of sensors and mobile devices has led to an explosive increase of the scale of spatio-temporal data in the form of trajectories. A trajectory aggregate query, as a fundamental functionality for measuring trajectory data, aims to retrieve the statistics of trajectories passing a user-specified spatio-temporal region. A large-scale spatio-temporal database with big disk-resident data takes very long time to produce exact answers to such queries. Hence, approximate query processing with a guaranteed error bound is a promising solution in many scenarios with stringent response-time requirements. In this paper, we study the problem of approximate query processing for trajectory aggregate queries. We show that it boils down to the distinct value estimation problem, which has been proven to be very hard with powerful negative results given that no index is built. By utilizing the well-established spatio-temporal index and introducing an inverted index to trajectory data, we are able to design random index sampling (RIS) algorithm to estimate the answers with a guaranteed error bound. To further improve system scalability, we extend RIS algorithm to concurrent random index sampling (CRIS) algorithm to process a number of trajectory aggregate queries arriving concurrently with overlapping spatio-temporal query regions. To demonstrate the efficacy and efficiency of our sampling and estimation methods, we applied them in a real large-scale user trajectory database collected from a cellular service provider in China. Our extensive evaluation results indicate that both RIS and CRIS outperform exhaustive search for single and concurrent trajectory aggregate queries by two orders of magnitude in terms of the query processing time, while preserving a relative error ratio lower than 10\%, with only 1% search cost of the exhaustive search method. Chi-Yin Chow, Mingxuan Yuan, Jia-Dong Zhang, Qiang Yang 0001, Zhi-Li Zhang |
CIKM | 2 |
| 2015 | ORec: An Opinion-Based Point-of-Interest Recommendation FrameworkabstractAs location-based social networks (LBSNs) rapidly grow, it is a timely topic to study how to recommend users with interesting locations, known as points-of-interest (POIs). Most existing POI recommendation techniques only employ the check-in data of users in LBSNs to learn their preferences on POIs by assuming a user's check-in frequency to a POI explicitly reflects the level of her preference on the POI. However, in reality users usually visit POIs only once, so the users' check-ins may not be sufficient to derive their preferences using their check-in frequencies only. Actually, the preferences of users are exactly implied in their opinions in text-based tips commenting on POIs. In this paper, we propose an opinion-based POI recommendation framework called ORec to take full advantage of the user opinions on POIs expressed as tips. In ORec, there are two main challenges: (i) detecting the polarities of tips (positive, neutral or negative), and (ii) integrating them with check-in data including social links between users and geographical information of POIs. To address these two challenges, (1) we develop a supervised aspect-dependent approach to detect the polarity of a tip, and (2) we devise a method to fuse tip polarities with social links and geographical information into a unified POI recommendation framework. Finally, we conduct a comprehensive performance evaluation for ORec using two large-scale real data sets collected from Foursquare and Yelp. Experimental results show that ORec achieves significantly superior polarity detection and POI recommendation accuracy compared to other state-of-the-art polarity detection and POI recommendation techniques. Jia-Dong Zhang, Chi-Yin Chow, Yu Zheng 0004 |
CIKM | 2 |
| 2015 | TaxiRec: recommending road clusters to taxi drivers using ranking-based extreme learning machinesabstractUtilizing large-scale GPS data to improve taxi services becomes a popular research problem in the areas of data mining, intelligent transportation, and the Internet of Things. In this paper, we utilize a large-scale GPS data set generated by over 7,000 taxis in a period of one month in Nanjing, China, and propose TaxiRec; a framework for discovering the passenger-finding potentials of road clusters, which is incorporated into a recommender system for taxi drivers to hunt passengers. In TaxiRec, we first construct the road network by defining the nodes and road segments. Then, the road network is divided into a number of road clusters through a clustering process on the mid points of the road segments. Afterwards, a set of features for each road cluster is extracted from real-life data sets, and a ranking-based extreme learning machine (ELM) model is proposed to evaluate the passenger-finding potential of each road cluster. Experimental results demonstrate the feasibility and effectiveness of the proposed framework. Ran Wang 0001, Chi-Yin Chow, Victor C. S. Lee, Sam Kwong |
SIGSPATIAL/GIS | 2 |
| 2015 | Coding-Based Cooperative Caching in Data Broadcast Environments
Houling Ji, Victor C. S. Lee, Chi-Yin Chow, Kai Liu 0001, Guoqing Wu 0004 |
ICA3PP (1) | 3 |
| 2015 | Growing the charging station network for electric vehicles with trajectory data analyticsabstractElectric vehicles (EVs) have undergone an explosive increase over recent years, due to the unparalleled advantages over gasoline cars in green transportation and cost efficiency. Such a drastic increase drives a growing need for widely deployed publicly accessible charging stations. Thus, how to strategically deploy the charging stations and charging points becomes an emerging and challenging question to urban planners and electric utility companies. In this paper, by analyzing a large scale electric taxi trajectory data, we make the first attempt to investigate this problem. We develop an optimal charging station deployment (OCSD) framework that takes the historical EV taxi trajectory data, road map data, and existing charging station information as input, and performs optimal charging station placement (OCSP) and optimal charging point assignment (OCPA). The OCSP and OCPA optimization components are designed to minimize the average time to the nearest charging station, and the average waiting time for an available charging point, respectively. To evaluate the performance of our OCSD framework, we conduct experiments on one-month real EV taxi trajectory data. The evaluation results demonstrate that our OCSD framework can achieve a 26%–94% reduction rate on average time to find a charging station, and up to two orders of magnitude reduction on waiting time before charging, over baseline methods. Moreover, our results reveal interesting insights in answering the question: “Super or small stations?”: When the number of deployable charging points is sufficiently large, more small stations are preferred; and when there are relatively few charging points to deploy, super stations is a wiser choice. Jun Luo 0007, Chi-Yin Chow, Kam-Lam Chan, Fan Zhang 0019 |
ICDE | 3 |
| 2015 | GeoSoCa: Exploiting Geographical, Social and Categorical Correlations for Point-of-Interest RecommendationsabstractRecommending users with their preferred points-of-interest (POIs), e.g., museums and restaurants, has become an important feature for location-based social networks (LBSNs), which benefits people to explore new places and businesses to discover potential customers. However, because users only check in a few POIs in an LBSN, the user-POI check-in interaction is highly sparse, which renders a big challenge for POI recommendations. To tackle this challenge, in this study we propose a new POI recommendation approach called GeoSoCa through exploiting geographical correlations, social correlations and categorical correlations among users and POIs. The geographical, social and categorical correlations can be learned from the historical check-in data of users on POIs and utilized to predict the relevance score of a user to an unvisited POI so as to make recommendations for users. First, in GeoSoCa we propose a kernel estimation method with an adaptive bandwidth to determine a personalized check-in distribution of POIs for each user that naturally models the geographical correlations between POIs. Then, GeoSoCa aggregates the check-in frequency or rating of a user's friends on a POI and models the social check-in frequency or rating as a power-law distribution to employ the social correlations between users. Further, GeoSoCa applies the bias of a user on a POI category to weigh the popularity of a POI in the corresponding category and models the weighed popularity as a power-law distribution to leverage the categorical correlations between POIs. Finally, we conduct a comprehensive performance evaluation for GeoSoCa using two large-scale real-world check-in data sets collected from Foursquare and Yelp. Experimental results show that GeoSoCa achieves significantly superior recommendation quality compared to other state-of-the-art POI recommendation techniques. Jia-Dong Zhang, Chi-Yin Chow |
SIGIR | 2 |
| 2015 | Learning ELM-Tree from big data based on uncertainty reduction
Ran Wang 0001, Yu-Lin He, Chi-Yin Chow, Fang-Fang Ou |
Fuzzy Sets Syst. | 3 |
| 2015 | MobiFeed: A location-aware news feed framework for moving users
Wenjian Xu, Chi-Yin Chow, Man Lung Yiu, Qing Li 0001, Chung Keung Poon |
GeoInformatica | 2 |
| 2015 | CoRe: Exploiting the personalized influence of two-dimensional geographic coordinates for location recommendations
Jia-Dong Zhang, Chi-Yin Chow |
Inf. Sci. | 2 |
| 2015 | REAL: A Reciprocal Protocol for Location Privacy in Wireless Sensor NetworksabstractK-anonymity has been used to protect location privacy for location monitoring services in wireless sensor networks (WSNs), where sensor nodes work together to report k-anonymized aggregate locations to a server. Each k-anonymized aggregate location is a cloaked area that contains at least k persons. However, we identify an attack model to show that overlapping aggregate locations still pose privacy risks because an adversary can infer some overlapping areas with less than k persons that violates the k-anonymity privacy requirement. In this paper, we propose a reciprocal protocol for location privacy (REAL) in WSNs. In REAL, sensor nodes are required to autonomously organize their sensing areas into a set of non-overlapping and highly accurate k-anonymized aggregate locations. To confront the three key challenges in REAL, namely, self-organization, reciprocity property and high accuracy, we design a state transition process, a locking mechanism and a time delay mechanism, respectively. We compare the performance of REAL with current protocols through simulated experiments. The results show that REAL protects location privacy, provides more accurate query answers, and reduces communication and computational costs. Jia-Dong Zhang, Chi-Yin Chow |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2015 | Spatiotemporal Sequential Influence Modeling for Location Recommendations: A Gravity-based ApproachabstractRecommending to users personalized locations is an important feature of Location-Based Social Networks (LBSNs), which benefits users who wish to explore new places and businesses to discover potential customers. In LBSNs, social and geographical influences have been intensively used in location recommendations. However, human movement also exhibits spatiotemporal sequential patterns, but only a few current studies consider the spatiotemporal sequential influence of locations on users’ check-in behaviors. In this article, we propose a new gravity model for location recommendations, called LORE, to exploit the spatiotemporal sequential influence on location recommendations. First, LORE extracts sequential patterns from historical check-in location sequences of all users as a Location-Location Transition Graph (L 2 TG), and utilizes the L 2 TG to predict the probability of a user visiting a new location through the developed additive Markov chain that considers the effect of all visited locations in the check-in history of the user on the new location. Furthermore, LORE applies our contrived gravity model to weigh the effect of each visited location on the new location derived from the personalized attractive force (i.e., the weight) between the visited location and the new location. The gravity model effectively integrates the spatiotemporal, social, and popularity influences by estimating a power-law distribution based on (i) the spatial distance and temporal difference between two consecutive check-in locations of the same user, (ii) the check-in frequency of social friends, and (iii) the popularity of locations from all users. Finally, we conduct a comprehensive performance evaluation for LORE using three large-scale real-world datasets collected from Foursquare, Gowalla, and Brightkite. Experimental results show that LORE achieves significantly superior location recommendations compared to other state-of-the-art location recommendation techniques. Jia-Dong Zhang, Chi-Yin Chow |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2015 | User-Defined Privacy Grid System for Continuous Location-Based ServicesabstractLocation-based services (LBS) require users to continuously report their location to a potentially untrusted server to obtain services based on their location, which can expose them to privacy risks. Unfortunately, existing privacy-preserving techniques for LBS have several limitations, such as requiring a fully-trusted third party, offering limited privacy guarantees and incurring high communication overhead. In this paper, we propose a user-defined privacy grid system called dynamic grid system (DGS); the first holistic system that fulfills four essential requirements for privacy-preserving snapshot and continuous LBS. (1) The system only requires a semi-trusted third party, responsible for carrying out simple matching operations correctly. This semi-trusted third party does not have any information about a user's location. (2) Secure snapshot and continuous location privacy is guaranteed under our defined adversary models. (3) The communication cost for the user does not depend on the user's desired privacy level, it only depends on the number of relevant points of interest in the vicinity of the user. (4) Although we only focus on range and k-nearest-neighbor queries in this work, our system can be easily extended to support other spatial queries without changing the algorithms run by the semi-trusted third party and the database server, provided the required search area of a spatial query can be abstracted into spatial regions. Experimental results show that our DGS is more efficient than the state-of-the-art privacy-preserving technique for continuous LBS. Roman Schlegel, Chi-Yin Chow, Qiong Huang 0001, Duncan S. Wong |
IEEE Trans. Mob. Comput. | 2 |
| 2015 | How Much to Coordinate? Optimizing In-Network Caching in Content-Centric NetworksabstractIn content-centric networks, it is challenging how to optimally provision in-network storage to cache contents, to balance the tradeoffs between the network performance and the provisioning cost. To address this problem, we first propose a holistic model for intradomain networks to characterize the network performance of routing contents to clients and the network cost incurred by globally coordinating the in-network storage capability. We then derive the optimal strategy for provisioning the storage capability that optimizes the overall network performance and cost, and analyze the performance gains via numerical evaluations on real network topologies. Our results reveal interesting phenomena; for instance, different ranges of the Zipf exponent can lead to opposite optimal strategies, and the tradeoffs between the network performance and the provisioning cost have great impacts on the stability of the optimal strategy. We also demonstrate that the optimal strategy can achieve significant gain on both the load reduction at origin servers and the improvement on the routing performance. Moreover, given an optimal coordination level ℓ*, we design a routing-aware content placement (RACP) algorithm that runs on a centralized server. The algorithm computes and assigns contents to each CCN router to store, which can minimize the overall routing cost, e.g., transmission delay or hop counts, to deliver contents to clients. By conducting extensive simulations using a large-scale trace dataset collected from a commercial 3G network in China, our results demonstrate that our caching scheme can achieve 4% to 22% latency reduction on average over the state-of-the-art caching mechanisms. Haiyong Xie 0001, Yonggang Wen 0001, Chi-Yin Chow, Zhi-Li Zhang |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2015 | iGeoRec: A Personalized and Efficient Geographical Location Recommendation FrameworkabstractGeographical influence has been intensively exploited for location recommendations in location-based social networks (LBSNs) due to the fact that geographical proximity significantly affects users’ check-in behaviors. However, current studies only model the geographical influence on all users’ check-in behaviors as auniversalway. We argue that the geographical influence on users’ check-in behaviors should bepersonalized. In this paper, we propose a personalized and efficient geographical location recommendation framework called iGeoRec to take full advantage of the geographical influence on location recommendations. In iGeoRec, there are mainly two challenges: (1) personalizing the geographical influence to accurately predict the probability of a user visiting a new location, and (2) efficiently computing the probability of each user to all new locations. To address these two challenges, (1) we propose a probabilistic approach to personalize the geographical influence as a personal distribution for each user and predict the probability of a user visiting any new location using her personal distribution. Furthermore, (2) we develop an efficient approximation method to compute the probability of any user to all new locations; the proposed method reduces the computational complexity of the exact computation method from$O(|L|n^3)$to$O(|L|n)$(where$|L|$is the total number of locations in an LBSN and$n$is the number of check-in locations of a user). Finally, we conduct extensive experiments to evaluate the recommendationaccuracyandefficiencyof iGeoRec using two large-scale real data sets collected from the two of the most popular LBSNs: Foursquare and Gowalla. Experimental results show that iGeoRec provides significantly superior performance compared to other state-of-the-art geographical recommendation techniques. Jia-Dong Zhang, Chi-Yin Chow |
IEEE Trans. Serv. Comput. | 2 |
| 2014 | Using multi-criteria decision making for personalized point-of-interest recommendationsabstractLocation-based business review (LBBR) sites (e.g., Yelp) provide us a possibility to recommend new points of interest (POIs) for users. The geographical position and category of POIs have been considered as two major factors in modeling users' preferences. However, it is argued that the user's visiting behaviors are also affected by the attributes of POIs, which reflect the basic features of the POIs. Besides, a user may have different preference levels on the same POI with regard to different criteria. To this end, we propose a new personalized POI recommendation framework using Multi-Criteria Decision Making (MCDM). Firstly, preference models are built for the user's geographical, category, and attribute preferences. Then, an MCDM-based recommendation framework is designed to iteratively combine the user's preferences on the three criteria and select the top-N POIs as a recommendation list. Experimental results show that our framework not only outperforms the state-of-the-art POI recommendation techniques, but also provides a better trade-off mechanism for MCDM than the weighted sum approach. Chi-Yin Chow, Ran Wang 0001, Victor C. S. Lee |
SIGSPATIAL/GIS | 2 |
| 2014 | Exploring cell tower data dumps for supervised learning-based point-of-interest predictionabstractExploring massive mobile data for location-based services (LBS) becomes one of the key challenges in mobile data mining. In this paper, we propose a framework that uses large-scale cell tower data dumps and extracts points-of-interest (POIs) from a social network web site called Weibo, and provides new LBS based on these two data sets, i.e., predicting the existence of POIs and the number of POIs in a certain area. We use Voronoi diagram to divide a city area into non-overlapping regions, and a k-means clustering algorithm to aggregate neighboring cell towers into region groups. A supervised learning algorithm is adopted to build up a model between the number of connections of cell towers and the POIs in different region groups, where a classification or regression model is used to predict the POI existence or the number of POIs, respectively. We studied 12 state-of-the-art classification and regression algorithms, and the experimental results demonstrate the feasibility and effectiveness of the proposed framework. Ran Wang 0001, Chi-Yin Chow, Sarana Nutanong, Mingxuan Yuan, Victor C. S. Lee |
SIGSPATIAL/GIS | 2 |
| 2014 | LORE: exploiting sequential influence for location recommendationsabstractProviding location recommendations becomes an important feature for location-based social networks (LBSNs), since it helps users explore new places and makes LBSNs more prevalent to users. In LBSNs, geographical influence and social influence have been intensively used in location recommendations based on the facts that geographical proximity of locations significantly affects users' check-in behaviors and social friends often have common interests. Although human movement exhibits sequential patterns, most current studies on location recommendations do not consider any sequential influence of locations on users' check-in behaviors. In this paper, we propose a new approach called LORE to exploit sequential influence on location recommendations. First, LORE incrementally mines sequential patterns from location sequences and represents the sequential patterns as a dynamic Location-Location Transition Graph (L2TG). LORE then predicts the probability of a user visiting a location by Additive Markov Chain (AMC) with L2TG. Finally, LORE fuses sequential influence with geographical influence and social influence into a unified recommendation framework; in particular the geographical influence is modeled as two-dimensional check-in probability distributions rather than one-dimensional distance probability distributions in existing works. We conduct a comprehensive performance evaluation for LORE using two large-scale real data sets collected from Foursquare and Gowalla. Experimental results show that LORE achieves significantly superior location recommendations compared to other state-of-the-art recommendation techniques. Jia-Dong Zhang, Chi-Yin Chow |
SIGSPATIAL/GIS | 2 |
| 2014 | Differentially Private Location Recommendations in Geosocial NetworksabstractLocation-tagged social media have an increasingly important role in shaping behavior of individuals. With the help of location recommendations, users are able to learn about events, products or places of interest that are relevant to their preferences. User locations and movement patterns are available from geosocial networks such as Foursquare, mass transit logs or traffic monitoring systems. However, disclosing movement data raises serious privacy concerns, as the history of visited locations can reveal sensitive details about an individual's health status, alternative lifestyle, etc. In this paper, we investigate mechanisms to sanitize location data used in recommendations with the help of differential privacy. We also identify the main factors that must be taken into account to improve accuracy. Extensive experimental results on real-world datasets show that a careful choice of differential privacy technique leads to satisfactory location recommendation results. Jia-Dong Zhang, Gabriel Ghinita, Chi-Yin Chow |
MDM (1) | 3 |
| 2013 | CALBA: capacity-aware location-based advertising in temporary social networksabstractA temporary social network (TSN) is confined to a specific place (e.g., hotel and shopping mall) or activity (e.g., concert and exhibition) in which the TSN service provider allows nearby third party vendors (e.g., restaurants and stores) to advertise their goods or services to its registered users. However, simply broadcasting all the vendors' advertisements to all the users in the TSN may cause the service provider to lose its fans. In this paper, we present Capacity-Aware Location-Based Advertising (CALBA), which is a framework designed for TSNs to select vendors as advertising sources for mobile users. In CALBA we measure the relevance of a vendor to a user by considering their geographical proximity and the user's preferences. Our goal is to maximize the overall relevance of selected vendors for a user with the constraint that the total advertising frequency of the selected vendors should not exceed the user's specified capacity. First, we model the snapshot selection problem as 0-1 knapsack and solve it using an approximation method. Then, CALBA keeps track of the selection result for moving users by employing a safe region technique that can reduce its computational cost. We also propose three pruning rules and a unique access order to effectively prune vendors which could not affect a safe region, in order to improve the efficiency of the client-side computation. We evaluate the performance of CALBA based on a real location-based social network data set crawled from Foursquare. Experimental results show that CALBA outperforms a naïve approach which periodically invokes the snapshot vendor selection. Wenjian Xu, Chi-Yin Chow, Jia-Dong Zhang |
SIGSPATIAL/GIS | 2 |
| 2013 | iGSLR: personalized geo-social location recommendation: a kernel density estimation approachabstractWith the rapidly growing location-based social networks (LBSNs), personalized geo-social recommendation becomes an important feature for LBSNs. Personalized geo-social recommendation not only helps users explore new places but also makes LBSNs more prevalent to users. In LBSNs, aside from user preference and social influence, geographical influence has also been intensively exploited in the process of location recommendation based on the fact that geographical proximity significantly affects users' check-in behaviors. Although geographical influence on users should be personalized, current studies only model the geographical influence on all users' check-in behaviors in a universal way. In this paper, we propose a new framework called iGSLR to exploit personalized social and geographical influence on location recommendation. iGSLR uses a kernel density estimation approach to personalize the geographical influence on users' check-in behaviors as individual distributions rather than a universal distribution for all users. Furthermore, user preference, social influence, and personalized geographical influence are integrated into a unified geo-social recommendation framework. We conduct a comprehensive performance evaluation for iGSLR using two large-scale real data sets collected from Foursquare and Gowalla which are two of the most popular LBSNs. Experimental results show that iGSLR provides significantly superior location recommendation compared to other state-of-the-art geo-social recommendation techniques. Jia-Dong Zhang, Chi-Yin Chow |
SIGSPATIAL/GIS | 2 |
| 2013 | SMashQ: spatial mashup framework for k-NN queries in time-dependent road networks
Detian Zhang, Chi-Yin Chow, Qing Li 0001, Xinming Zhang 0001, Yinlong Xu 0001 |
Distributed Parallel Databases | 2 |
| 2013 | Efficient Index-Based Approaches for Skyline Queries in Location-Based ApplicationsabstractEnriching many location-based applications, various new skyline queries are proposed and formulated based on the notion of locational dominance, which extends conventional one by taking objects' nearness to query positions into account additional to objects' nonspatial attributes. To answer a representative class of skyline queries for location-based applications efficiently, this paper presents two index-based approaches, namely, augmented R-tree and dominance diagram. Augmented R-tree extends R-tree by including aggregated nonspatial attributes in index nodes to enable dominance checks during index traversal. Dominance diagram is a solution-based approach, by which each object is associated with a precomputed nondominance scope wherein query points should have the corresponding object not locationally dominated by any other. Dominance diagram enables skyline queries to be evaluated via parallel and independent comparisons between nondominance scopes and query points, providing very high search efficiency. The performance of these two approaches is evaluated via empirical studies, in comparison with other possible approaches. Ken C. K. Lee, Baihua Zheng, Cindy X. Chen, Chi-Yin Chow |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2012 | MobiFeed: a location-aware news feed system for mobile usersabstractA location-aware news feed system enables mobile users to share geo-tagged user-generated messages, e.g., a user can receive nearby messages that are the most relevant to her. In this paper, we present MobiFeed that is a framework designed for scheduling news feeds for mobile users. MobiFeed consists of three key functions, location prediction, relevance measure, and news feed scheduler. The location prediction function is designed to predict a mobile user's locations based on an existing path prediction algorithm. The relevance measure function is implemented by combining the vector space model with non-spatial and spatial factors to determine the relevance of a message to a user. The news feed scheduler works with the other two functions to generate news feeds for a mobile user at her current and predicted locations with the best overall quality. To ensure that MobiFeed can scale up to a larger number of messages, we design a heuristic news feed scheduler. Wenjian Xu, Chi-Yin Chow, Man Lung Yiu, Qing Li 0001, Chung Keung Poon |
SIGSPATIAL/GIS | 2 |
| 2012 | Dash: A Novel Search Engine for Database-Generated Dynamic Web PagesabstractDatabase-generated dynamic web pages (db-pages, in short), whose contents are created on the fly by web applications and databases, are now prominent in the web. However, many of them cannot be searched by existing search engines. Accordingly, we develop a novel search engine named Dash, which stands for Db-pAge Search, to support db-page search. Dash determines db-pages possibly generated by a target web application and its database through exploring the application code and the related database content and supports keyword search on those db-pages. In this paper, we present its system design and focus on the efficiency issue. To minimize costs incurred for collecting, maintaining, indexing and searching a massive number of db-pages that possibly have overlapped contents, Dash derives and indexes db-page fragments in place of db-pages. Each db-page fragment carries a disjointed part of a db-page. To efficiently compute and index db-page fragments from huge datasets, Dash is equipped with MapReduce based algorithms for database crawling and db-page fragment indexing. Besides, Dash has a top-k search algorithm that can efficiently assemble db-page fragments into db-pages relevant to search keywords and return the k most relevant ones. The performance of Dash is evaluated via extensive experimentation. Ken C. K. Lee, Kanchan Bankar, Baihua Zheng, Chi-Yin Chow, Honggang Wang 0001 |
ICDCS | 4 |
| 2012 | GeoFeed: A Location Aware News Feed SystemabstractThis paper presents the Geo Feed system, a location-aware news feed system that provides a new platform for its users to get spatially related message updates from either their friends or favorite news sources. Geo Feed distinguishes itself from all existing news feed systems in that it takes into account the spatial extents of messages and user locations when deciding upon the selected news feed. Geo Feed is equipped with three different approaches for delivering the news feed to its users, namely, spatial pull, spatial push, and shared push. Then, the main challenge of Geo Feed is to decide on when to use each of these three approaches to which users. Geo Feed is equipped with a smart decision model that decides about using these approaches in a way that: (a) minimizes the system overhead for delivering the location-aware news feed, and (b) guarantees a certain response time for each user to obtain the requested location-aware news feed. Experimental results, based on real and synthetic data, show that Geo Feed outperforms existing news feed systems in terms of response time and maintenance cost. Jie Bao 0003, Mohamed F. Mokbel, Chi-Yin Chow |
ICDE | 3 |
| 2011 | Efficient Evaluation of k-NN Queries Using Spatial Mashups
Detian Zhang, Chi-Yin Chow, Qing Li 0001, Xinming Zhang 0001, Yinlong Xu 0001 |
SSTD | 2 |
| 2011 | Query-aware location anonymization for road networks
Chi-Yin Chow, Mohamed F. Mokbel, Jie Bao 0003 |
GeoInformatica | 1 |
| 2011 | Spatial cloaking for anonymous location-based services in mobile peer-to-peer environments
Chi-Yin Chow, Mohamed F. Mokbel |
GeoInformatica | 1 |
| 2011 | A Privacy-Preserving Location Monitoring System for Wireless Sensor NetworksabstractMonitoring personal locations with a potentially untrusted server poses privacy threats to the monitored individuals. To this end, we propose a privacy-preserving location monitoring system for wireless sensor networks. In our system, we design two in-network location anonymization algorithms, namely, resource and quality-aware algorithms, that aim to enable the system to provide high-quality location monitoring services for system users, while preserving personal location privacy. Both algorithms rely on the well-established k-anonymity privacy concept, that is, a person is indistinguishable among k persons, to enable trusted sensor nodes to provide the aggregate location information of monitored persons for our system. Each aggregate location is in a form of a monitored area A along with the number of monitored persons residing in A, where A contains at least k persons. The resource-aware algorithm aims to minimize communication and computational cost, while the quality-aware algorithm aims to maximize the accuracy of the aggregate locations by minimizing their monitored areas. To utilize the aggregate location information to provide location monitoring services, we use a spatial histogram approach that estimates the distribution of the monitored persons based on the gathered aggregate location information. Then, the estimated distribution is used to provide location monitoring services through answering range queries. We evaluate our system through simulated experiments. The results show that our system provides high-quality location monitoring services for system users and guarantees the location privacy of the monitored persons. Chi-Yin Chow, Mohamed F. Mokbel, Tian He 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2011 | On Efficient and Scalable Support of Continuous Queries in Mobile Peer-to-Peer EnvironmentsabstractIn this paper, we propose an efficient and scalable query processing framework for continuous spatial queries (range and k-nearest-neighbor queries) in mobile peer-to-peer (P2P) environments, where no fixed communication infrastructure or centralized/distributed servers are available. Due to the limitations in mobile P2P environments, for example, user mobility, limited battery power, limited communication range, and scarce communication bandwidth, it is costly to maintain the exact answer of continuous spatial queries. To this end, our framework enables the user to find an approximate answer with quality guarantees. In particular, we design two key features to adapt continuous spatial query processing to mobile P2P environments. 1) Each mobile user can specify his or her desired quality of services (QoS) for a query answer in a personalized QoS profile. The QoS profile consists of two parameters, namely, coverage and accuracy. The coverage parameter indicates the desired level of completeness of the available information for computing an approximate answer, and the accuracy parameter indicates the desired level of accuracy of the approximate answer. 2) We design a continuous answer maintenance scheme to enable the user to collaborate with other peers to continuously maintain a query answer. With these two features in our framework, the user can obtain a query answer from a local cache if the answer satisfies his or her QoS requirements. Otherwise, the user enlists neighbors for help to share their cached information to refine the answer. If the refined answer still cannot satisfy the QoS requirements, the user broadcasts the query to the peers residing within the required search area of the query to find the most accurate answer. Experiment results show that our framework is efficient and scalable and provides an effective trade-off between the communication overhead and the quality of query answers. Chi-Yin Chow, Mohamed F. Mokbel, Hong Va Leong |
IEEE Trans. Mob. Comput. | 1 |
| 2010 | Efficient Evaluation of k-Range Nearest Neighbor Queries in Road NetworksabstractA k-Range Nearest Neighbor (or kRNN for short) query in road networks finds the k nearest neighbors of every point on the road segments within a given query region based on the network distance. The kRNN query is significantly important for location-based applications in many realistic scenarios. For example, (1) the user’s location is uncertain, i.e., user’s location is modeled by a spatial region, and (2) the user is not willing to reveal her exact location to preserve her privacy, i.e., her location is blurred into a spatial region. However, the existing solutions for kRNN queries simply apply the traditional k-nearest neighbor query processing algorithm multiple times, which poses a huge redundant searching overhead. To this end, we propose an efficient kRNN query processing algorithm in this paper. Our algorithm (1) employs a shared execution approach to eliminate the redundant searching overhead, and (2) provides a parameter that can be tuned to achieve a tradeoff between the query processing performance and the storage overhead, while guaranteeing the user’s exact k-nearest neighbors are included in the query answers. The experimental results show that our algorithm always outperforms the existing solution in terms of query response time, and the introduced tuning parameter is an effective way to achieve the tradeoff between the query response time and the storage overhead. Jie Bao 0003, Chi-Yin Chow, Mohamed F. Mokbel, Wei-Shinn Ku |
Mobile Data Management | 2 |
| 2009 | Aggregate Location Monitoring for Wireless Sensor Networks: A Histogram-Based ApproachabstractLocation monitoring systems are used to detect human activities and provide monitoring services, e.g., aggregate queries. In this paper, we consider an aggregate location monitoring system where wireless sensor nodes are counting sensors that are only capable of detecting the number of objects within their sensing areas. As traditional query processors rely on the knowledge of users' exact locations, they cannot provide any monitoring services based on the readings reported from counting sensors. To this end, we propose an adaptive spatio-temporal histogram to enable monitoring services without the need of users' exact locations. The main idea of the histogram is to keep statistics about the distribution of moving objects. At the core of the histogram, we propose three techniques, memorization, locality awareness and packing, to improve monitoring accuracy and efficiency. Furthermore, the histogram is designed in a way that achieves a trade-off between the energy and bandwidth consumption of the sensor network and the accuracy of monitoring services. Experimental results show that the proposed histogram provides high-quality location monitoring services (i.e., 90% accuracy for both skewed and uniform mobility patterns) and outperforms a basic histogram and the state-of-the-art spatio-temporal histogram by two orders of magnitude in most cases. Chi-Yin Chow, Mohamed F. Mokbel, Tian He 0001 |
Mobile Data Management | 1 |
| 2009 | Approximate Evaluation of Range Nearest Neighbor Queries with Quality Guarantee
Chi-Yin Chow, Mohamed F. Mokbel, Joseph L. Naps, Suman Nath |
SSTD | 1 |
| 2009 | Casper*: Query processing for location services without compromising privacyabstractIn this article, we present a new privacy-aware query processing framework, Capser *, in which mobile and stationary users can obtain snapshot and/or continuous location-based services without revealing their private location information. In particular, we propose a privacy-aware query processor embedded inside a location-based database server to deal with snapshot and continuous queries based on the knowledge of the user's cloaked location rather than the exact location. Our proposed privacy-aware query processor is completely independent of how we compute the user's cloaked location. In other words, any existing location anonymization algorithms that blur the user's private location into cloaked rectilinear areas can be employed to protect the user's location privacy. We first propose a privacy-aware query processor that not only supports three new privacy-aware query types, but also achieves a trade-off between query processing cost and answer optimality. Then, to improve system scalability of processing continuous privacy-aware queries, we propose a shared execution paradigm that shares query processing among a large number of continuous queries. The proposed scalable paradigm can be tuned through two parameters to trade off between system scalability and answer optimality. Experimental results show that our query processor achieves high quality snapshot and continuous location-based services while supporting queries and/or data with cloaked locations. Chi-Yin Chow, Mohamed F. Mokbel, Walid G. Aref |
ACM Trans. Database Syst. | 1 |
| 2009 | Scalable processing of snapshot and continuous nearest-neighbor queries over one-dimensional uncertain data
Jinchuan Chen, Reynold Cheng, Mohamed F. Mokbel, Chi-Yin Chow |
VLDB J. | 4 |
| 2008 | Probabilistic Verifiers: Evaluating Constrained Nearest-Neighbor Queries over Uncertain DataabstractIn applications like location-based services, sensor monitoring and biological databases, the values of the database items are inherently uncertain in nature. An important query for uncertain objects is the probabilistic nearest-neighbor query (PNN), which computes the probability of each object for being the nearest neighbor of a query point. Evaluating this query is computationally expensive, since it needs to consider the relationship among uncertain objects, and requires the use of numerical integration or Monte-Carlo methods. Sometimes, a query user may not be concerned about the exact probability values. For example, he may only need answers that have sufficiently high confidence. We thus propose the constrained nearest-neighbor query (C-PNN), which returns the IDs of objects whose probabilities are higher than some threshold, with a given error bound in the answers. The C-PNN can be answered efficiently with probabilistic verifiers. These are methods that derive the lower and upper bounds of answer probabilities, so that an object can be quickly decided on whether it should be included in the answer. We have developed three probabilistic verifiers, which can be used on uncertain data with arbitrary probability density functions. Extensive experiments were performed to examine the effectiveness of these approaches. Reynold Cheng, Jinchuan Chen, Mohamed F. Mokbel, Chi-Yin Chow |
ICDE | 4 |
| 2008 | Tinycasper: a privacy-preserving aggregate location monitoring system in wireless sensor networksabstractThis demo presents a privacy-preserving aggregate location monitoring system, namely, TinyCasper, in which we can monitor moving objects in wireless sensor networks while preserving their location privacy. TinyCasper consists of two main modules, in-network location anonymization and aggregate query processing over anonymized locations. In the first module, trusted wireless sensor nodes collaborate with each other to anonymize users' exact locations by a cloaked spatial region that satisfies a prespecified privacy requirement. On the other side, the aggregate query processing module collects and analyzes the cloaked spatial regions reported from the wireless sensor nodes to support aggregate and alarm queries over anonymized locations. The prototype of TinyCasper is implemented on a physical test-bed on the TinyOS/Mote platform with 39 MICAz motes. Chi-Yin Chow, Mohamed F. Mokbel, Tian He 0001 |
SIGMOD Conference | 1 |
| 2007 | The New Casper: A Privacy-Aware Location-Based Database ServerabstractThis demo presents Casper; a framework in which users entertain anonymous location-based services. Casper consists of two main components; the location anonymizer that blurs the users' exact location into cloaked spatial regions and the privacy-aware query processor that is responsible on providing location-based services based on the cloaked spatial regions. While the location anonymizer is implemented as a stand alone application, the privacy-aware query processor is embedded into PLACE; a research prototype for location-based database servers. Mohamed F. Mokbel, Chi-Yin Chow, Walid G. Aref |
ICDE | 2 |
| 2007 | Enabling Private Continuous Queries for Revealed User Locations
Chi-Yin Chow, Mohamed F. Mokbel |
SSTD | 1 |
| 2007 | GroCoca: group-based peer-to-peer cooperative caching in mobile environmentabstractIn a mobile cooperative caching environment, we observe the need for cooperating peers to cache useful data items together, so as to improve cache hit from peers. This could be achieved by capturing the data requirement of individual peers in conjunction with their mobility pattern, for which we realized via a GROup-based COoperative CAching scheme (GroCoca). In GroCoca, we define a tightly-coupled group (TCG) as a collection of peers that possess similar mobility pattern and display similar data affinity. A family of algorithms is proposed to discover and maintain all TCGs dynamically. Furthermore, two cooperative cache management protocols, namely, cooperative cache admission control and replacement, are designed to control data replicas and improve data accessibility in TCGs. A cache signature scheme is also adopted in GroCoca in order to provide information for the mobile clients to determine whether their TCG members are likely caching their desired data items and to perform cooperative cache replacement Experimental results show that GroCoca outperforms the conventional caching scheme and standard COoperative CAching scheme (COCA) in terms of access latency and global cache hit ratio. However, GroCoca generally incurs higher power consumption. Chi-Yin Chow, Hong Va Leong, Alvin Chan Toong Shoon |
IEEE J. Sel. Areas Commun. | 1 |
| 2006 | A peer-to-peer spatial cloaking algorithm for anonymous location-based serviceabstractThis paper tackles a major privacy threat in current location-based services where users have to report their exact locations to the database server in order to obtain their desired services. For example, a mobile user asking about her nearest restaurant has to report her exact location. With untrusted service providers, reporting private location information may lead to several privacy threats. In this paper, we present a peer-to-peer (P2P)spatial cloaking algorithm in which mobile and stationary users can entertain location-based services without revealing their exact location information. The main idea is that before requesting any location-based service, the mobile user will form a group from her peers via single-hop communication and/or multi-hop routing. Then,the spatial cloaked area is computed as the region that covers the entire group of peers. Two modes of operations are supported within the proposed P2P s patial cloaking algorithm, namely, the on-demand mode and the proactive mode. Experimental results show that the P2P spatial cloaking algorithm operated in the on-demand mode has lower communication cost and better quality of services than the proactive mode, but the on-demand incurs longer response time. Chi-Yin Chow, Mohamed F. Mokbel |
GIS | 1 |
| 2006 | The New Casper: Query Processing for Location Services without Compromising Privacy
Mohamed F. Mokbel, Chi-Yin Chow, Walid G. Aref |
VLDB | 2 |
| 2005 | Distributed group-based cooperative caching in a mobile broadcast environmentabstractCaching is a key technique for improving data retrieval performance of mobile clients. The emergence of state-of-the-art peer-to-peer communication technologies now brings to reality what we call "cooperative caching" in which mobile clients not only can retrieve data items from mobile support stations, but also from the cache in their peers, thereby inducing a new dimension for mobile data caching. In this paper, we propose a distributed group-based cooperative caching scheme, in which we define the concept of a tightly-coupled group (TCG) by capturing the data affinity of individual peers and their mobility patterns, in a mobile broadcast environment. A distributed stable peer discovery protocol is proposed for discovering all TCGs dynamically. In addition, a cache signature scheme is adopted to provide hints for the mobile clients to determine whether their required data items are cached by their neighboring peers, and to perform cooperative cache replacement to increase overall data availability. Simulation studies are conducted to evaluate the effectiveness of our distributed group-based cooperative caching scheme. Chi-Yin Chow, Hong Va Leong, Alvin Chan Toong Shoon |
Mobile Data Management | 1 |
| 2004 | Cache Signatures for Peer-to-Peer Cooperative Caching in Mobile EnvironmentsabstractCaching is a key technique for improving data retrieval performance of mobile clients in mobile environments. The emergence of robust and reliable peer-to-peer (P2P) technologies now brings to reality what we call "cooperative caching" in which mobile clients can access data items from the cache in their neighboring peers. This paper considers a COoperative CAching scheme for mobile systems, called COCA. A cache signature scheme is devised for COCA that provides hints for the mobile clients to determine whether a required data item is cached by their neighboring peers based on their local state. The trade-off between the improvement in system performance and the overheads of the cache signature scheme in COCA is discussed. The performance of COCA with and without the cache signature scheme is evaluated through a number of simulated experiments. COCA is shown to be capable of effectively reducing the number of server requests and power consumption, as well as shortening the access latency as the number of neighboring peers increases. The inclusion of cache signature scheme further improves on the access latency. Chi-Yin Chow, Hong Va Leong, Alvin Chan Toong Shoon |
AINA (1) | 1 |
| 2004 | Group-Based Cooperative Cache Management for Mobile Clients in a Mobile EnvironmentabstractCaching is a key technique for improving data retrieval performance of mobile clients. The emergence of robust and reliable peer-to-peer (P2P) communication technologies now brings to reality what we call "cooperating caching" in which mobile clients not only can retrieve data items from mobile support stations, but also can access them from the cache in their neighboring peers, thereby inducing a new dimension for mobile data caching. This work extends a cooperative caching scheme, called COCA, in a pull-based mobile environment. Built upon the COCA framework, we propose a group-based cooperative caching scheme, called GroCoca, in which we define a tightly-coupled group (TCG) as a set of peers that possess similar movement pattern and exhibit similar data affinity. In GroCoca, a centralized incremental clustering algorithm is used to discover all TCGs dynamically, and the MHs in same TCG manage their cached data items cooperatively. In the simulated experiments, GroCoca is shown to reduce the access latency and server request ratio effectively. Chi-Yin Chow, Hong Va Leong, Alvin Chan Toong Shoon |
ICPP | 1 |