EDBT 2026 Demo / reviewers in the wild / expert
Kwan Hui Lim 0001
dblp:115/6484-1
· DBLP profile ↗
32ranked-venue papers in the field
8as first author
15since 2021 · last 2025
0000-0002-4569-0901ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 12 (3 first)Big Data, Cloud & Distributed Data Systems · 11 (2 first)Information Retrieval & Web Search · 8 (2 first)Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BGM-HAN: A Hierarchical Attention Network for Accurate and Fair Decision Assessment on Semi-structured Profiles
Junhua Liu 0002, Roy Ka-Wei Lee, Kwan Hui Lim 0001 |
ASONAM (2) | 3 |
| 2025 | Understanding Fairness-Accuracy Trade-offs in Machine Learning Models: Does Promoting Fairness Undermine Performance?
Junhua Liu 0002, Roy Ka-Wei Lee, Kwan Hui Lim 0001 |
ASONAM (2) | 3 |
| 2025 | From Intents to Conversations: Generating Intent-Driven Dialogues with Contrastive Learning for Multi-Turn Classification
Junhua Liu 0002, Tan Yong Keat, Kwan Hui Lim 0001 |
CIKM | 4 |
| 2025 | Get Global Guarantees: On the Probabilistic Nature of Perturbation RobustnessabstractRobustness is a critical requirement for deploying machine learning models in safety-sensitive domains, where even imperceptible input perturbations can lead to hazardous outcomes. However, existing robustness assessment techniques prior to deployment often face a trade-off between computational feasibility and measurement precision, limiting their effectiveness in practice. To address these limitations, we provide a systematic comparative study of prevailing robustness definitions and their corresponding evaluation methodologies. Building on this analysis, we propose tower robustness, which is a novel and practical concept setting out from a global perspective. Further, we provide upper and lower bounds of tower robustness, based on hypothesis testing, for quantitative evaluation, enabling more rigorous and efficient pre-deployment assessments. Through empirical investigation, we demonstrate that our approach provides reliable robustness assessments. These findings advance the systematic understanding of robustness and contribute a practical framework for enhancing the safety of machine learning models in safety-critical applications. Wenchuan Mu, Kwan Hui Lim 0001 |
CIKM | 2 |
| 2025 | Bayesian Privacy Guarantee for User History in Sequential Recommendation Using Randomised ResponseabstractSequential recommendation systems play an important role in delivering personalised user experiences, yet they rely heavily on detailed user history, raising serious privacy concerns. In this work, we introduce a novel framework that integrates a randomised response mechanism into sequential recommendation to provide strong privacy guarantees while preserving recommendation effectiveness. By obfuscating user history through controlled probabilistic item substitution based on semantic similarity, our approach ensures that released sequences protect individual behaviour with provable Bayesian posterior privacy. We further propose training strategies tailored for privacy-filtered data, including a frequency-based vocabulary expansion method inspired by subword tokenisation. Experiments on four real-world datasets demonstrate that our approach preserves recommendation quality under strong privacy constraints and outperforms existing baselines even without applying privacy filters. Wenchuan Mu, Kwan Hui Lim 0001 |
CIKM | 2 |
| 2025 | Deep Learning of Dynamic POI Generation and Optimisation for Itinerary RecommendationabstractItinerary recommendation involves suggesting a sequence of Points of Interests (POIs) that users obtain maximum satisfaction under a time budget. Existing models have three challenges. First, they model user interest as non-time dependent, which cannot capture user interest appropriately because user interest can be contextual on time, e.g., interest in restaurants are likely higher during typical meal times. Second, they model the distance dependency of user interest as a linear one, which does not always adequately capture this relationship, e.g., it could be a cubic decay relationship. Finally, existing studies treat POI recommendation and itinerary optimisation as two separate problems, which can result in sub-optimal itinerary recommendations. In this paper, we propose a deep learning model that recommends POIs and constructs the itinerary simultaneously and in an integrated manner. It captures user dynamic interest and non-linear spatial dependencies in itinerary recommendations. The proposed model has two steps, where the candidate selection policy generates a set of personalised candidate POIs based on user interest and the itinerary construction step maximises user interest within budget time. To recommend an appropriate candidate set, we propose a multi-head, attention-based transformer to leverage periodic trends and recent activities to capture user dynamic preferences. We also introduce a new co-visiting patterns-based graph convolutional network (GCN) model to capture user non-linear spatial dependencies. To construct the full itinerary from the dynamic candidate sets, we apply greedy policy that incrementally constructs itineraries within the budget time which aims to maximise user interest and minimize queuing time. Experimental results show that the proposed deep learning model outperforms state-of-the-art baselines in itinerary recommendation in four theme parks and four cities datasets. The proposed model outperforms the baselines in itinerary recommendation from 7.79% to 26.28% on various datasets in terms of F1-score value. We also show that the proposed candidate generation approach outperforms the state-of-the-art next POI recommendation models in eight real datasets. The proposed model outperforms the baselines on average by 11.29 % in terms of F1-score@5 values and 9.08% in terms of F1-score@10 values. We have publicly shared our source code at GitHub 1 for the reproducibility of our proposed model. Sajal Halder, Kwan Hui Lim 0001, Jeffrey Chan, Xiuzhen Zhang 0001 |
Trans. Recomm. Syst. | 2 |
| 2023 | Modelling Text Similarity: A SurveyabstractOnline social networking services such as Twitter and Instagram have become pervasive platforms for engaging in discussions on a wide array of topics. These platforms cater to both mainstream subjects, like music and movies, as well as more specialized areas, such as politics. With the growing volume of textual data generated on these platforms, the ability to define and identify similar texts becomes crucial for effective investigation and clustering. In this paper, we explore the challenges and significance of text similarity regression models in the context of online social networking services. We delve into the methods and techniques employed to define and find similarities among texts, enabling the extraction of meaningful patterns and insights. Specifically, we categorize text similarity regression models into four distinct types: set-theoretic, sequence-theoretic, real-vector, and end-to-end methods. This categorization is based on the mathematical formalisation of similarity used by each model. Ultimately, our survey aims to provide a comprehensive overview of the interlinkages between independently proposed methods for text similarity. By understanding the strengths and weaknesses of these methods, researchers can make informed decisions when designing novel approaches and algorithms. We hope this survey serves as a valuable resource for advancing the state-of-the-art in addressing the complex problem of text similarity. Wenchuan Mu, Kwan Hui Lim 0001 |
ASONAM | 2 |
| 2023 | Photozilla: An Image Dataset of Photography Styles and its Application to Visual Embedding and Style DetectionabstractThe widespread sharing of digital photography and images have led to the rapid development of various vision-related applications, such as photography style detection. Towards this effort, we introduce a photography style dataset termed Photozilla, which comprises over 990k images belonging to 10 different photographic styles. We used Photozilla to train 3 classification models for categorizing images into the relevant style and achieve an accuracy of ~96%. To better detect new photography styles that are constantly emerging, we also present a Siamese-based network that uses the trained classification models as the base architecture to adapt and classify unseen styles with only 25 training samples. Experiment results show an accuracy of over 68% in terms of identifying 10 additional distinct categories of photography styles. This dataset can be found at https://trisha025.github.io/Photozilla/. Trisha Singhal, Junhua Liu 0002, Wenchuan Mu, Luciënne T. M. Blessing, Kwan Hui Lim 0001 |
ASONAM | 5 |
| 2023 | SBTREC - A Transformer Framework for Personalized Tour Recommendation Problem with Sentiment AnalysisabstractWhen traveling to an unfamiliar city for holidays, tourists often rely on guidebooks, travel websites, or recommendation systems to plan their daily itineraries and explore popular points of interest (POIS). However, these approaches may lack optimization in terms of time feasibility, localities, and user preferences. In this paper, we propose the SBTREC algorithm: a BERT-based Trajectory Recommendation with sentiment analysis, for recommending personalized sequences of POIS as itineraries. Considering the locations, sightseeing, and travel time between consecutive Pots, our approach incorporates individual user preferences through the utilization of historical data. The key contributions of this work include analyzing users’ check-ins and uploaded photos to understand the relationship between Pot visits and distance. In addition, SBTREC also encompasses sentiment analysis to improve recommendation accuracy by understanding users’ preferences and satisfaction levels from reviews and comments about different Pots. Our proposed algorithms are evaluated against other sequence prediction methods using datasets from 8 cities. The results demonstrate that SBTREC achieves an average $\mathcal{F}_{1}$ score of 61.45%, outperforming baseline algorithms. The paper further discusses the flexibility of the SBTREC algorithm, its ability to adapt to different scenarios and cities without modification, and its potential for extension by incorporating additional information for more reliable predictions. Overall, SBTREC provides personalized and relevant Pot recommendations, enhancing tourists’ overall trip experiences. Future work includes fine-tuning personalized embeddings for users, with evaluation of users’ comments on Pots, to further enhance prediction accuracy. Ngai Lam Ho, Roy Ka-Wei Lee, Kwan Hui Lim 0001 |
IEEE Big Data | 3 |
| 2023 | A Transformer-Based Framework for POI-Level Social Post Geolocation
Kwan Hui Lim 0001, Teng Guo 0002, Junhua Liu 0002 |
ECIR (1) | 2 |
| 2022 | POIBERT: A Transformer-based Model for the Tour Recommendation ProblemabstractTour itinerary planning and recommendation are challenging problems for tourists visiting unfamiliar cities. Many tour recommendation algorithms only consider factors such as the location and popularity of Points of Interest (POIs) but their solutions may not align well with the user's own preferences and other location constraints. Additionally, these solutions do not take into consideration of the users' preference based on their past POIs selection. In this paper, we propose POIBERT, an algorithm for recommending personalized itineraries using the BERT language model on POIs. POIBERT builds upon the highly successful BERT language model with the novel adaptation of a language model to our itinerary recommendation task, alongside an iterative approach to generate consecutive POIs.Our recommendation algorithm is able to generate a sequence of POIs that optimizes time and users’ preference in POI categories based on past trajectories from similar tourists. Our tour recommendation algorithm is modeled by adapting the itinerary recommendation problem to the sentence completion problem in natural language processing (NLP). We also innovate an iterative algorithm to generate travel itineraries that satisfies the time constraints which is most likely from past trajectories. Using a Flickr dataset of seven cities, experimental results show that our algorithm out-performs many sequence prediction algorithms based on measures in recall, precision and F1-scores. Ngai Lam Ho, Kwan Hui Lim 0001 |
IEEE Big Data | 2 |
| 2022 | POI recommendation with queuing time and user interest awarenessabstractPoint-of-interest (POI) recommendation is a challenging problem due to different contextual information and a wide variety of human mobility patterns. Prior studies focus on recommendation that considers user travel spatiotemporal and sequential patterns behaviours. These studies do not pay attention to user personal interests, which is a significant factor for POI recommendation. Besides user interests, queuing time also plays a significant role in affecting user mobility behaviour, e.g., having to queue a long time to enter a POI might reduce visitor's enjoyment. Recently, attention-based recurrent neural networks-based approaches show promising performance in the next POI recommendation task. However, they are limited to single head attention, which can have difficulty in finding the appropriate user mobility behaviours considering complex relationships among POI spatial distances, POI check-in time, user interests and POI queuing times. In this research work, we are the first to consider queuing time and user interest awareness factors for next POI recommendation. We demonstrate how it is non-trivial to recommend a next POI and simultaneously predict its queuing time. To solve this problem, we propose a multi-task, multi-head attention transformer model called TLR-M_UI. The model recommends the next POIs to the target users and predicts queuing time to access the POIs simultaneously by considering user mobility behaviours. The proposed model utilises POIs description-based user personal interest that can also solve the new categorical POI cold start problem. Extensive experiments on six real-world datasets show that the proposed models outperform the state-of-the-art baseline approaches in terms of precision, recall, and F1-score evaluation metrics. The model also predicts and minimizes the queuing time. For the reproducibility of the proposed model, we have publicly shared our implementation code at GitHub (https://github.com/sajalhalder/TLR-M_UI). Sajal Halder, Kwan Hui Lim 0001, Jeffrey Chan, Xiuzhen Zhang 0001 |
Data Min. Knowl. Discov. | 2 |
| 2022 | Efficient itinerary recommendation via personalized POI selection and pruningabstractAbstract Personalized itinerary recommendation has garnered wide research interests for their ubiquitous applications. Recommending personalized itineraries is complex because of the large number of points of interest (POI) to consider in order to construct an itinerary based on visitors’ interest and preference, time budget and uncertain queuing time. Previous studies typically aim to plan itineraries that maximize POI popularity, visitors’ interest and minimize queuing time. However, existing solutions may not reflect visitor preferences because when creating itineraries, they prefer to recommend POIs with short prior visiting periods. These recommendations can conflict with real-life scenarios as visitors typically spend less time at POIs that they do not enjoy, thus leading to the inclusion of unsuitable POIs. Moreover, constructing itineraries based on selected POIs is a challenging and time-consuming process. Existing approaches involve searching through a large number of non-optimal, duplicate itineraries that are time-consuming to review and generate. To address these issues, we propose an adaptive Monte Carlo tree search (MCTS)-based reinforcement learning algorithmEffiTourRecusing an effective POI selection strategy by giving preference to POIs with long visiting times and short queuing times along with high POI popularity and visitor interest. In addition, to reduce non-optimal and duplicated itineraries generation, we propose an efficient MCTS search pruning technique to explore a smaller, more promising portion of solution space. Experiment results in real theme park datasets show clear advantages of our proposed method over baselines, where our method outperforms the current state-of-the-art by 20.89 to 52.32% in precision, 8.36 to 21.35% in F1-score and 40.00 to 67.64% in execution time. Sajal Halder, Kwan Hui Lim 0001, Jeffrey Chan, Xiuzhen Zhang 0001 |
Knowl. Inf. Syst. | 2 |
| 2021 | Analyzing Scientific Publications using Domain-Specific Word Embedding and Topic ModellingabstractThe scientific world is changing a tarapid pace, with new technology being developed and new trends being set at an increasing frequency. This paper presents a framework for conducting scientific analyses of academic publications, which is crucial to monitor research trends and identify potential innovations. This framework adopts and combines various techniques of Natural Language Processing, such as word embedding and topic modelling. Word embedding is used to capture semantic meanings of domain-specific words. We propose two novel scientific publication embedding, i.e., P UB-G and P UB-W, which are capable of learning semantic meanings of general as well as domain-specific words in various research fields. Thereafter, topic modelling is used to identify clusters of research topics within these larger research fields. We curated apublication dataset consisting of two conferences and two journals from 1995 to 2020 from two research domains. Experimental results show that our PUB-G and PUB-W embeddings are superior in comparison to other baseline embeddings by a margin of ~0.18-1.03 based on topic coherence. Trisha Singhal, Junhua Liu 0002, Luciënne T. M. Blessing, Kwan Hui Lim 0001 |
IEEE BigData | 4 |
| 2021 | Transformer-Based Multi-task Learning for Queuing Time Aware Next POI Recommendation
Sajal Halder, Kwan Hui Lim 0001, Jeffrey Chan, Xiuzhen Zhang 0001 |
PAKDD (2) | 2 |
| 2020 | Understanding Public Sentiments, Opinions and Topics about COVID-19 using TwitterabstractThe COVID-19 pandemic has caused widespread devastation throughout the world. In addition to the health and economical impacts, there is an enormous emotional toll associated with the constant stress of daily life with the numerous restrictions in place to combat the pandemic. To better understand the impact of COVID-19, we proposed a framework that utilizes public tweets to derive the sentiments, emotions and discussion topics of the general public in various regions and across multiple timeframes. Using this framework, we study and discuss various research questions relating to COVID-19, namely: (i) how sentiments/emotions change during the pandemic? (ii) how sentiments/emotions change in relation to global events? and (iii) what are the common topics discussed during the pandemic? Jolin Shaynn-Ly Kwan, Kwan Hui Lim 0001 |
ASONAM | 2 |
| 2020 | Urban Crowdsensing using Social Media: An Empirical Study on Transformer and Recurrent Neural NetworksabstractAn important aspect of urban planning is understanding crowd levels at various locations, which typically require the use of physical sensors. Such sensors are potentially costly and time consuming to implement on a large scale. To address this issue, we utilize publicly available social media datasets and use them as the basis for two urban sensing problems, namely event detection and crowd level prediction. One main contribution of this work is our collected dataset from Twitter and Flickr, alongside ground truth events. We demonstrate the usefulness of this dataset with two preliminary supervised learning approaches: firstly, a series of neural network models to determine if a social media post is related to an event and secondly a regression model using social media post counts to predict actual crowd levels. We discuss preliminary results from these tasks and highlight some challenges. Jerome Heng, Junhua Liu 0002, Kwan Hui Lim 0001 |
IEEE BigData | 3 |
| 2020 | EPIC30M: An Epidemics Corpus of Over 30 Million Relevant TweetsabstractSince the start of COVID-19, there has been several relevant corpora from various sources that were released to support research in this area. While these corpora are valuable in supporting analysis for this specific pandemic, researchers will benefit from additional benchmark corpora that contain other epidemics for better generalizability and to facilitate cross-epidemic pattern recognition and trend analysis tasks. During our research, we discover little disease related corpora in the literature that are sizable and rich enough to support such cross-epidemic analysis tasks. To address this issue, we present EPIC30M, a large-scale epidemic corpus that contains more than 30 million micro-blog posts, i.e., tweets crawled from Twitter, from year 2006 to 2020. EPIC30M contains a subset of 26.2 million tweets related to three general diseases, namely Ebola, Cholera and Swine Flu, and another subset of 4.7 million tweets of six global epidemic outbreaks, including the 2009 H1N1 Swine Flu, 2010 Haiti Cholera, 2012 Middle-East Respiratory Syndrome (MERS), 2013 West African Ebola, 2016 Yemen Cholera and 2018 Kivu Ebola. Furthermore, we explore and discuss the properties of this corpus with statistics of key terms and hashtags and trends analysis for each subset. Finally, we discuss the potential value and impact that EPIC30M could generate through a discussion of multiple use cases of cross-epidemic research topics that attract growing interest in recent years. These use cases span multiple research areas, such as epidemiological modeling, pattern recognition, natural language understanding and economical modeling. The corpus is publicly available at https://www.github.com/junhua/epic. Junhua Liu 0002, Trisha Singhal, Luciënne T. M. Blessing, Kristin L. Wood, Kwan Hui Lim 0001 |
IEEE BigData | 5 |
| 2020 | Urban Heat Islands: Beating the Heat with Multi-Modal Spatial AnalysisabstractIn today's highly urbanized environment, the Urban Heat Island (UHI) phenomenon is increasingly prevalent where surface temperatures in urbanized areas are found to be much higher than surrounding rural areas. Excessive levels of heat stress leads to problems at various levels, ranging from the individual to the world. At the individual level, UHI could lead to the human body being unable to cope and break-down in terms of core functions. At the world level, UHI potentially contributes to global warming and adversely affects the environment. Using a multi-modal dataset comprising remote sensory imagery, geospatial data and population data, we proposed a framework for investigating how UHI is affected by a city's urban form characteristics through the use of statistical modelling. Using Singapore as a case study, we demonstrate the usefulness of this framework and discuss our main findings i n u nderstanding the effects of UHI and urban form characteristics. Marcus Yong, Kwan Hui Lim 0001 |
IEEE BigData | 2 |
| 2020 | Strategic and Crowd-Aware Itinerary Recommendation
Junhua Liu 0002, Kristin L. Wood, Kwan Hui Lim 0001 |
ECML/PKDD (4) | 3 |
| 2019 | Spatio-temporal Event Detection using Poisson Model and Quad-tree on Geotagged Social MediaabstractIdentifying events happening in a specific locality is important as an early warning for accidents, protests, elections or breaking news. However, this location-specific event detection is challenging as the locations and types of events are not known beforehand. To address this problem, we propose an online spatio-temporal event detection system using social media that is able to detect events at different time and space resolutions. First, we exploit a quad-tree method to split the geographical space into multiscale regions based on the density of social media data. Then, we implement a statistical unsupervised approach using Poisson distribution and a smoothing method for highlighting regions with unexpected density of social posts. Further, event duration is estimated by merging events happening in the same region at consecutive time intervals. A post processing stage is introduced to filter out events that are spam, fake or wrong. Finally, we incorporate simple semantics by using social media entities to assess the integrity, and accuracy of detected events. The proposed method is evaluated using Twitter and Flickr for the city of Melbourne based on recall and precision measures. We also propose a new quality measure named strength index, which automatically measures how accurate the reported event is. Yasmeen M. George, Shanika Karunasekera, Aaron Harwood, Kwan Hui Lim 0001 |
IEEE BigData | 4 |
| 2019 | Sentiment-Aware and Personalized Tour RecommendationabstractItinerary planning is one of the most important tasks in tourism. A well-planned itinerary enhances the tourist experience and their visit satisfaction in new cities. However, the task of planning personalized tour itineraries is complicated by tourists with different interest preferences. Furthermore, there is an added complexity of recommending an itinerary with discrete budget, time and cost. Due to an increase in web-technologies and online geo-location services, there is emerging research targeting itinerary recommendation based on each tourist's interest, preferences and trip constraints. While several research works consider tourist interest, they adopt a simple measure based on the number of times a tourist has visited a place or the number of photos taken by the tourist at a place. Our research proposes an improved sentiment-aware personalized tour planner that considers each tourist's interests based on his/her sentiments on specific categories relative to his/her overall preferences. Unlike the previous approaches that do not consider the actual opinion based preferences, our proposed approach determines user interests based on their sentiments associated with their written text about a place of their visit. This interest measure is based on the intuition that users are more likely to post favorable comments about places they like. Using a dataset from Twitter, we compare our proposed algorithm against the baseline and experimental results show that our algorithm obtained superior performance in terms of tour precision, recall, Fl-score and overall popularity. Prarthana Padia, Kwan Hui Lim 0001, Jeffrey Chan, Aaron Harwood |
IEEE BigData | 2 |
| 2019 | Identifying and Understanding Business Trends using Topic Models with Word EmbeddingabstractTopic modelling and trend analysis are increasingly important in today's digital world, especially for identifying promising business ideas and trends. With the increasing amount of data being generated daily, a key challenge is to effectively identify emerging business ideas/topics and trends from this large volume of data. Towards this effort, we introduce a framework that allows us to identify promising business ideas from a large stream of academic papers. Academic papers are suitable for this purpose as they study emerging areas and problems in different domains. Our framework comprises three main components, namely: (i) a data collection component that retrieves academic papers and their meta-data; (ii) a topic modelling algorithm that combines traditional topic modelling techniques with recent advances in word embeddings; and (iii) a trend analysis component that allows us to visualize the popularity of different business trends/topics across time. Results on a corpus of 287k academic papers show that our proposed methods outperform the standard baselines based on topic coherence scores and also allows us to understand key temporal trends. Yun Ning Pek, Kwan Hui Lim 0001 |
IEEE BigData | 2 |
| 2019 | Tour recommendation and trip planning using location-based social media: a survey
Kwan Hui Lim 0001, Jeffrey Chan, Shanika Karunasekera, Christopher Leckie |
Knowl. Inf. Syst. | 1 |
| 2018 | RAPID: Real-time Analytics Platform for Interactive Data Mining
Kwan Hui Lim 0001, Sachini Jayasekara, Shanika Karunasekera, Aaron Harwood, Lucia Falzon, John Dunn, Glenn Burgess |
ECML/PKDD (3) | 1 |
| 2018 | Personalized trip recommendation for tourists based on user interests, points of interest visit durations and visit recency
Kwan Hui Lim 0001, Jeffrey Chan, Christopher Leckie, Shanika Karunasekera |
Knowl. Inf. Syst. | 1 |
| 2017 | ClusTop: A clustering-based topic modelling algorithm for twitter using word networksabstractTwitter is a popular microblogging service, where users frequently engage in discussions about various topics of interest, ranging from popular topics (e.g., music) to niche topics (e.g., politics). With the large amount of tweets, a key challenge is to automatically model and determine the discussion topics without having prior knowledge of the types and number of topics, or requiring the technical expertise to define various algorithmic parameters. For this purpose, we propose the Clustering-based Topic Modelling (ClusTop) algorithm that constructs various types of word network and automatically determines the discussion topics using community detection approaches. Unlike traditional topic models, ClusTop is able to automatically determine the appropriate number of topics and does not require numerous parameters to be set. The ClusTop algorithm is also able to capture the syntactic meaning in tweets via the use of bigrams, trigrams and other word combinations in constructing the word network graph. Using three Twitter datasets with labelled crises and events as topics, ClusTop has been shown to outperform various baselines in terms of topic coherence, pointwise mutual information, precision, recall and F-score. Kwan Hui Lim 0001, Shanika Karunasekera, Aaron Harwood |
IEEE BigData | 1 |
| 2017 | Spatial-based topic modelling using wikidata knowledge baseabstractTopic modelling is a well-studied field that aims to identify topics from traditional documents such as news articles and reports. More recently, Latent Dirichlet Allocation (LDA) and its variants, have been applied on social media platforms to model and study topics relating to sports, politics and companies. While these applications were able to successfully identify the general topics, we posit that standard LDA can be augmented with spatial and temporal considerations based on the geo-coordinates and timestamps of social media posts. Towards this effort, we propose a spatial and temporal variant of LDA to better detect more specific topics, such as a particular art exhibit held at a museum or a security incident happening on a particular day. We validate our approach on a Twitter dataset and find that the detected topics are well-aligned to real-life events happening on the specific days and locations. Kwan Hui Lim 0001, Shanika Karunasekera, Aaron Harwood, Lucia Falzon |
IEEE BigData | 1 |
| 2017 | Personalized Itinerary Recommendation with Queuing Time AwarenessabstractPersonalized itinerary recommendation is a complex and time-consuming problem, due to the need to recommend popular attractions that are aligned to the interest preferences of a tourist, and to plan these attraction visits as an itinerary that has to be completed within a specific time limit. Furthermore, many existing itinerary recommendation systems do not automatically determine and consider queuing times at attractions in the recommended itinerary, which varies based on the time of visit to the attraction, e.g., longer queuing times at peak hours. To solve these challenges, we propose the PersQ algorithm for recommending personalized itineraries that take into consideration attraction popularity, user interests and queuing times. We also implement a framework that utilizes geo-tagged photos to derive attraction popularity, user interests and queuing times, which PersQ uses to recommend personalized and queue-aware itineraries. We demonstrate the effectiveness of PersQ in the context of five major theme parks, based on a Flickr dataset spanning nine years. Experimental results show that PersQ outperforms various state-of-the-art baselines, in terms of various queuing-time related metrics, itinerary popularity, user interest alignment, recall, precision and F1-score. Kwan Hui Lim 0001, Jeffrey Chan, Shanika Karunasekera, Christopher Leckie |
SIGIR | 1 |
| 2016 | Improving Personalized Trip Recommendation by Avoiding CrowdsabstractThere has been a growing interest in recommending trips for tourists using location-based social networks. The challenge of trip recommendation not only lies in searching for relevant points-of-interest (POIs) to form a personalized trip, but also selecting the best time of day to visit the POIs. Popular POIs can be too crowded during peak times, resulting in long queues and delays. In this work, we propose the Personalized Crowd-aware Trip Recommendation (PersCT) algorithm to recommend personalized trips that also avoid the most crowded times of the POIs. We model the problem as an extension of the Orienteering Problem with multiple constraints. We extract user interests by collaborative filtering and we propose an extension of the Ant Colony Optimisation algorithm to merge user interests with POI popularity and crowdedness data to recommend trips. We evaluate our algorithm using foot traffic information obtained from a real-life pedestrian sensor dataset and user travel histories extracted from a Flickr photo dataset. We show that our algorithm out-performs several benchmarks in achieving a balance between conflicting objectives by satisfying user interests while reducing the crowdedness of the trips. Christopher Leckie, Jeffrey Chan, Kwan Hui Lim 0001, Tharshan Vaithianathan |
CIKM | 4 |
| 2015 | Detecting Location-Centric Communities Using Social-Spatial Links with Temporal Constraints
Kwan Hui Lim 0001, Jeffrey Chan, Christopher Leckie, Shanika Karunasekera |
ECIR | 1 |
| 2012 | Tweets Beget Propinquity: Detecting Highly Interactive Communities on Twitter Using Tweeting LinksabstractMany community detection algorithms have been developed to detect communities on Online Social Networks (OSN). However, these algorithms are based only on topological links and researchers have observed that many topological links do not translate to actual user interaction. As such, many members of the detected communities do not communicate frequently to each other. This inactivity creates a problem in targeted advertising and viral marketing which requires the community to be highly active so as to allow the diffusion of product/service information. We propose an approach to detect highly interactive Twitter communities that share common interests, based on the frequency and patterns of direct tweeting among users, rather than the topological information implicit in follower/following links. From a topological aspect, we show that our method detects communities that are more cohesive and connected within different interest groups. We also show that the detected communities interact actively about the specific interests, based on the high frequency of #hash tags and @mentions related to this interest. In addition, we study the trends in their tweeting patterns such as how they follow and unfollow other users. Kwan Hui Lim 0001, Amitava Datta |
Web Intelligence | 1 |