VLDB 2026 Research / reviewers in the wild / expert
Huayu Wu 0001
dblp:99/4453
· DBLP profile ↗
44ranked-venue papers
15as first author
1since 2021 · last 2023
0000-0001-9942-6885ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 39 · 15 first-author · 1 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
12 papers |
Information retrieval · 30% Data mining · 26% Recommender systems · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Smart cities and intelligent transportation · 95% Computational social science and digital humanities · 5% | |
| Network and information security
2 papers |
Privacy and data protection · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Ubiquitous computing and smart environments · 100% | |
| Artificial intelligence
2 papers |
Graph learning · 82% Information extraction and text analysis · 18% |
Topics — the 30 heaviest of 40, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › spatiotemporal data mining
trajectory data mining |
0.9 | 3 | 2019 | TourSense: A Framework for Tourist Identification and Analytics Using Transport Data · IEEE Trans. Knowl. Data Eng. 2019 From Raw Footprints to Personal Interests: Bridging the Semantic Gap via Trip Intention Aggregation · ICDE 2017 A General and Parallel Platform for Mining Co-Movement Patterns over Large-scale Trajectories · Proc. VLDB Endow. 2016 |
Information retrieval › ranking
learning to rank |
0.7 | 1 | 2023 | A Semantic Search Framework for Similar Audit Issue Recommendation in Financial Industry · WSDM 2023 |
Information retrieval
ranking |
0.7 | 1 | 2023 | A Semantic Search Framework for Similar Audit Issue Recommendation in Financial Industry · WSDM 2023 |
Information retrieval › search engines
semantic search |
0.7 | 1 | 2023 | A Semantic Search Framework for Similar Audit Issue Recommendation in Financial Industry · WSDM 2023 |
Machine learning › Graph learning
graph propagation |
0.4 | 1 | 2019 | TourSense: A Framework for Tourist Identification and Analytics Using Transport Data · IEEE Trans. Knowl. Data Eng. 2019 |
Recommender systems › user interest modeling
user interest inference |
0.3 | 1 | 2018 | CO2: Inferring Personal Interests From Raw Footprints by Connecting the Offline World with the Online World · ACM Trans. Inf. Syst. 2018 |
Ubiquitous computing and smart environments › mobile crowdsourcing › crowdsensing
mobile crowdsensing |
0.3 | 1 | 2018 | Smartphone Sensing Meets Transport Data: A Collaborative Framework for Transportation Service Analytics · IEEE Trans. Mob. Comput. 2018 |
Ubiquitous computing and smart environments › mobile sensing
smartphone sensing |
0.3 | 1 | 2018 | Smartphone Sensing Meets Transport Data: A Collaborative Framework for Transportation Service Analytics · IEEE Trans. Mob. Comput. 2018 |
Knowledge graphs
semantic enrichment |
0.3 | 1 | 2017 | From Raw Footprints to Personal Interests: Bridging the Semantic Gap via Trip Intention Aggregation · ICDE 2017 |
Spatial and temporal data management
trajectory data |
0.3 | 1 | 2017 | From Raw Footprints to Personal Interests: Bridging the Semantic Gap via Trip Intention Aggregation · ICDE 2017 |
Recommender systems › user interest modeling
user interest mining |
0.3 | 1 | 2017 | From Raw Footprints to Personal Interests: Bridging the Semantic Gap via Trip Intention Aggregation · ICDE 2017 |
Data mining › pattern mining › temporal pattern mining
co-movement pattern mining |
0.2 | 1 | 2016 | A General and Parallel Platform for Mining Co-Movement Patterns over Large-scale Trajectories · Proc. VLDB Endow. 2016 |
Data mining
pattern mining |
0.2 | 1 | 2016 | A General and Parallel Platform for Mining Co-Movement Patterns over Large-scale Trajectories · Proc. VLDB Endow. 2016 |
Data stream processing
streaming analytics |
0.2 | 1 | 2016 | Mercury: Metro density prediction with recurrent neural network on streaming CDR data · ICDE 2016 |
Spatial and temporal data management
trajectory data management |
0.2 | 1 | 2016 | A General and Parallel Platform for Mining Co-Movement Patterns over Large-scale Trajectories · Proc. VLDB Endow. 2016 |
Visualization and visual analytics › geospatial visualization
urban data visualization |
0.2 | 1 | 2016 | An Intelligent System for Taxi Service Monitoring, Analytics and Visualization · IJCAI 2016 |
Privacy and data protection
anonymization |
0.2 | 1 | 2016 | Fuzzy trajectory linking · ICDE 2016 |
Privacy and data protection
de-anonymization |
0.2 | 1 | 2016 | Fuzzy trajectory linking · ICDE 2016 |
Privacy and data protection › location privacy
trajectory privacy |
0.2 | 1 | 2016 | Fuzzy trajectory linking · ICDE 2016 |
Data models and query languages
XML data management |
0.2 | 2 | 2012 | Labeling Dynamic XML Documents: An Order-Centric Approach · IEEE Trans. Knowl. Data Eng. 2012 DDE: from dewey to a fully dynamic XML labeling scheme · SIGMOD Conference 2009 |
Data models and query languages › XML data management
XML labeling scheme |
0.2 | 2 | 2012 | Labeling Dynamic XML Documents: An Order-Centric Approach · IEEE Trans. Knowl. Data Eng. 2012 DDE: from dewey to a fully dynamic XML labeling scheme · SIGMOD Conference 2009 |
Smart cities and intelligent transportation
public transit |
0.2 | 1 | 2015 | FTT: A System for Finding and Tracking Tourists in Public Transport Services · SIGMOD Conference 2015 |
Data mining › spatiotemporal data mining › trajectory data mining
trajectory pattern mining |
0.2 | 1 | 2015 | FTT: A System for Finding and Tracking Tourists in Public Transport Services · SIGMOD Conference 2015 |
Smart cities and intelligent transportation › public transit
public transit analytics |
0.2 | 1 | 2014 | Identifying tourists from public transport commuters · KDD 2014 |
Data mining › predictive modeling
classification |
0.2 | 1 | 2014 | Identifying tourists from public transport commuters · KDD 2014 |
Data stream processing › stream processing systems
privacy-preserving stream processing |
0.2 | 1 | 2013 | A privacy preserving framework for managing vehicle data in road pricing systems · KDD 2013 |
Privacy and data protection
privacy-preserving data management |
0.2 | 1 | 2013 | A privacy preserving framework for managing vehicle data in road pricing systems · KDD 2013 |
Indexing and storage engines › index maintenance
dynamic indexing |
0.1 | 1 | 2012 | Labeling Dynamic XML Documents: An Order-Centric Approach · IEEE Trans. Knowl. Data Eng. 2012 |
Information retrieval
indexing |
0.1 | 1 | 2012 | Labeling Dynamic XML Documents: An Order-Centric Approach · IEEE Trans. Knowl. Data Eng. 2012 |
Ubiquitous computing and smart environments
mobile sensing |
0.1 | 1 | 2018 | Smartphone Sensing Meets Transport Data: A Collaborative Framework for Transportation Service Analytics · IEEE Trans. Mob. Comput. 2018 |
Methods — techniques the papers use, named apart from their topics
interactive user interface · 0.8graph-based iterative propagation · 0.8queuing analytics · 0.7data fusion · 0.7cross-encoder · 0.7bi-encoder · 0.7TF-IDF · 0.7probabilistic generative model · 0.6latent dirichlet allocation · 0.6parallel stream processing · 0.5trip intention inference · 0.3probabilistic framework · 0.3weight sharing · 0.2spatial data processing · 0.2recurrent neural network · 0.2parallel framework · 0.2naive bayes · 0.2hypothesis testing · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Semantic Search Framework for Similar Audit Issue Recommendation in Financial IndustryabstractAudit issues summarize the findings during audit reviews and provide valuable insights of risks and control gaps in a financial institute. Despite the wide use of data analytics and NLP in financial services, due to the diverse coverage and lack of annotations, there are very few use cases that analyze audit issue writing and derive insights from it. In this paper, we propose a deep learning based semantic search framework to search, rank and recommend similar past issues based on new findings. We adopt a two-step approach. First, a TF-IDF based search algorithm and a Bi-Encoder are used to shortlist a set of issue candidates based on the input query. Then a Cross-Encoder will re-rank the candidates and provide the final recommendation. We will also demonstrate how the models are deployed and integrated with the existing workbench to benefit auditors in their daily work. Chuchu Zhang, Can Song, Samarth Agarwal, Huayu Wu 0001, John Jianan Lu |
WSDM | 4 |
| 2019 | TourSense: A Framework for Tourist Identification and Analytics Using Transport DataabstractWe advocate for and presentTourSense, a framework for tourist identification and preference analytics using city-scale transport data (bus, subway, etc.). Our work is motivated by the observed limitations of utilizing traditional data sources (e.g., social media data and survey data) that commonly suffer from the limited coverage of tourist population and unpredictable information delay.TourSensedemonstrates how the transport data can overcome these limitations and provide better insights for different stakeholders, typically including tour agencies, transport operators, and tourists themselves. Specifically, we first propose a graph-based iterative propagation learning algorithm to recognize tourists from public commuters. Taking advantage of the trace data from the identified tourists, we then design a tourist preference analytics model to learn and predict their next tour, where an interactive user interface is implemented to ease the information access and gain the insights from the analytics results. Experiments with real-world datasets (from over 5.1 million commuters and their 462 million trips) show the promise and effectiveness of the proposed framework: the Macro and Micro F1 scores of the tourist identification system achieve 0.8549 and 0.7154, respectively, whereas the tourist preference analytics system improves the baselines by at least 23.53 and 11.44 percent in terms of precision and recall. Yu Lu 0003, Huayu Wu 0001, Xin Liu 0027, Penghe Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2018 | Fast and Scalable Big Data Trajectory Clustering for Understanding Urban MobilityabstractClustering of large-scale vehicle trajectories is an important aspect for understanding urban traffic patterns, particularly for optimizing public transport routes and frequencies and improving the decisions made by authorities. Existing trajectory clustering schemes are not well suited to large numbers of trajectories in dense city road networks due to the difficulty in finding a representative distance measure between trajectories that can scale to very large datasets. In this paper, we propose a novel Dijkstra-based dynamic time warping distance measure, trajDTW between two trajectories, which is suitable for large numbers of overlapping trajectories in a dense road network as found in major cities around the world. We also propose a novel fast-clusiVAT algorithm that can suggest the number of clusters in a trajectory dataset and identify and visualize the trajectories belonging to each cluster. We conduct experiments on a large-scale taxi trajectory dataset consisting of 3.28 million trajectories obtained from the GPS traces of 15 061 taxis within Singapore over a period of one month. Our analysis finds 13 trajectory clusters spanning the major expressways of Singapore, each of which can be further divided into two sub-clusters based on the travel direction. For each cluster, we provide a time-based distribution of trajectories to yield insights into how urban mobility patterns change with the time of day. We compare the trajectory clusters obtained using our approach with those obtained using popular general and trajectory specific clustering frameworks: DBSCAN, OPTICS, NETSCAN, and NEAT. We demonstrate that the clusters obtained using our novel fast-clusiVAT framework are better than those obtained using other clustering schemes, evaluated based on two internal cluster validity measures: Dunn's and Silhouette indices. Moreover, our fast-clusiVAT algorithm achieves significant speedup over a comparable approach without loss of cluster quality. Huayu Wu 0001, Sutharshan Rajasegarar, Christopher Leckie, Shonali Krishnaswamy, Marimuthu Palaniswami |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Smartphone Sensing Meets Transport Data: A Collaborative Framework for Transportation Service AnalyticsabstractWe advocate for and introduce TRANSense, a framework for urban transportation service analytics that combines participatory smartphone sensing data with city-scale transportation-related transactional data (taxis, trains, etc.). Our work is driven by the observed limitations of using each data type in isolation: (a) commonly-used anonymous city-scale datasets (such as taxi bookings and GPS trajectories) provide insights into the aggregate behavior of transport infrastructure, but fail to reveal individual-specific transport experiences (e.g., wait times in taxi queues); while (b) mobile sensing data can capture individual-specific commuting-related activities, but suffers from accuracy and energy overhead challenges due to usage artefacts and lack of appropriate sensing triggers. TRANSense demonstrates how a judicious fusion of such disparate data sources can overcome these challenges and offer novel insights. We detail two examples: (a) Taxi Service Analyzer that provides accurate detection of commuter queuing for taxis and estimates their wait time, by using taxi trip records to identify potential taxi locations with high demand and subsequently selectively triggering mobile sensing-based queuing analytics on nearby commuters; and (b) Subway Boarding Analyzer that identifies instances when passengers fail to board arriving trains, by first estimating train arrivals from temporal patterns of passenger egress at station gantries, and then using mobile sensing-based analysis of commuter movement behavior on platforms. Experiments with real-world datasets (from over 20,000 taxis and 1.7 million commuters in Singapore) show the power of this approach: the taxi service analyzer detects commuter queuing with over 90 percent accuracy with negligible energy overhead and estimates wait times with error margins below 15 percent, whereas the subway boarding analyzer can detect failed boarding events with a precision of over 90 percent (more than thrice what is achievable through purely mobile sensing). Yu Lu 0003, Archan Misra, Wen Sun 0004, Huayu Wu 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2018 | CO2: Inferring Personal Interests From Raw Footprints by Connecting the Offline World with the Online WorldabstractUser-generated trajectories (UGTs), such as travel records from bus companies, capture rich information of human mobility in the offline world. However, some interesting applications of these raw footprints have not been exploited well due to the lack of textual information to infer the subject’s personal interests. Although there is rich semantic information contained in the spatial- and temporal-aware user-generated contents (STUGC) published in the online world, such as Twitter, less effort has been made to utilize this information to facilitate the interest discovery process. In this article, we design an effective probabilistic framework named CO 2 to connect the offline world with the online world in order to discover users’ interests directly from their raw footprints in UGT. CO 2 first infers trip intentions by utilizing the semantic information in STUGC and then discovers user interests by aggregating the intentions. To evaluate the effectiveness of CO 2 , we use two large-scale real-world datasets as a case study and further conduct a questionnaire survey to show the superior performance of CO 2 . Long Guo, Dongxiang Zhang, Yuan Wang 0003, Huayu Wu 0001, Bin Cui 0001, Kian-Lee Tan |
ACM Trans. Inf. Syst. | 4 |
| 2017 | STA: A Spatio-Temporal Thematic Analytics Framework for Urban Ground Sensing
Guizi Chen, Liang Yu 0005, Wee Siong Ng, Huayu Wu 0001, Usha Nanthani Kunasegaran |
ADMA | 4 |
| 2017 | Mobile Robot Scheduling with Multiple Trips and Time Windows
Shudong Liu 0003, Huayu Wu 0001, Shili Xiang, Xiaoli Li 0001 |
ADMA | 2 |
| 2017 | From Raw Footprints to Personal Interests: Bridging the Semantic Gap via Trip Intention AggregationabstractUser-generated trajectories (UGT), such as GPS footprints from wearable devices or travel records from bus companies, capture rich information of human mobility and urban dynamics in the offline world. In this paper, our objective is to enrich these raw footprints and discover the users' personal interests by utilizing the semantic information contained in the spatial-and temporal-aware user-generated contents (STUGC) published in the online world. We design a novel probabilistic framework named CO2to connect the offline world with the online world in order to discover the users' interests directly from their raw footprints in UGT. In particular, we first propose a latent probabilistic generative model named STLDA to infer the intention attached with each trip, and then aggregate the extracted trip intentions to discover the users' personal interests. To tackle the inherent sparsity and noisiness problems of the tags in STUGC, STLDA considers the inner correlation between tags (i.e., semantic, spatial and temporal correlation) on the topic-level. To evaluate the effectiveness of CO2, we utilize a dataset containing three months of data with 5.3 billion bus records and a Twitter dataset with 1.5 million tweets published in 6 months in Singapore as a case study. Experimental results on these two real-world datasets show that CO2is effective in discovering user interests and improves the precision of the state-of-the-art method by 280%. In addition, we also conduct a questionnaire survey in Singapore to evaluate the effectiveness of CO2. The results further validate the superiority of CO2. Long Guo, Dongxiang Zhang, Huayu Wu 0001, Bin Cui 0001, Kian-Lee Tan |
ICDE | 3 |
| 2016 | Bus Routes Design and Optimization via Taxi Data AnalyticsabstractPublic bus services are often planned in the context of urban planning. For a city with efficient and extensive network of public transportation system like Singapore, enhancing the existing coverage of bus service to meet the dynamic mobility needs of the population requires data mining approach. Specifically, frequent taxi rides between two locations at a period of time may suggest possible poor coverage of public transport service, if not lacking of the public transport service. In this paper, we describe a proof of concept effort to discover this weakness and its improvement in public transportation system via mining of taxi ride dataset. We cluster taxi rides dataset to determine some popular taxi rides in Singapore. From the clustered taxi rides, we filter and select only the clusters whose commuting via existing public transport are tortuous if not unreachable door-to-door. Based on the discovered travel pattern, we propose new bus routes that serve the passengers of these clusters. We formulate the bus planning problem as an optimization of directed cycle graph, and present it's preliminary solution and results. We showcase our idea in the case of Singapore. Seong-Ping Chuah, Huayu Wu 0001, Yu Lu 0003, Liang Yu 0005, Stéphane Bressan |
CIKM | 2 |
| 2016 | Routing an Autonomous Taxi with Reinforcement LearningabstractSingapore's vision of a Smart Nation encompasses the development of effective and efficient means of transportation. The government's target is to leverage new technologies to create services for a demand-driven intelligent transportation model including personal vehicles, public transport, and taxis. Singapore's government is strongly encouraging and supporting research and development of technologies for autonomous vehicles in general and autonomous taxis in particular. The design and implementation of intelligent routing algorithms is one of the keys to the deployment of autonomous taxis. In this paper we demonstrate that a reinforcement learning algorithm of the Q-learning family, based on a customized exploration and exploitation strategy, is able to learn optimal actions for the routing autonomous taxis in a real scenario at the scale of the city of Singapore with pick-up and drop-off events for a fleet of one thousand taxis. Miyoung Han, Pierre Senellart, Stéphane Bressan, Huayu Wu 0001 |
CIKM | 4 |
| 2016 | Cost Minimization and Social Fairness for Spatial Crowdsourcing Tasks
Qing Liu 0020, Talel Abdessalem, Huayu Wu 0001, Zihong Yuan, Stéphane Bressan |
DASFAA (1) | 3 |
| 2016 | Mercury: Metro density prediction with recurrent neural network on streaming CDR dataabstractTelecommunication companies possess mobility information of their phone users, containing accurate locations and velocities of commuters travelling in public transportation system. Although the value of telecommunication data is well believed under the smart city vision, there is no existing solution to transform the data into actionable items for better transportation, mainly due to the lack of appropriate data utilization scheme and the limited processing capability on massive data. This paper presents the first ever system implementation of real-time public transportation crowd prediction based on telecommunication data, relying on the analytical power of advanced neural network models and the computation power of parallel streaming analytic engines. By analyzing the feeds of caller detail record (CDR) from mobile users in interested regions, our system is able to predict the number of metro passengers entering stations, the number of waiting passengers on the platforms and other important metrics on the crowd density. New techniques, including geographical-spatial data processing, weight-sharing recurrent neural network, and parallel streaming analytical programming, are employed in the system. These new techniques enable accurate and efficient prediction outputs, to meet the real-world business requirements from public transportation system. Victor C. Liang, Richard T. B. Ma, Wee Siong Ng, Marianne Winslett, Huayu Wu 0001, Shanshan Ying |
ICDE | 6 |
| 2016 | Fuzzy trajectory linkingabstractToday, people can access various services with smart carry-on devices, e.g., surf the web with smart phones, make payments with credit cards, or ride a bus with commuting cards. In addition to the offered convenience, the access of such services can reveal their traveled trajectory to service providers. Very often, a user who has signed up for multiple services may expose her trajectory to more than one service providers. This state of affairs raises a privacy concern, but also an opportunity. On one hand, several colluding service providers, or a government agency that collects information from such service providers, may identify and reconstruct users' trajectories to an extent that can be threatening to personal privacy. On the other hand, the processing of such rich data may allow for the development of better services for the common good. In this paper, we take a neutral standpoint and investigate the potential for trajectories accumulated from different sources to be linked so as to reconstruct a larger trajectory of a single person. We develop a methodology, called fuzzy trajectory linking (FTL) that achieves this goal, and two instantiations thereof, one based on hypothesis testing and one on Naïve-Bayes. We provide a theoretical analysis for factors that affect FTL and use two real datasets to demonstrate that our algorithms effectively achieve their goals. Huayu Wu 0001, Mingqiang Xue, Jianneng Cao, Panagiotis Karras, Wee Siong Ng, Kee Kiat Koo |
ICDE | 1 |
| 2016 | An Intelligent System for Taxi Service Monitoring, Analytics and Visualization
Yu Lu 0003, Gim Guan Chua, Huayu Wu 0001, Clement Shi Qi Ong |
IJCAI | 3 |
| 2016 | Understanding Urban Mobility via Taxi Trip ClusteringabstractClustering of a large amount of taxi GPS mobility data helps to understand the spatio-temporal dynamics for the applications of urban planning and transportation. In this paper we cluster the origin-destination pairs of the passenger taxi rides to provide useful insight into the city mobility patterns, urban hot-spots, road network usage and general patterns of the crowd movement within the city of Singapore. We perform experiments on a large scale Singapore taxi dataset consisting of more than 10 million passenger origin-destination GPS points. We use the clusi VAT sampling scheme to obtain the sample trips which return coarse clusters describing the major crowd movement and reduce the data points that are not captured by the coarse clusters and may bring in noises during fine-grained clustering. After the sampling step we use the well known density based clustering algorithm DBSCAN to find cluster structure in the sampled data points and later extend it to the rest of the dataset using nearest prototype rule. We report 24 trip clusters from the dataset which are compact enough to draw meaningful conclusions about the city mobility patterns and the number of trips in each cluster is large enough to be representative of the general traffic movement. Huayu Wu 0001, Yu Lu 0003, Shonali Krishnaswamy, Marimuthu Palaniswami |
MDM | 2 |
| 2016 | A General and Parallel Platform for Mining Co-Movement Patterns over Large-scale TrajectoriesabstractDiscovering co-movement patterns from large-scale trajectory databases is an important mining task and has a wide spectrum of applications. Previous studies have identified several types of interesting co-movement patterns and show-cased their usefulness. In this paper, we make two key contributions to this research field. First, we propose a more general co-movement pattern to unify those defined in the past literature. Second, we propose two types of parallel and scalable frameworks and deploy them on Apache Spark. To the best of our knowledge, this is the first work to mine co-movement patterns in real life trajectory databases with hundreds of millions of points. Experiments on three real life large-scale trajectory datasets have verified the efficiency and scalability of our proposed solutions. Dongxiang Zhang, Huayu Wu 0001, Kian-Lee Tan |
Proc. VLDB Endow. | 3 |
| 2015 | Expressing and Processing Path-Centric XML Queries
Huayu Wu 0001, Dongxu Shao, Ruiming Tang, Tok Wang Ling, Stéphane Bressan |
DEXA (2) | 1 |
| 2015 | Flexible Data Management across XML and Relational Models: A Semantic Approach
Huayu Wu 0001, Tok Wang Ling, Wee Siong Ng |
ER | 1 |
| 2015 | Locating Self-Collection Points for Last-Mile Logistics Using Public Transport Data
Huayu Wu 0001, Dongxu Shao, Wee Siong Ng |
PAKDD (1) | 1 |
| 2015 | FTT: A System for Finding and Tracking Tourists in Public Transport ServicesabstractThe tourism industry is a key economic driver for many cities. To understand tourists' traveling patterns can help both public and private relevant sectors design and improve their services to serve tourists better and get additional values from it. The existing approaches to discover tourists' traveling pattern focus on small sets of known tourists extracted from social media or other channels. The accuracy of the mining result cannot be guaranteed due to the small and bias set of samples. Huayu Wu 0001, Jo-Anne Tan, Wee Siong Ng, Mingqiang Xue, Wei Chen 0025 |
SIGMOD Conference | 1 |
| 2015 | Truth Finding with Attribute PartitioningabstractTruth finding is the problem of determining which of the statements made by contradictory sources is correct, in the absence of prior information on the trustworthiness of the sources. A number of approaches to truth finding have been proposed, from simple majority voting to elaborate iterative algorithms that estimate the quality of sources by corroborating their statements. In this paper, we consider the case where there is an inherent structure in the statements made by sources about real-world objects, that imply different quality levels of a given source on different groups of attributes of an object. We do not assume this structuring given, but instead find it automatically, by exploring and weighting the partitions of the sets of attributes of an object, and applying a reference truth finding algorithm on each subset of the optimal partition. Our experimental results on synthetic and real-world datasets show that we obtain better precision at truth finding than baselines in cases where data has an inherent structure. Mouhamadou Lamine Ba, Roxana Horincar, Pierre Senellart, Huayu Wu 0001 |
WebDB | 4 |
| 2014 | A*DAX: A Platform for Cross-Domain Data Linking, Sharing and Analytics
Narayanan Amudha, Gim Guan Chua, Eric Siew Khuan Foo, Shen-Tat Goh, Shuqiao Guo, Paul Min Chim Lim, Mun-Thye Mak, Muhammad Cassim Mahmud Munshi, See-Kiong Ng, Wee Siong Ng, Huayu Wu 0001 |
DASFAA (2) | 11 |
| 2014 | Parallelizing Structural Joins to Process Queries over Big XML Data Using MapReduce
Huayu Wu 0001 |
DEXA (2) | 1 |
| 2014 | Identifying tourists from public transport commutersabstractTourism industry has become a key economic driver for Singapore. Understanding the behaviors of tourists is very important for the government and private sectors, e.g., restaurants, hotels and advertising companies, to improve their existing services or create new business opportunities. In this joint work with Singapore's Land Transport Authority (LTA), we innovatively apply machine learning techniques to identity the tourists among public commuters using the public transportation data provided by LTA. On successful identification, the travelling patterns of tourists are then revealed and thus allow further analyses to be carried out such as on their favorite destinations, region of stay, etc. Technically, we model the tourists identification as a classification problem, and design an iterative learning algorithm to perform inference with limited prior knowledge and labeled data. We show the superiority of our algorithm with performance evaluation and comparison with other state-of-the-art learning algorithms. Further, we build an interactive web-based system for answering queries regarding the moving patterns of the tourists, which can be used by stakeholders to gain insight into tourists' travelling behaviors in Singapore. Mingqiang Xue, Huayu Wu 0001, Wei Chen 0025, Wee Siong Ng, Gin Howe Goh |
KDD | 2 |
| 2014 | HipStream: A Privacy-Preserving System for Managing Mobility Data StreamsabstractPersonal mobile data are being extensively collected by various service providers, in the form of data stream. Most service providers promise their customers for not misusing their data by paper-based agreement. However, the customers have no way to know whether the agreements are strictly followed or not, unless any scandals of private data misuse are revealed. To guarantee the correct use of customers' personal data and assure them of the service safety, system-level data privacy control between the data owners (i.e., Customers) and the data users (i.e., Service providers) is in compelling need. Inspired by the concept of Hippocratic data management, we design and implement a system, Hip Stream to systemically enforce different Hippocratic principles to preserve data providers' privacy when they send their data stream for services. In this paper, we describe the architecture of the Hip Stream system and demonstrate how it meets those privacy principles. Huayu Wu 0001, Shili Xiang, Wee Siong Ng, Wei Wu 0020, Mingqiang Xue |
MDM (1) | 1 |
| 2013 | Querying Semi-structured Data with Mutual Exclusion
Huayu Wu 0001, Ruiming Tang, Tok Wang Ling |
DASFAA (1) | 1 |
| 2013 | Discovering Semantics from Data-Centric XML
Luochen Li, Thuy Ngoc Le, Huayu Wu 0001, Tok Wang Ling, Stéphane Bressan |
DEXA (1) | 3 |
| 2013 | The Price Is Right - Models and Algorithms for Pricing DataabstractData is a modern commodity. Yet the pricing models in useon electronic data markets either focus on the usage of computing resources,or are proprietary, opaque, most likely ad hoc, and not conduciveof a healthy commodity market dynamics. In this paper we propose ageneric data pricing model that is based on minimal provenance, i.e. minimalsets of tuples contributing to the result of a query.We show that theproposed model fulfills desirable properties such as contribution monotonicity,bounded-price and contribution arbitrage-freedom. We presenta baseline algorithm to compute the exact price of a query based onour pricing model. We show that the problem is NP-hard. We thereforedevise, present and compare several heuristics. We conduct a comprehensiveexperimental study to show their effectiveness and efficiency. Ruiming Tang, Huayu Wu 0001, Zhifeng Bao, Stéphane Bressan, Patrick Valduriez |
DEXA (2) | 2 |
| 2013 | From Structure-Based to Semantics-Based: Towards Effective XML Keyword Search
Thuy Ngoc Le, Huayu Wu 0001, Tok Wang Ling, Luochen Li, Jiaheng Lu |
ER | 2 |
| 2013 | A privacy preserving framework for managing vehicle data in road pricing systemsabstractThe Electronic Road Pricing (ERP) system was implemented by the Land Transport Authority of Singapore to control traffic by road pricing since 1998. To better understand the traffic condition and improve the pricing scheme, the government initiated the next generation ERP (ERP 2) project, which aims to use the Global Navigation Satellite System (GNSS) collecting positional data from vehicles for analysis. However, most drivers fear of being monitored once the government installs the devices in their vehicles to collect GPS data. The existing data stream management systems (DSMS) centralize both data management and privacy control at server site. This framework assumes DSMS server is secure and trustable, and protects providers' data from illegal access by data users. In ERP 2, the DSMS server is maintained by the government, i.e., data user. Thus, the existing framework is not adoptable. We propose a novel framework in which privacy protection is pushed to data provider site. By doing this, the system could be safer and more efficient. Our framework can be used for the situations such as ERP 2, i.e., data providers would like to control their own privacy policies and/or the workload of DSMS server needs to be reduced. Huayu Wu 0001, Wee Siong Ng, Kian-Lee Tan, Wei Wu 0020, Shili Xiang, Mingqiang Xue |
KDD | 1 |
| 2012 | Limiting Disclosure for Data Streams in the Cloud
Wee Siong Ng, Huayu Wu 0001, Wei Wu 0020, Shili Xiang |
CLOSER | 2 |
| 2012 | A Framework for Conditioning Uncertain Relational Data
Ruiming Tang, Reynold Cheng, Huayu Wu 0001, Stéphane Bressan |
DEXA (2) | 3 |
| 2012 | Processing XML Twig Pattern Query with Wildcards
Huayu Wu 0001, Chunbin Lin, Tok Wang Ling, Jiaheng Lu |
DEXA (1) | 1 |
| 2012 | A Hybrid Approach for General XML Query Processing
Huayu Wu 0001, Ruiming Tang, Tok Wang Ling, Stéphane Bressan |
DEXA (1) | 1 |
| 2012 | Privacy Preservation in Streaming Data CollectionabstractBig data management and analysis has become a hot topic in academic and industrial research. In fact, a large portion of big data in service today are initially streaming data. To preserve the privacy of such data that are collected from data streams, the most efficient way is to control the process of data collection according to corresponding privacy polices. In this paper, we design a framework to support data stream management with privacy-preserving capabilities. In particular, we focus on two premier principles of data privacy, limited disclosure and limited collection. With these two principles guaranteed, the archived data will not necessarily be checked for privacy protection, before analysis and other operations can be done. Wee Siong Ng, Huayu Wu 0001, Wei Wu 0020, Shili Xiang, Kian-Lee Tan |
ICPADS | 2 |
| 2012 | Labeling Dynamic XML Documents: An Order-Centric ApproachabstractDynamic XML labeling schemes have important applications in XML Database Management Systems. In this paper, we explore dynamic XML labeling schemes from a novel order-centric perspective. We compare the various labeling schemes proposed in the literature with a special focus on their orders of labels. We show that the order of labels fundamentally impacts the update performance of a labeling scheme and develop an order-based framework to classify and characterize XML labeling schemes. Although there are dynamic XML labeling schemes that can completely avoid relabeling, the gain in update performance all come with considerable costs such as larger label size and lower query performance, even if the XML documents are hardly updated. We introduce vector order which is the foundation of the dynamic labeling schemes we propose. Compared with previous solutions that are based on natural order or lexicographical order, vector order is a simple, yet most effective solution to process updates in XML DBMS. We show that vector order can be gracefully applied to both range-based and prefix-based labeling schemes with little overhead introduced. Moreover, vector order-based labeling schemes are not only efficient to process, but also resilient to skewed insertions. Qualitative and experimental evaluations confirm the benefits of our approach compared to previous solutions. Tok Wang Ling, Huayu Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2011 | Edit Distance between XML and Probabilistic XML Documents
Ruiming Tang, Huayu Wu 0001, Sadegh Heyrani-Nobari, Stéphane Bressan |
DEXA (1) | 2 |
| 2011 | Object-Oriented XML Keyword Search
Huayu Wu 0001, Zhifeng Bao |
ER | 1 |
| 2010 | An Effective Object-Level XML Keyword Search
Zhifeng Bao, Jiaheng Lu, Tok Wang Ling, Huayu Wu 0001 |
DASFAA (1) | 5 |
| 2010 | Efficient Label Encoding for Range-Based Dynamic XML Labeling Schemes
Tok Wang Ling, Zhifeng Bao, Huayu Wu 0001 |
DASFAA (1) | 4 |
| 2010 | Reducing Graph Matching to Tree Matching for XML Queries with ID References
Huayu Wu 0001, Tok Wang Ling, Gillian Dobbie, Zhifeng Bao |
DEXA (2) | 1 |
| 2009 | DDE: from dewey to a fully dynamic XML labeling schemeabstractLabeling schemes lie at the core of query processing for many XML database management systems. Designing labeling schemes for dynamic XML documents is an important problem that has received a lot of research attention. Existing dynamic labeling schemes, however, often sacrifice query performance and introduce additional labeling cost to facilitate arbitrary updates even when the documents actually seldom get updated. Since the line between static and dynamic XML documents is often blurred in practice, we believe it is important to design a labeling scheme that is compact and efficient regardless of whether the documents are frequently updated or not. In this paper, we propose a novel labeling scheme called DDE (for Dynamic DEwey) which is tailored for both static and dynamic XML documents. For static documents, the labels of DDE are the same as those of dewey which yield compact size and high query performance. When updates take place, DDE can completely avoid re-labeling and its label quality is most resilient to the number and order of insertions compared to the existing approaches. In addition, we introduce Compact DDE (CDDE) which is designed to optimize the performance of DDE for insertions. Both DDE and CDDE can be incorporated into existing systems and applications that are based on dewey labeling scheme with minimum efforts. Experiment results demonstrate the benefits of our proposed labeling schemes over the previous approaches. Tok Wang Ling, Huayu Wu 0001, Zhifeng Bao |
SIGMOD Conference | 3 |
| 2009 | Performing grouping and aggregate functions in XML queriesabstractSince more and more business data are represented in XML format, there is a compelling need of supporting analytical operations in XML queries. Particularly, the latest version of XQuery proposed by W3C, XQuery 1.1, introduces a new construct to explicitly express grouping operation in FLWOR expression. Existing works in XML query processing mainly focus on physically matching query structure over XML document. Given the explicit grouping operation in a query, how to efficiently compute grouping and aggregate functions over XML document is not well studied yet. In this paper, we extend our previous XML query processing algorithm, VERT, to efficiently perform grouping and aggregate function in queries. The main technique of our approach is introducing relational tables to index values. Query pattern matching and aggregation computing are both conducted with table indices. We also propose two semantic optimizations to further improve the query performance. Finally we present experimental results to validate the efficiency of our approach, over other existing approaches. Huayu Wu 0001, Tok Wang Ling, Zhifeng Bao |
WWW | 1 |
| 2007 | VERT: A Semantic Approach for Content Search and Content Extraction in XML Query Processing
Huayu Wu 0001, Tok Wang Ling |
ER | 1 |