VLDB 2026 Research / reviewers in the wild / expert
Raymond K. Wong 0001
dblp:53/3032 · also Raymond Kwok-Kay Wong
· DBLP profile ↗
70ranked-venue papers in the field
7as first author
12since 2021 · last 2025
0000-0002-9814-6029ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 23 (5 first)Database Systems & Data Management · 20 (2 first)Data Mining & Knowledge Discovery · 18Other / Interdisciplinary · 6Knowledge Engineering, Semantic Web & Information Systems · 2Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tailored Bipartite Graph Anomaly Detection via Fixed-Attention Network
Shaowen Tang, Raymond K. Wong 0001 |
PAKDD (1) | 2 |
| 2025 | Reimagining Attentional Copulas: A Transformer-Based Approach with Proportional Dependency Learning
Yankin Chi, Raymond K. Wong 0001 |
WISE (2) | 2 |
| 2024 | Estimation-based optimizations for the semantic compression of RDF knowledge basesabstractStructured knowledge bases are critical for the interpretability of AI techniques. RDF KBs, which are the dominant representation of structured knowledge, are expanding extremely fast to increase their knowledge coverage, enhancing the capability of knowledge reasoning while bringing heavy burdens to downstream applications. Recent studies employ semantic compression to detect and remove knowledge redundancies via semantic models and use the induced model for further applications, such as knowledge completion and error detection. However, semantic models that are sufficiently expressive for semantic compression cannot be efficiently induced, especially for large-scale KBs, due to the hardness of logic induction. In this article, we present estimation-based optimizations for the semantic compression of RDF KBs from the perspectives of input and intermediate data involved in the induction of first-order logic rules. The negative sampling technique selects a representative subset of all negative tuples with respect to the closed-world assumption, reducing the cost of evaluating the quality of a logic rule used for knowledge inference. The number of logic inference operations used during a compression procedure is reduced by a statistical estimation technique that prunes logic rules of low quality. The evaluation results show that the two techniques are feasible for the purpose of semantic compression and accelerate the compression algorithm by up to 47x compared to the state-of-the-art system. • Negative sampling reduces the cost of a single logic inference. • Negative sampling speeds up semantic compression of RDF KBs by more than 2x. • Entailment estimation reduces the number of logic inferences in logic rule mining. • Entailment estimation speeds up semantic compression of RDF KBs by up to 47x. • Entailment estimation reduces up to 99% of memory cost of semantic compression. Ruoyu Wang 0004, Raymond K. Wong 0001, Daniel Sun 0004 |
Inf. Process. Manag. | 2 |
| 2023 | Symbolic Minimization on Relational DataabstractThe current wave of AI is heavily driven by data, especially for cognitive capabilities. Minimization of data semantics not only reveals core information but also becomes a guide in a wide range of domains. However, scalability is theoretically weak in pure semantic methodologies. In order to cooperate with large DBs, expressiveness is over-sacrificed in existing techniques. Thus, the quality of discovered patterns and redundancies are far from satisfactory. In this article, we formalize symbolic minimization on relational DBs and prove its NP-Completeness. A lossless technique is proposed by inducing generic first-order Horn rules that infer a subset of records from the others. More importantly, we further improve the scalability via effective caching and pruning without sacrificing the expressiveness of first-order Horn rules. A concrete system is implemented and comprehensively evaluated. Experiments show that our technique removes up to 70% contents and outperforms the state-of-the-art on minimization and scalability. The optimizations reduce up to 96% memory consumption and accelerate the performance by two orders. Our technique shows the practicality of pure semantic approaches in database mining. Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Popularity Forecasting for Emerging Research Topics at Its Early Stage of Evolution
Yankin Chi, Raymond K. Wong 0001, John Shepherd 0001 |
ADMA (1) | 2 |
| 2022 | Incorporating neighborhood features in RNNs for popularity forecasting for emerging research fieldsabstractModelling popularity for academic fields have been an ongoing study. This is especially important for emerging fields as accurate models allow efforts and resources to be efficiently utilised on promising research directions. Existing modelling methods mainly face at least one of the following three challenges: Using domain specific binary classifications on topics to be ether emerging or non-emerging leading to low generalizability. Having a biased and restricted scope of investigation due to topic terms requiring manual mining from a limited number of documents. Neglecting the effect of "cold start" when utilising a field’s historical features as inputs for popularity forecasting especially when the field is emerging and possesses limited historical data. In this paper, we build upon existing work to propose a forecasting algorithm addressing all three challenges. Firstly, we define popularity forecasting as a multivariate regression problem. Next, by combining the utilisation of existing academic databases, time specific node embeddings, and dynamic time warping, we extract concurrently trending neighbour fields whose trending pattern are similar to the field of interest. Lastly, multivariate forecasting is conducted using long short-term memory (LSTM) and dual attention recurrent neural networks (DA-RNN). Experimental results on 10 emerging and non-emerging fields of study showcases the existence and various dynamics of "cold start". Additionally, the proposed algorithm is also shown to greatly reduce RMSE, MAE, and MAPE against traditional methods for emerging fields while retaining similar performance for non-emerging fields. This validates the significance of these challenges against existing methods and provides insight on the dependency structure of emerging topics with their historical features. Yankin Chi, Raymond K. Wong 0001, Hongkuan Wang, John Shepherd 0001 |
DSAA | 2 |
| 2022 | RDF Knowledge Base Summarization by Inducing First-Order Horn Rules
Ruoyu Wang 0004, Daniel Sun 0004, Raymond K. Wong 0001 |
ECML/PKDD (2) | 3 |
| 2022 | Hybrid Variational Autoencoder for Recommender SystemsabstractE-commerce platforms heavily rely on automatic personalized recommender systems, e.g., collaborative filtering models, to improve customer experience. Some hybrid models have been proposed recently to address the deficiency of existing models. However, their performances drop significantly when the dataset is sparse. Most of the recent works failed to fully address this shortcoming. At most, some of them only tried to alleviate the problem by considering either user side or item side content information. In this article, we propose a novel recommender model called Hybrid Variational Autoencoder (HVAE) to improve the performance on sparse datasets. Different from the existing approaches, we encode both user and item information into a latent space for semantic relevance measurement. In parallel, we utilize collaborative filtering to find the implicit factors of users and items, and combine their outputs to deliver a hybrid solution. In addition, we compare the performance of Gaussian distribution and multinomial distribution in learning the representations of the textual data. Our experiment results show that HVAE is able to significantly outperform state-of-the-art models with robust performance. Hangbin Zhang, Raymond K. Wong 0001, Victor W. Chu |
ACM Trans. Knowl. Discov. Data | 2 |
| 2021 | Maximizing Explainability with SF-Lasso and Selective Inference for Video and Picture Ads
Eunkyung Park 0004, Raymond K. Wong 0001, Junbum Kwon, Victor W. Chu |
PAKDD (1) | 2 |
| 2021 | Anchoring-and-Adjustment to Improve the Quality of Significant Features
Eunkyung Park 0004, Raymond K. Wong 0001, Junbum Kwon, Victor W. Chu |
WISE (1) | 2 |
| 2021 | Online force-directed algorithms for visualization of dynamic graphs
Se-Hang Cheong, Yain-Whar Si, Raymond K. Wong 0001 |
Inf. Sci. | 3 |
| 2021 | Feature extraction for chart pattern classification in financial time series
Yuechu Zheng, Yain-Whar Si, Raymond K. Wong 0001 |
Knowl. Inf. Syst. | 3 |
| 2020 | Big Data Quality Prediction on Banking Applications: Extended AbstractabstractBig data has been transformed into knowledge by information systems to add value in businesses. Enterprises relying on it benefit from risk management to a certain extent. The value, however, depends on the quality of data. The quality needs to be verified before any use of the data. Specifically, measuring the quality by simulating the real life situation and even forecast it accurately turns into a hot topic. In recent years, there have been numerous researches on the measurement and assessment of data quality. These are yet to utilize a scientific computational method for the measurement and prediction. Current methods either fail to make an accurate prediction or do not consider the correlation and time sequence factors of the data. To address this, we design a model to extend machine learning technique to business applications predicting this. Firstly, we implement the model to detect data noises from a risk dataset according to an international data quality standard from banking industry and then estimate their impacts with Gaussian and Bayesian methods. Secondly, we direct sequential learning in multiple deep neural networks for the prediction with an attention mechanism. The model is experimented with various network methodologies to show the predictive power of machine learning technique and is evaluated by validation data to confirm the model effectiveness. The model is scalable to apply to any industries utilizing big data other than the banking industry. Ka Yee Wong, Raymond K. Wong 0001 |
DSAA | 2 |
| 2020 | Dual-PISA: An index for aggregation operations on time series data
Jialin Qiao, Xiangdong Huang 0001, Jianmin Wang 0001, Raymond K. Wong 0001 |
Inf. Syst. | 4 |
| 2019 | Multitask Learning for Sparse Failure Prediction
Simon Luo, Victor W. Chu, Zhidong Li, Yang Wang 0002, Jianlong Zhou, Fang Chen 0001, Raymond K. Wong 0001 |
PAKDD (1) | 7 |
| 2019 | Enhancing portfolio return based on sentiment-of-topic
Victor W. Chu, Raymond K. Wong 0001, Fang Chen 0001, Ivan Ho, Joe Lee |
Data Knowl. Eng. | 2 |
| 2018 | Leveraging Local Interactions for Geolocating Social Media Users
Elaheh ShafieiBavani, Raymond K. Wong 0001, Fang Chen 0001 |
PAKDD (3) | 3 |
| 2018 | Twitter user geolocation by filtering of highly mentioned usersabstractGeolocated social media data provide a powerful source of information about places and regional human behavior. Because only a small amount of social media data have been geolocation‐annotated, inference techniques play a substantial role to increase the volume of annotated data. Conventional research in this area has been based on the text content of posts from a given user or the social network of the user, with some recent crossovers between the text‐ and network‐based approaches. This paper proposes a novel approach to categorize highly‐mentioned users (celebrities) into Local and Global types, and consequently use Local celebrities as location indicators. A label propagation algorithm is then used over the refined social network for geolocation inference. Finally, we propose a hybrid approach by merging a text‐based method as a back‐off strategy into our network‐based approach. Empirical experiments over three standard Twitter benchmark data sets demonstrate that our approach outperforms state‐of‐the‐art user geolocation methods. Elaheh ShafieiBavani, Raymond K. Wong 0001, Fang Chen 0001 |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2017 | Exploring Celebrities on Inferring User Geolocation in Twitter
Elaheh ShafieiBavani, Raymond K. Wong 0001, Fang Chen 0001 |
PAKDD (1) | 3 |
| 2017 | Active Learning with Density-Initialized Decision Tree for Record MatchingabstractOne of the fundamental problem in data management and data integration fields is Record Matching, which refers to identifying records that relate to the same entities across different data sources. In recent literature, active learning has demonstrated to be effective for record matching. One of the key steps of active learning is to build a proper initial classifier, with which active learning algorithms can quickly locate informative examples for training accurate models. However, in this process, example labelling for model training is usually expensive. Even worse, if a weak initial classifier is used, the labelling cost can be significantly increased. In this paper, we propose an unsupervised algorithm to determine the initial classifier. The process of classifier initialization requires no labelling cost. Then on our proposed algorithm, we present an active sampling method for selecting informative examples. The experiments show that our approach achieves competitive learning performance with much less labelling cost than other approaches of active learning. Chenxiao Dou, Daniel Sun 0004, Guoqiang Li 0001, Raymond K. Wong 0001 |
SSDBM | 4 |
| 2016 | Interrelationships of Service Orchestrations
Victor W. Chu, Raymond K. Wong 0001, Fang Chen 0001, Chihung Chi |
ADMA | 2 |
| 2016 | Real-Time Stream Mining Electric Power Consumption Data Using Hoeffding Tree with Shadow Features
Simon Fong 0001, Meng Yuen, Raymond K. Wong 0001, Wei Song 0004, Kyungeun Cho |
ADMA | 3 |
| 2016 | Adaptive Multi-objective Swarm Crossover Optimization for Imbalanced Data Classification
Jinyan Li 0002, Simon Fong 0001, Raymond K. Wong 0001 |
ADMA | 4 |
| 2016 | PISA: An Index for Aggregating Big Time Series DataabstractAggregation operation plays an important role in time series database management. As the amount of data increases, current solutions such as summary table and MapReduce-based methods struggle to respond to such queries with low latency. Other approaches such as segment tree based methods have a poor insertion performance when the data size exceeds the available memory. This paper proposes a new segment tree based index called PISA, which has fast insertion performance and low latency for aggregation queries. PISA uses a forest to overcome the performance disadvantages of insertions in traditional segment trees. By defining two kinds of tags, namely code number and serial number, we propose an algorithm to accelerate queries by avoiding reading unnecessary data on disk. The index is stored on disk and only takes a few hundred bytes of memory for billions of data points. PISA can be easily implemented on both traditional databases and NoSQL systems, examples including MySQL and Cassandra. It handles aggregation queries within milliseconds on a commodity server for a time range that may contain tens of billions of data points. Xiangdong Huang 0001, Jianmin Wang 0001, Raymond K. Wong 0001, Chen Wang 0018 |
CIKM | 3 |
| 2016 | Effective Local Metric Learning for Water Pipe Assessment
Mojgan Ghanavati, Raymond K. Wong 0001, Fang Chen 0001, Yang Wang 0002, Simon Fong 0001 |
PAKDD (1) | 2 |
| 2016 | Unsupervised Blocking of Imbalanced Datasets for Record Matching
Chenxiao Dou, Daniel Sun 0004, Raymond K. Wong 0001 |
WISE (2) | 3 |
| 2014 | Microblog Topic Contagiousness Measurement and Emerging Outbreak MonitoringabstractA recent study on collective attention in Twitter shows that an epidemic spreading of hashtags is predominantly driven by external factors. We extend a time-series form of susceptible-infectious-recovered (SIR) model to monitor microblog emerging outbreaks by considering both endogenous and exogenous drivers. In addition, we adopt partially labeled Dirichlet allocation (PLDA) model to generate both background latent topics and hashtag topics. It overcomes the problem of small available samples in hashtag analysis by including related but unlabeled tweets through inference. We standardize hashtag topic contagiousness measure as the estimated effective-reproduction-number R derived from epidemiology. It is obtained by Bayesian parameter estimation. Guided by R, one can profile and categorize emerging topics, and generate alerts on potential outbreaks. Experiment results confirm the effectiveness of this approach. Victor W. Chu, Raymond K. Wong 0001, Fang Chen 0001, Chihung Chi |
CIKM | 2 |
| 2014 | Causal Structure Discovery for Spatio-temporal Data
Victor W. Chu, Raymond K. Wong 0001, Wei Liu 0007, Fang Chen 0001 |
DASFAA (1) | 2 |
| 2014 | CPL+: An improved approach for evaluating the local completeness of event logs
Hedong Yang, Lijie Wen 0001, Jianmin Wang 0001, Raymond K. Wong 0001 |
Inf. Process. Lett. | 4 |
| 2013 | Scalable context-aware role mining with MapReduceabstractCloud computing platforms facilitate efficiently processing complicated computing problems of which the time cost used to be unacceptable. Recent research has attempted to use role-based approaches for context-aware service recommendation, yet role mining problem has been proven to be difficult to compute. Currently proposed role-mining algorithms are inefficient and may not scale to cope with the huge amount of data in the real-world. This paper proposes a novel algorithm with much better runtime complexity, and in MapReduce style to take advantage of popular distributed computing platforms. Experiments running on a medium-sized high performance computing cluster demonstrate that our proposed algorithm works well with both running time complexity and scalability. Raymond K. Wong 0001, Chihung Chi |
IEEE BigData | 2 |
| 2012 | Full-text search on multi-byte encoded documentsabstractThe Burrows Wheeler transform (BWT) has become popular in text compression, full-text search, XML representation, and DNA sequence matching. It is very efficient to perform a full-text search on BWT encoded text using backward search. This paper aims to study different approaches for applying BWT on multi-byte encoded (e.g. UTF-16) text documents. While previous work has studied BWT on word-based models, and BWT can be applied directly on multi-byte encodings (by treating the document as single-byte coded), there has been no extensive study on how to utilize BWT on multi-byte encoded documents for efficient full-text search. Therefore, in this paper, we propose several ways to efficiently backward search multi-byte text documents. We demonstrate our findings using Chinese text documents. Our experiment results show that our extensions to the standard BWT method offer faster search performance and use less runtime memory. Raymond K. Wong 0001, Fengming Shi, Nicole Lam |
ACM Symposium on Document Engineering | 1 |
| 2010 | Structure vs. content in hierarchical corpora
Alex Penev, Raymond K. Wong 0001 |
Inf. Retr. | 2 |
| 2009 | Framework for timely and accurate ads on mobile devicesabstractWe propose a framework for mobile advertising covering Value-Added Services where an ad-database is maintained on the device and both selection and display are dictated by the device. Advantages over existing mobile marketing are that ads are more timely, viable on a variety of use cases, can be both location-sensitive and personalized with minimal privacy concerns, and provide an obvious means for subsidizing users' service costs. We construct a suitable selection algorithm and evaluate its execution, accuracy and scalability. We show that ad-serving can be done under the processing constraints imposed by mobiles, which may lead to improvements in mobile marketing effectiveness. Alex Penev, Raymond K. Wong 0001 |
CIKM | 2 |
| 2009 | Efficient Data Structure for XML Keyword Search
Ryan H. Choi, Raymond K. Wong 0001 |
DASFAA | 2 |
| 2009 | Efficient Filtering of Branch Queries for High-Performance XML Data ServicesabstractEfficient XML filtering has been the fundamental technique in recent Web service and XML publish/subscribe applications. In this article, we consider the problem of filtering a streaming XML data efficiently against a large number of branch XPath queries. To improve the performance of XML filtering, branch queries are grouped into similar queries, and the common paths between queries in the same group are identified. After performing structural matching of queries, queries are organized in a way that multiple queries can be evaluated simultaneously in the post-processing phase. In the post-processing phase, join operations are executed in a pipeline fashion, and intermediate join results are shared amongst the queries in the same group. As a result, the total number of join operations performed in the post-processing phase is significantly reduced. In addition, we also present how to efficiently return all matching elements for each matching branch query. Experiments show that our proposal is efficient and scalable compared to previous work. Ryan H. Choi, Raymond K. Wong 0001 |
J. Database Manag. | 2 |
| 2008 | Grouping Categorical AnomaliesabstractWe present an approach for discovery of groups of unusual data points that are anomalous for similar reasons. This differs from clustering in that the points that are grouped may be quite 'distant' and can use categorical attributes, and differs from anomaly detection in that we are not looking for individual outliers. Matthew Gebski, Alex Penev, Raymond K. Wong 0001 |
Web Intelligence | 3 |
| 2008 | TagScore: Approximate Similarity Using Tag SynopsesabstractCollaborative tagging is the aggregate effort by a community of online users to annotate web content with metadata labels called tags. It is a simple activity that enriches our knowledge about digital content, and has gained popularity with services such as Del.icio.us. Del.icio.us has a large repository that evolves daily, presenting interesting new problems for IR. We present TagScore, a scoring function to rate the goodness of Del.icio.us tags for their associated web page. It gives us a succinct synopsis for a page that we can use to efficiently find similar pages. Using real Del.icio.us data, we show that our approach gives good correlation to cosine similarity but is several hundred times faster and requires minimal storage overhead. Alex Penev, Raymond K. Wong 0001 |
Web Intelligence | 2 |
| 2008 | XML Storage and Processing on Mobile Devices
Raymond K. Wong 0001 |
WISE | 1 |
| 2008 | Finding similar pages in a social tagging repositoryabstractSocial tagging describes a community of users labeling web content with tags. It is a simple activity that enriches our knowledge about resources on the web. For a computer to help users search the tagged repository, it must know when tags are good or bad. We describe TagScore, a scoring function that rates the goodness of tags. The tags and their ratings give us a succinct synopsis for a page. We `find similar' pages in Del.icio.us by comparing synopses. Our approach gives good correlation to the full cosine similarity but is hundreds of times faster. Alex Penev, Raymond K. Wong 0001 |
WWW | 2 |
| 2007 | Adaptive Distance Measurement for Time Series Databases
Van Munin Chhieng, Raymond K. Wong 0001 |
DASFAA | 2 |
| 2007 | An Efficient Histogram Method for Outlier Detection
Matthew Gebski, Raymond K. Wong 0001 |
DASFAA | 2 |
| 2007 | Querying and maintaining a compact XML storageabstractAs XML database sizes grow, the amount of space used for storing the data and auxiliary data structures becomes a major factor in query and update performance. This paper presents a new storage scheme for XML data that supports all navigational operations in near constant time. In addition to supporting efficient queries, the space requirement of the proposed scheme is within a constant factor of the information theoretic minimum, while insertions and deletions can be performed in near constant time as well. As a result, the proposed structure features a small memory footprint that increases cache locality, whilst still supporting standard APIs, such as DOM, and necessary database operations, such as queries and updates, efficiently. Analysis and experiments show that the proposed structure is space and time efficient. Raymond K. Wong 0001, Franky Lam, William M. Shui |
WWW | 1 |
| 2006 | Adaptively Indexing Dynamic XML
Damien K. Fisher, Raymond K. Wong 0001 |
DASFAA | 2 |
| 2006 | Classification of Hidden Network Streams
Matthew Gebski, Alex Penev, Raymond K. Wong 0001 |
DaWaK | 3 |
| 2006 | Visual Specification and Optimization of XQuery Using VXQ
Ryan H. Choi, Raymond K. Wong 0001, Wei Wang 0011 |
DEXA | 2 |
| 2006 | Topic Distillation in Desktop Search
Alex Penev, Matthew Gebski, Raymond K. Wong 0001 |
DEXA | 3 |
| 2006 | Protocol Identification of Encrypted Network TrafficabstractNew means of communication are constantly emerging, some of which may constitute resource misuse of an organisation's network system. Identifying the protocols used is straight-forward when inspecting network logs, but we focus on the problem of identifying the underlying protocol present in an unknown TCP connection. Actions are difficult to detect if the underlying protocol is encrypted and tunneled through a proxy server or SSH. We use a graph-comparison approach to build profiles of several protocols, and attempt to classify an unknown, encrypted protocol against these profiles using only the visible behaviour of the protocol being tunneled - the size, timing and direction of packets Matthew Gebski, Alex Penev, Raymond K. Wong 0001 |
Web Intelligence | 3 |
| 2005 | Intrusion Detection via Analysis and Modelling of User Commands
Matthew Gebski, Raymond K. Wong 0001 |
DaWaK | 2 |
| 2005 | A New Approach for Cluster Detection for Large Datasets with High Dimensionality
Matthew Gebski, Raymond K. Wong 0001 |
DaWaK | 2 |
| 2004 | Algebraic Transformation and Optimization for XQuery
Damien K. Fisher, Franky Lam, Raymond K. Wong 0001 |
APWeb | 3 |
| 2004 | Skipping Strategies for Efficient Structural Joins
Franky Lam, William M. Shui, Damien K. Fisher, Raymond K. Wong 0001 |
DASFAA | 4 |
| 2004 | Effective Clustering Schemes for XML Databases
William M. Shui, Damien K. Fisher, Franky Lam, Raymond K. Wong 0001 |
DEXA | 4 |
| 2004 | Dynamic Resource Selection For Service Composition in The GridabstractWhile numerous efforts have focused on service composition in the Grid environment, service selection among similar services from multiple providers has not been addressed. In particular, all service composition work done so far are based on a given selection of services under a well set environment. As a result, uncertainty (e.g., server load, network traffic, computation time of the services due to changing memory and other unexpected conditions) under a real, dynamic environment has never been considered. This paper prototypes the service selection under a Grid environment and proposes an uncertainty framework to address the issue. Experimental results show that our considerations are valid and our preliminary solution works well in our Globus Grid network. William Kwok-Wai Cheung, Jiming Liu 0001, Kevin H. Tsang, Raymond K. Wong 0001 |
Web Intelligence | 4 |
| 2004 | Management of Web Data Models Based on Graph TransformationabstractBased on the Reserved Graph Grammar (RGG), this paper presents a unified framework to manage model-based information on the Web in a hierarchical structure. The framework allows models, schemas, and data instances to be represented explicitly and uniformly. The uniform representation of the framework also enables simple user-defined graph transformation rules for different Web data models to translate schemas and data instances between different formats. In addition, the framework implements a set of prototype tools for users to identify meta-primitives at the meta-model level, to define a model or schema by specifying a set of graph grammar rules and to draw the structure of data instances. These features promote a wide scope of Web-related applications, such as information exchange between different organizations, and integration of data coming from heterogeneous information sources. Guang-Lei Song, Kang Zhang 0001, Raymond K. Wong 0001 |
Web Intelligence | 3 |
| 2003 | An Efficient Path Index for Querying Semi-structured Data
Michael Barg, Raymond K. Wong 0001, Franky Lam |
APWeb | 2 |
| 2003 | A Fast Index for XML Document Version Management
Nicole Lam, Raymond K. Wong 0001 |
APWeb | 2 |
| 2003 | Efficient ordering for XML dataabstractWith the increasing popularity of XML, there arises the need for managing and querying information in this form. Several query languages, such as XQuery, have been proposed which return their results in document order. However, most recent efforts focused on query optimization have disregarded order. This paper presents a simple yet elegant method to maintain document ordering for XML data. Analysis of our method shows that it is indeed efficient and scalable, even for changing data. Damien K. Fisher, Franky Lam, William M. Shui, Raymond K. Wong 0001 |
CIKM | 4 |
| 2003 | Fast and Versatile ath Index for Querying Semi-Structured DataabstractThe richness of semi-structured data allows data of varied and inconsistent structures to be stored in a single database. Such data can be represented as a graph, and queries can be constructed using path expressions, which describe traversals through the graph. Instead of providing optimal performance for a limited range of path expressions, we propose a mechanism which is shown to have consistent and high performance for path expressions of any complexity, including those with descendant operators (path wildcards). We further detail mechanisms which employ our index to perform more complex processing, such as evaluating both path expressions containing links and entire (sub) queries containing path based predicates. Performance is shown to be independent of the number of terms in the path expression, even where these contain wildcards. Experiments show that our index is faster than conventional methods by up to two orders of magnitude for certain query types, is small, and scales well. Michael Barg, Raymond K. Wong 0001 |
DASFAA | 2 |
| 2003 | Integrating, Managing and Analyzing Protein Structures with XML DatabasesabstractThe development of high-throughput genome sequencing and protein structure determination techniques have provided researchers with a wealth of biological data. Integrated analysis of such data is difficult due to the disparate nature of the repositories used to store this biological data and of the software used for its analysis. This paper presents a framework based upon the use of semi-structured database management systems that would provide an integrated interface for the collection, storage and retrieval of biological data from existing repositories and of biological information generated by existing analysis programs. A simple implementation that integrates information from databases and analytical programs is presented as a proof of concept. In particular, this paper focuses on the data transformation, data integration, and the support of active rules for biological data. William M. Shui, Raymond K. Wong 0001, Stephen C. Graham, Lawrence K. Lee, W. Bret Church |
DASFAA | 2 |
| 2003 | Efficient Re-construction of Document Versions Based on Adaptive Forward and Backward Change Deltas
Raymond K. Wong 0001, Nicole Lam |
DEXA | 1 |
| 2002 | Efficient synchronization for mobile XML dataabstractMany handheld applications receive data from a primary database server and operate in an intermittently connected environment these days. They maintain data consistency with data sources through sychronization. In certain applications such as sales force automation, it is highly desirable if updates on the data source can be reflected at the handheld applications immediately. This paper proposes an efficient method to synchronize XML data on multiple mobile devices. Each device retrieves and caches a local copy of data from the database source based on a regular path expression. These local copies may be overlapping or disjoint with each other. An efficient mechanism is proposed to find all the disjoint copies to avoid unnecessary synchronizations. Each update to the data source will then be checked to identify all handheld applications which are affected by the update. Communication costs can be further reduced by eliminating the forwarding of unnecessary operations to groups of mobile clients. Franky Lam, Nicole Lam, Raymond K. Wong 0001 |
CIKM | 3 |
| 2002 | Inter-system Triggering for Heterogeneous Database Systems
Walter Guan, Raymond K. Wong 0001 |
DEXA | 2 |
| 2002 | Managing and querying multi-version XML data with update loggingabstractWith the increasing popularity of storing content on the WWW and intranet in XML form, there arises the need for the control and management of this data. As this data is constantly evolving, users want to be able to query previous versions, query changes in documents, as well as to retrieve a particular document version efficiently. This paper proposes a version management system for XML data that can manage and query changes in an effective and meaningful manner. Raymond K. Wong 0001, Nicole Lam |
ACM Symposium on Document Engineering | 1 |
| 2001 | Structural Proximity Searching for Large Collections of Semi-Structured DataabstractThe richness of the XML data format allows data to be structured in a way which precisely captures the semantics required by the author. It is the structure of the data, however, which forms the basis of all XML query languages. Without at least some notion of the structure, a user cannot meaningfully query the data. This problem is compounded when one considers that heterogeneous data adhering to different schema are likely to exist in the database(s) being queried. This paper proposes a solution based on an efficient proximity index. In particular, we describe a family of encoding and compression schemes which enable us to build an index to efficiently implement the proximity search. Our index is extremely small, and can reflect updates in the underlying database in modest time. Experiments show that our algorithm and implementation are fast and scale well. Michael Barg, Raymond K. Wong 0001 |
CIKM | 2 |
| 2001 | Structural Inference for Semistructured DataabstractA08- "+\t-'(+-(+> -+3CB@D E\t"!F GH( 8- 28530 - %#J+HI$-(+\tK-%0L7\t"-+?\t->+\t-%\t%$&$, J+08M 28530 - \t$O&0P-%%\t-(!08-QR+$2+%\t35SL 'TUEVDW708N 7W0P-\t)+:X.++-("Y\t"W 7"+K'7%\t-(E-(F+$& F ;%$\t&-Z-%0#\tHIY&0P[\t"+- 3J\\]Z-('^ -+) :XMH:_#^R\t#! $,\t.$&%\t+'E\t-[\t"M %"@ #% -%06'%\t&$L0P+3`B@\t"=(%+\tEJ7\t"-A08- H(%$&%\t'\t" $,\t-%0%A0P+[- M0a-(b-(N \t\t-(#%$1$&+'(+& $,\t-%0%A0P+[- M0a-(b-(N 29000-3 %$'(-,\t"+)++$#++'.Ff-\t"7:Z\t"+ag+h -++MF M+ %#\t"-( K08-(i+H-(+:5-/f3MjX-("+ HI6RfN \t\t-(EHI$<-(+k+:l" F+&\t"-m)F Yf-(\t$, HI$-f?-+\t%\t-(2\t"+ag+)\t-WFfn\t"h-( \t$, 28810 1. Jason Sankey, Raymond K. Wong 0001 |
CIKM | 2 |
| 2001 | Modeling and Manipulating Multidimensional Data in Semistructured DatabasesabstractMultidimensional information is pervasive in many computer applications including time series, spatial information, data warehousing, and visual data. While semistructured data or XML is becoming more and more popular for information integration and exchange, not much research work has been done in the design and implementation of semistructured database systems to manage multidimensional information efficiently. In this paper, dimension operators have been defined based on a multidimensional logic which we call ML(/spl omega/). It can be used in applications such as multidimensional spreadsheets and multidimensional databases usually found in decision support systems and data warehouses. Finally a multidimensional, an XML database system is prototyped and described in detail. Technologies such as XSL are used to transform or visualise data from different dimensions. Franky Lam, Raymond K. Wong 0001, Mehmet A. Orgun |
DASFAA | 2 |
| 2001 | The extended XQL for querying and updating large XML databasesabstractXQL has been argued as just a model for asking for specific sets of elements with very limited query capability. This paper proposes several extensions of XQL to address the issues. The extensions include full-text indexed search, path variables, joins, session-based navigations, and updates. Effort has been spent to preserve the conciseness of the language syntax. Its corresponding query processor with optimization mechanism has been prototyped and available online. Finally, implementation issues are discussed. Raymond K. Wong 0001 |
ACM Symposium on Document Engineering | 1 |
| 2000 | A Role-Based Access Control Model for XML RepositoriesabstractXML has recently attracted many software vendors, business organizations and research institutions over the world as a promising way to represent Web information. The availability of large amounts of XML data on the Web raises several issues that the XML standard does not address. In particular, access control and authorization issues have seldom been addressed. Since XML are usually stored in multiple sources or repositories, discretionary access control is desirable, if it is not necessary. In this paper, we propose a role-based access control (RBAC) model for XML repositories. The resultant RBAC scheme is in XML and heavily relies on XPath. Raymond K. Wong 0001 |
WISE | 2 |
| 1999 | Multifaceted Object Modeling with Roles: A Comprehensive Approach
Qing Li 0001, Raymond K. Wong 0001 |
Inf. Sci. | 2 |
| 1997 | A Data Model and Semantics of Objects with Dynamic RolesabstractAlthough the concept of roles is becoming a popular research issue in object oriented databases and has been proven to be useful for dynamic and evolving applications, it has only been described conceptually in most of the previous work. Moreover, the important issues such as the semantics of roles (e.g., message passing) are seldom discussed. Furthermore, none of the previous work has investigated the idea of role player qualification, which models the fact that not every object is qualified to play a particular role. We present a data model and the semantics of roles. We discuss each of the above issues and illustrate the ideas with examples. From these examples, we can easily see that the problems we discussed are fundamental and indeed exist in many complex applications. Raymond K. Wong 0001, H. Lewis Chau, Frederick H. Lochovsky |
ICDE | 1 |