Raymond Y. K. Lau

dblp:l/RaymondYKLau · also Raymond Lau 0001, Raymond Yiu-Keung Lau · DBLP profile ↗
← Back
99ranked-venue papers
27as first author
14since 2021 · last 2026
0000-0002-5751-4550ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 16 first-author · 2 since 2021Databases, data management, data science and information retrieval · 38 · 10 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 3 since 2021Computer networks · 3 · 2 first-authorTheory of computation · 2 · 1 first-author
YearPublicationVenuePosition
2026 Beyond interlocks: Board networks and corporate social responsibility performance similarity
Qin Su, Xin Bao, Raymond Y. K. Lau
Decis. Support Syst.5
2024 QTU-Net: Quaternion Transformer-Based U-Net for Water Body Extraction of RGB Satellite Image
abstract
Deep learning models have achieved great success in water body extraction (WBE) from remote sensing images. However, the existing deep learning-based extraction methods exhibit limitations in their ability to fully explore the intricate interconnections inherent in RGB color satellite imagery and to enhance semantic representation across diverse regions. Furthermore, these methods often struggle with challenges posed by the uneven distribution of water bodies at different scales within the image, as well as substantial color disparities between water and land areas. In this article, we tackle WBE task from quaternion domain and introduce a novel approach called quaternion transformer-based U-Net (QTU-Net) to address these challenges. Our method specifically leverages quaternion convolution operations to capture the holistic relationships among RGB channels, thereby enhancing the semantic representation of WBE. Additionally, we propose a quaternion initialization module (QIM) to determine optimal RGB weights and facilitate the generation of quaternion data. To further improve the accuracy of water body delineation, we incorporate an innovative multiscale similarity aggregation attention (MSAA) component that enhances local similarity capture across various scales. Finally, we evaluate the proposed QTU-Net based on three publicly available benchmark datasets. The experimental results demonstrate that the proposed QTU-Net outperforms state-of-the-art baseline methods.
Chunshan Li, Xiaofei Yang 0002, Zhiquan Zhou 0002, Raymond Y. K. Lau
IEEE Trans. Geosci. Remote. Sens.6
2023 Copula Guided Parallel Gibbs Sampling for Nonparametric and Coherent Topic Discovery (Extended Abstract)
abstract
In terms of the generative process, the Gamma-Gamma-Poisson Process (G2PP) is equivalent to the nonparametric topic model of Hierarchical Dirichlet Process (HDP). Considering the high computational cost of estimating parameters in HDP, a parallel G2PP was developed to generate topics efficiently via multi-threading. Unfortunately, the above model needs to predefine the number of topics. To address this issue, we first propose a Topic Self-Adaptive Model (TSAM) for nonparametric and parallel topic discovery. In TSAM, a monitor-executor mechanism is developed to manage the global topic information using a hierarchical structure of threads. Based on the apparatus of copulas, we further extend our TSAM to TSAMcop for coherent topic modeling by exploiting a copula guided parallel Gibbs sampling algorithm. Extensive experiments validate the effectiveness of both TSAM and TSAMcop.
Lihui Lin, Yanghui Rao, Haoran Xie 0001, Raymond Y. K. Lau, Jian Yin 0001, Fu Lee Wang, Qing Li 0001
ICDE4
2023 Leveraging Emotional Features and Machine Learning for Predicting Startup Funding Success
abstract
Analyzing the crucial factors which help predict startups' funding amounts is important for senior executives of these firms to formulate effective business strategies, which leads to the ultimate startup success. In this research, we crawled real-world startup funding data from the well-known “AngelList” platform which disseminates information about the company profiles of startups, the specific business sectors, and potential investors. Potential investors browse the information posted on AngelList, which may in turn influence their decisions in funding certain startups. Our work aims to evaluate the bundle of factors (e.g., sentiments and emotions embedded in startup profile descriptions, startups' fundamentals, etc.) that may influence startups' funding successes. Moreover, we have examined a variety of state-of-the-art machine learning-based prediction models. In particular, we applied TextCNN, a well-known deep learning method to extract sentimental and emotional features from company profile text to enhance the startup funding prediction task. Our experimental results show that the emotion feature can significantly boost startup funding prediction performance by 12% in terms of F-score, and it is also among the key factors that influence startup funding amounts. To our best knowledge, this work represents the first successful research on examining the relationship emotions captured in company profile text and startup funding success.
Raymond Y. K. Lau
TENCON2
2023 Emotion-regulatory chatbots for enhancing consumer servicing: An interpersonal emotion management approach
Bei Luo, Raymond Y. K. Lau, Chunping Li
Inf. Manag.2
2023 DeepEmotionNet: Emotion mining for corporate performance analysis and prediction
Qiping Wang 0002, Tingxuan Su, Raymond Y. K. Lau, Haoran Xie 0001
Inf. Process. Manag.3
2023 Blockchain-Enhanced Smart Contract for Cost-Effective Insurance Claims Processing
abstract
Blockchain-enabled smart contracts have revolutionized the insurance industry due to their potential to streamline backend operations, mitigate fraudulent claims, and enhance data security and transparency. Guided by the design science methodology, the authors propose two specific smart contract frameworks to enhance insurance claims processing related to vehicle damage claims and personal injury claims. These proposed frameworks can improve the overall efficiency and effectiveness of insurance claims processing by automating claims submission, review, analysis, and payment, while reducing fraud and data leakage, by merging various data sources and disintermediation. Furthermore, the authors design a smart contract template supported by eight operational algorithms to facilitate the processing of insurance claims with the help of smart contracts. This template provides practitioners with a standardized prototype for the development of secure and efficient insurance applications.
Qiping Wang 0002, Raymond Y. K. Lau, Yain-Whar Si, Haoran Xie 0001, Xiaohui Tao 0001
J. Glob. Inf. Manag.2
2023 Semi-Supervised Sentiment Classification and Emotion Distribution Learning Across Domains
abstract
In this study, sentiment classification and emotion distribution learning across domains are both formulated as a semi-supervised domain adaptation problem, which utilizes a small amount of labeled documents in the target domain for model training. By introducing a shared matrix that captures the stable association between document clusters and word clusters, non-negative matrix tri-factorization (NMTF) is robust to the labeled target domain data and has shown remarkable performance in cross-domain text classification. However, the existing NMTF-based models ignore the incompatible relationship of sentiment polarities and the relatedness among emotions. Besides, their applications on large-scale datasets are limited by the high computation complexity. To address these issues, we propose a semi-supervised NMTF framework for sentiment classification and emotion distribution learning across domains. Based on a many-to-many mapping between document clusters and sentiment polarities (or emotions), we first incorporate the prior information of label dependency to improve the model performance. Then, we develop a parallel algorithm based on message passing interface (MPI) to further enhance the model scalability. Extensive experiments on real-world datasets validate the effectiveness of our method.
Yufu Chen, Yanghui Rao, Shurui Chen, Zhiqi Lei, Haoran Xie 0001, Raymond Y. K. Lau, Jian Yin 0001
ACM Trans. Knowl. Discov. Data6
2023 Granularity-Aware Area Prototypical Network With Bimargin Loss for Few Shot Relation Classification
abstract
Relation Classification is one of the most important tasks in text mining. Previous methods either require large-scale manually-annotated data or rely on distant supervision approaches which suffer from the long-tail problem. To reduce the expensive manually-annotating cost and solve the long-tail problem, prototypical networks are widely used in few-shot RC tasks. Despite their remarkable performance, current prototypical networks ignore the different granularities of relations, which degrades the classification performance dramatically. Moreover, the optimization of current prototypical networks simply relies on the cross-entropy loss, which cannot consider the intra-relation compactness and the dispersion among relations in a semantic space. It is not robust enough for current prototypical network in real-world and complicated scenarios. In this paper, we propose an area prototypical network with a granularity-aware measurement, aiming to considering the different granularities of relations. Each relation is represented as an area whose width can reflect the granularity level of relation. Moreover, to improve the robustness, bimargin loss is designed to force area prototypical network to improve the intra-relation compactness and inter-relation dispersion for the feature representation in a semantic space. Extensive experiments on two public datasets are conducted and evaluate the effectiveness of our proposed model.
Haopeng Ren, Yi Cai 0001, Raymond Y. K. Lau, Ho-fung Leung, Qing Li 0001
IEEE Trans. Knowl. Data Eng.3
2022 PsyCredit: An interpretable deep learning-based credit assessment approach facilitated by psychometric natural language processing
Kai Yang 0007, Raymond Y. K. Lau
Expert Syst. Appl.3
2022 A deep data augmentation framework based on generative adversarial networks
Qiping Wang 0002, Haoran Xie 0001, Yanghui Rao, Raymond Y. K. Lau, Detian Zhang
Multim. Tools Appl.5
2022 Copula Guided Parallel Gibbs Sampling for Nonparametric and Coherent Topic Discovery
abstract
Hierarchical Dirichlet Process (HDP) has attracted much attention in the research community of natural language processing. Given a corpus, HDP is able to determine the number of topics automatically, possessing an important feature dubbed nonparametric that overcomes the challenging issue of manually specifying a suitable topic number in parametric topic models, such as Latent Dirichlet Allocation (LDA). Nevertheless, HDP requires a much higher computational cost than LDA for parameter estimation. By taking the advantage of multi-threading, a parallel Gibbs sampling algorithm is proposed to estimate parameters for HDP based on the equivalence between HDP and Gamma-Gamma Poisson Process (G2PP) in terms of the generative process. Unfortunately, the above parallel Gibbs sampling algorithm requires to apply the finite approximation on the number of topics manually (i.e., predefine the topic number), thus can not retain the nonparametric feature of HDP. Another drawback of the above models is the lack of capturing the semantic dependencies between words, because the topic assignment of words is independent with each other. Although some works have been done in phrase-based topic modelling, these existing methods are still limited by either enforcing the entire phrase to share a common topic or requiring much complex and time-consuming phrase mining methods. In this paper, we aim to develop a copula guided parallel Gibbs sampling algorithm for HDP which can adjust the number of topics dynamically and capture the latent semantic dependencies between words that compose a coherent segment. Extensive experiments on real-world datasets indicate that our method achieves low perplexities and high topic coherence scores with a small time cost. In addition, we validate the effectiveness of our method on the modelling of word semantic dependencies by comparing the extracted topical phrases with those learned by state-of-the-art phrase-based baselines.
Lihui Lin, Yanghui Rao, Haoran Xie 0001, Raymond Y. K. Lau, Jian Yin 0001, Fu Lee Wang, Qing Li 0001
IEEE Trans. Knowl. Data Eng.4
2022 Leveraging statistical information in fine-grained financial sentiment analysis
Han Zhang 0043, Zongxi Li, Haoran Xie 0001, Raymond Y. K. Lau, Gary Cheng 0001, Qing Li 0001, Dian Zhang 0001
World Wide Web4
2021 On entropy-based term weighting schemes for text categorization
Tao Wang 0036, Yi Cai 0001, Ho-fung Leung, Raymond Y. K. Lau, Haoran Xie 0001, Qing Li 0001
Knowl. Inf. Syst.4
2020 Does the interplay between the personality traits of CEOs and CFOs influence corporate mergers and acquisitions intensity? An econometric analysis with machine learning-based constructs
Qiping Wang 0002, Raymond Y. K. Lau, Kai Yang 0007
Decis. Support Syst.2
2020 Healthcare informatics and analytics in big data
Md. Ileas Pramanik, Raymond Y. K. Lau, Md. Abul Kalam Azad 0004, Md. Sakir Hossain, Md. Kamal Hossain Chowdhury, B. K. Karmaker
Expert Syst. Appl.2
2019 ITWF: A framework to apply term weighting schemes in topic model
Kai Yang 0007, Yi Cai 0001, Ho-fung Leung, Raymond Y. K. Lau, Qing Li 0001
Neurocomputing4
2019 Utility-based feature selection for text classification
Heyong Wang, Ming Hong, Raymond Y. K. Lau
Knowl. Inf. Syst.3
2019 On the Effectiveness of Least Squares Generative Adversarial Networks
abstract
Unsupervised learning with generative adversarial networks (GANs) has proven to be hugely successful. Regular GANs hypothesize the discriminator as a classifier with the sigmoid cross entropy loss function. However, we found that this loss function may lead to the vanishing gradients problem during the learning process. To overcome such a problem, we propose in this paper the Least Squares Generative Adversarial Networks (LSGANs) which adopt the least squares loss for both the discriminator and the generator. We show that minimizing the objective function of LSGAN yields minimizing the Pearson$\chi ^2$divergence. We also show that the derived objective function that yields minimizing the Pearson$\chi ^2$divergence performs better than the classical one of using least squares for classification. There are two benefits of LSGANs over regular GANs. First, LSGANs are able to generate higher quality images than regular GANs. Second, LSGANs perform more stably during the learning process. For evaluating the image quality, we conduct both qualitative and quantitative experiments, and the experimental results show that LSGANs can generate higher quality images than regular GANs. Furthermore, we evaluate the stability of LSGANs in two groups. One is to compare between LSGANs and regular GANs without gradient penalty. We conduct three experiments, including Gaussian mixture distribution, difficult architectures, and a newly proposed method — datasets with small variability, to illustrate the stability of LSGANs. The other one is to compare between LSGANs with gradient penalty (LSGANs-GP) and WGANs with gradient penalty (WGANs-GP). The experimental results show that LSGANs-GP succeed in training for all the difficult architectures used in WGANs-GP, including 101-layer ResNet.
Xudong Mao, Qing Li 0001, Haoran Xie 0001, Raymond Y. K. Lau, Zhen Wang 0004, Stephen Paul Smolley
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Road Detection and Centerline Extraction Via Deep Recurrent Convolutional Neural Network U-Net
abstract
Road information extraction based on aerial images is a critical task for many applications, and it has attracted considerable attention from researchers in the field of remote sensing. The problem is mainly composed of two subtasks, namely, road detection and centerline extraction. Most of the previous studies rely on multistage-based learning methods to solve the problem. However, these approaches may suffer from the well-known problem of propagation errors. In this paper, we propose a novel deep learning model, recurrent convolution neural network U-Net (RCNN-UNet), to tackle the aforementioned problem. Our proposed RCNN-UNet has three distinct advantages. First, the end-to-end deep learning scheme eliminates the propagation errors. Second, a carefully designed RCNN unit is leveraged to build our deep learning architecture, which can better exploit the spatial context and the rich low-level visual features. Thereby, it alleviates the detection problems caused by noises, occlusions, and complex backgrounds of roads. Third, as the tasks of road detection and centerline extraction are strongly correlated, a multitask learning scheme is designed so that two predictors can be simultaneously trained to improve both effectiveness and efficiency. Extensive experiments were carried out based on two publicly available benchmark data sets, and nine state-of-the-art baselines were used in a comparative evaluation. Our experimental results demonstrate the superiority of the proposed RCNN-UNet model for both the road detection and the centerline extraction tasks.
Xiaofei Yang 0002, Xutao Li 0003, Yunming Ye, Raymond Y. K. Lau, Xiaofeng Zhang 0002, Xiaohui Huang 0003
IEEE Trans. Geosci. Remote. Sens.4
2018 Learning Dual Preferences with Non-negative Matrix Tri-Factorization for Top-N Recommender System
Xiangsheng Li, Yanghui Rao, Haoran Xie 0001, Yufu Chen, Raymond Y. K. Lau, Fu Lee Wang, Jian Yin 0001
DASFAA (1)5
2018 Enhancing Binary Classification by Modeling Uncertain Boundary in Three-Way Decisions (Extended Abstract)
abstract
Text classification techniques are playing a crucial role in identifying relevant texts from a large data set, e.g., various online crimes such as Cyberbullying, terrorist recruiting, propaganda or attack planning. Until now, supervised deep learning has brought about breakthroughs in processing multimedia data; however, there was no good practical way to harvest this opportunity for text classification because acquiring and maintaining a massive amount of training examples are too expensive for a large number of categories (e.g., Yahoo! taxonomy contains nearly 300,000 categories and the Library of Congress Subject Headings (LCSH) contains 394,070 subjects). Therefore, the question of how to effectively learn from sparse or small set of training examples is crucial for the true success of text classification. Semi-supervised approaches have been proposed for this challenge, which usually use a pair or several existing classifiers to extend a small training set. However, extracted pseudo training samples are uncertain because they are determined by a machine rather than people. Also, the massive volume and high variability of text data are creating a number of challenging issues such as the scalability and complicated relations between words. There are two fundamental issues with regards to the performance of existing classifiers: overlook and overload. Overlook means that some objects relevant to a class have been omitted, whereas overload means that some objects assigned to a class are actually not relevant to that class. The two issues are even more serious in the following two cases: (1) large uncertain boundary - the decision boundary between two classes includes many mixed examples (e.g., relevant and nonrelevant documents together), and (2) unbalanced classes - one class (e.g., information about terrorist attacks) is much smaller than another class (e.g., normal descriptions). We propose a three-way decision model [1] for dealing with the uncertain boundary for improving text classification performance based on rough set techniques and centroid solution. It aims to understand the uncertain boundary through partitioning the training samples into three regions (the positive, boundary and negative regions) by two main boundary vectors created from the labeled positive and negative training subsets, respectively, and further resolve the objects in the boundary region by two derived boundary vectors produced according to the structure of the boundary region. Four decision rules are proposed from the training process and applied to the incoming documents for more precise classification. The experimental results on the standard data sets RCV1 and Reuters-21578 show that the usage of boundary vectors is very effective and efficient for dealing with uncertainties of the decision boundary, and the proposed model has significantly improved the performance of binary text classification in terms of F1 measure and AUC area compared with six other popular baseline models.
Yuefeng Li 0001, Libiao Zhang, Yue Xu 0001, Yiyu Yao, Raymond Y. K. Lau, Yutong Wu 0001
ICDE5
2018 Universal affective model for Readers' emotion classification over short texts
Weiming Liang, Haoran Xie 0001, Yanghui Rao, Raymond Y. K. Lau, Fu Lee Wang
Expert Syst. Appl.4
2018 Hyperspectral Image Classification With Deep Learning Models
abstract
Deep learning has achieved great successes in conventional computer vision tasks. In this paper, we exploit deep learning techniques to address the hyperspectral image classification problem. In contrast to conventional computer vision tasks that only examine the spatial context, our proposed method can exploit both spatial context and spectral correlation to enhance hyperspectral image classification. In particular, we advocate four new deep learning models, namely, 2-D convolutional neural network (2-D-CNN), 3-D-CNN, recurrent 2-D CNN (R-2-D-CNN), and recurrent 3-D-CNN (R-3-D-CNN) for hyperspectral image classification. We conducted rigorous experiments based on six publicly available data sets. Through a comparative evaluation with other state-of-the-art methods, our experimental results confirm the superiority of the proposed deep learning models, especially the R-3-D-CNN and the R-2-D-CNN deep learning models.
Xiaofei Yang 0002, Yunming Ye, Xutao Li 0003, Raymond Y. K. Lau, Xiaofeng Zhang 0002, Xiaohui Huang 0003
IEEE Trans. Geosci. Remote. Sens.4
2018 Automatic Approach of Sentiment Lexicon Generation for Mobile Shopping Reviews
abstract
The dramatic increase in the use of smartphones has allowed people to comment on various products at any time. The analysis of the sentiment of users’ product reviews largely depends on the quality of sentiment lexicons. Thus, the generation of high‐quality sentiment lexicons is a critical topic. In this paper, we propose an automatic approach for constructing a domain‐specific sentiment lexicon by considering the relationship between sentiment words and product features in mobile shopping reviews. The approach first selects sentiment words and product features from original reviews and mines the relationship between them using an improved pointwise mutual information algorithm. Second, sentiment words that are related to mobile shopping are clustered into categories to form sentiment dimensions. At each sentiment dimension, each sentiment word can take the value of 0 or 1, where 1 indicates that the word belongs to a particular category whereas 0 indicates that it does not belong to that category. The generated lexicon is evaluated by constructing a sentiment classification task using several product reviews written in both Chinese and English. Two popular non‐domain‐specific sentiment lexicons as well as state‐of‐the‐art machine‐learning and deep‐learning models are chosen as benchmarks, and the experimental results show that our sentiment lexicons outperform the benchmarks with statistically significant differences, thus proving the effectiveness of the proposed approach.
Jun Feng 0001, Xiaodong Li 0007, Raymond Y. K. Lau
Wirel. Commun. Mob. Comput.4
2017 Least Squares Generative Adversarial Networks
abstract
Unsupervised learning with generative adversarial networks (GANs) has proven hugely successful. Regular GANs hypothesize the discriminator as a classifier with the sigmoid cross entropy loss function. However, we found that this loss function may lead to the vanishing gradients problem during the learning process. To overcome such a problem, we propose in this paper the Least Squares Generative Adversarial Networks (LSGANs) which adopt the least squares loss function for the discriminator. We show that minimizing the objective function of LSGAN yields minimizing the Pearson X2 divergence. There are two benefits of LSGANs over regular GANs. First, LSGANs are able to generate higher quality images than regular GANs. Second, LSGANs perform more stable during the learning process. We evaluate LSGANs on LSUN and CIFAR-10 datasets and the experimental results show that the images generated by LSGANs are of better quality than the ones generated by regular GANs. We also conduct two comparison experiments between LSGANs and regular GANs to illustrate the stability of LSGANs.
Xudong Mao, Qing Li 0001, Haoran Xie 0001, Raymond Y. K. Lau, Zhen Wang 0004, Stephen Paul Smolley
ICCV4
2017 Toward a social sensor based framework for intelligent transportation
abstract
Intelligent Transportation Systems (ITSs) are supposed to constantly provide drivers with useful navigation information (e.g., the shortest driving paths, real-time traffic conditions, and other road condition data) by collecting raw data from physical sensors and various media. Previous work mainly considers how to develop ITS frameworks and how to extract useful information from physical sensors. Meanwhile, the relatively rich user-contributed postings which contain valuable information about traffic and road conditions was not leveraged to enhance the effectiveness of ITSs. To address such a research gap, this paper presents a relatively new approach of extracting and analyzing useful driving navigation information from the big data archived on online social media, the so-called social sensors. In particular, we advocate a topic model-based computational method to extract relevant semantics (e.g., traffic conditions, road conditions, drivers' conditions, etc.) from the relatively noisy user-contributed postings on online social media. Then, a classification ensemble method is proposed to automatically identify specific traffic related events, and transmit these useful navigation hints back to an ITS for dissemination to other drivers. Evaluation based on real-world user postings from Twitter and Sina Weibo demonstrates that the proposed social sensor analytics method is effective in identifying main traffic events when compared to other baseline methods. Our work opens the door for the development of the next generation of ITSs which can leverage the big data archived on online social media to enhance the richness and the quality of traffic navigation aids provided by ITSs.
Raymond Y. K. Lau
WoWMoM1
2017 Relationship Identification Across Heterogeneous Online Social Networks
abstract
In the era of the social web, many people manage their social relationships through various online social networking services. It has been found that identifying the types of social relationships among users in online social networks facilitates the marketing of products via electronic “word of mouth.” However, it is a great challenge to identify the types of social relationships, given very limited information in a social network. In this article, we study how to identify the types of relationships across multiple heterogeneous social networks and examine if combining certain information from different social networks can help improve the identification accuracy. The main contribution of our research is that we develop a novel decision tree initiated random walk model, which takes into account both global network structure and local user behavior to bootstrap the performance of relationship identification. Experiments conducted based on two real‐world social networks, Sina Weibo and Jiepang, demonstrate that the proposed model achieves an average accuracy of 92.0%, significantly outperforming other baseline methods. Our experiments also confirm the effectiveness of combining information from multiple social networks. Moreover, our results reveal that human mobility features indicating location categories, coincidence, and check‐in patterns are among the most discriminative features for relationship identification.
Jiangning He, Hongyan Liu 0002, Raymond Y. K. Lau, Jun He 0008
Comput. Intell.3
2017 Smart health: Big data enabled health paradigm within smart cities
Md. Ileas Pramanik, Raymond Y. K. Lau, Haluk Demirkan, Md. Abul Kalam Azad 0004
Expert Syst. Appl.2
2017 Multi-view learning via multiple graph regularized generative model
Shaokai Wang, Ke Wang 0068, Xutao Li 0003, Yunming Ye, Raymond Y. K. Lau, Xiaolin Du
Knowl. Based Syst.5
2017 Semi-supervised Collective Classification in Multi-attribute Network Data
Shaokai Wang, Yunming Ye, Xutao Li 0003, Xiaohui Huang 0003, Raymond Y. K. Lau
Neural Process. Lett.5
2017 Bootstrapping Social Emotion Classification with Semantically Rich Hybrid Neural Networks
abstract
Social emotion classification aims to predict the aggregation of emotional responses embedded in online comments contributed by various users. Such a task is inherently challenging because extracting relevant semantics from free texts is a classical research problem. Moreover, online comments are typically characterized by a sparse feature space, which makes the corresponding emotion classification task very difficult. On the other hand, though deep neural networks have been shown to be effective for speech recognition and image analysis tasks because of their capabilities of transforming sparse low-level features to dense high-level features, their effectiveness on emotion classification requires further investigation. The main contribution of our work reported in this paper is the development of a novel model of semantically rich hybrid neural network (HNN) which leverages unsupervised teaching models to incorporate semantic domain knowledge into the neural network to bootstrap its inference power and interpretability. To our best knowledge, this is the first successful work of incorporating semantics into neural networks to enhance social emotion classification and network interpretability. Through empirical studies based on three real-world social media datasets, our experimental results confirm that the proposed hybrid neural networks outperform other state-of-the-art emotion classification methods.
Xiangsheng Li, Yanghui Rao, Haoran Xie 0001, Raymond Y. K. Lau, Jian Yin 0001, Fu Lee Wang
IEEE Trans. Affect. Comput.4
2017 Finding Semantically Valid and Relevant Topics by Association-Based Topic Selection Model
abstract
Topic modelling methods such as Latent Dirichlet Allocation (LDA) have been successfully applied to various fields, since these methods can effectively characterize document collections by using a mixture of semantically rich topics. So far, many models have been proposed. However, the existing models typically outperform on full analysis on the whole collection to find all topics but difficult to capture coherent and specifically meaningful topic representations. Furthermore, it is very challenging to incorporate user preferences into existing topic modelling methods to extract relevant topics. To address these problems, we develop a novel personalized Association-based Topic Selection (ATS) model, which can identify semantically valid and relevant topics from a set of raw topics based on the semantical relatedness between users’ preferences and the structured patterns captured in topics. The advantage of the proposed ATS model is that it enables an interactive topic modelling process driven by users’ specific interests. Based on three benchmark datasets, namely, RCV1, R8, and WT10G under the context of information filtering (IF) and information retrieval (IR), our rigorous experiments show that the proposed ATS model can effectively identify relevant topics with respect to users’ specific interests, and hence to improve the performance of IF and IR.
Yang Gao 0016, Yuefeng Li 0001, Raymond Y. K. Lau, Yue Xu 0001, Md. Abul Bashar
ACM Trans. Intell. Syst. Technol.3
2017 Enhancing Binary Classification by Modeling Uncertain Boundary in Three-Way Decisions
abstract
Text classification is a process of classifying documents into predefined categories through different classifiers learned from labelled or unlabelled training samples. Many researchers who work on binary text classification attempt to find a more effective way to separate relevant texts from a large data set. However, current text classifiers cannot unambiguously describe the decision boundary between positive and negative objects because of uncertainties caused by text feature selection and the knowledge learning process. This paper proposes a three-way decision model for dealing with the uncertain boundary to improve the binary text classification performance based on therough settechniques and centroid solution. It aims to understand the uncertain boundary through partitioning the training samples into three regions (the positive, boundary, and negative regions) by two main boundary vectors$\vec{C_{P}}$and$\vec{C_{N}}$, created from the labeled positive and negative training subsets, respectively, and further resolve the objects in the boundary region by two derived boundary vectors$\vec{B_{P}}$and$\vec{B_{N}}$, produced according to the structure of the boundary region. It involves an indirect strategy which is composed of two successive steps in the whole classification process: ‘two-way to three-way’ and ‘three-way to two-way’. Four decision rules are proposed from the training process and applied to the incoming documents for more precise classification. A large number of experiments have been conducted based on the standard data sets RCV1 and Reuters-21578. The experimental results show that the usage of boundary vectors is very effective and efficient for dealing with uncertainties of the decision boundary, and the proposed model has significantly improved the performance of binary text classification in terms of$F_{1}$measure and$AUC$area compared with six other popular baseline models.
Yuefeng Li 0001, Libiao Zhang, Yue Xu 0001, Yiyu Yao, Raymond Y. K. Lau, Yutong Wu 0001
IEEE Trans. Knowl. Data Eng.5
2016 Exploring Topic Discriminating Power of Words in Latent Dirichlet Allocation
abstract
Latent Dirichlet Allocation (LDA) and its variants have been widely used to discover latent topics in textual documents. However, some of topics generated by LDA may be noisy with irrelevant words scattering across these topics. We name this kind of words as topic-indiscriminate words, which tend to make topics more ambiguous and less interpretable by humans. In our work, we propose a new topic model named TWLDA, which assigns low weights to words with low topic discriminating power (ability). Our experimental results show that the proposed approach, which effectively reduces the number of topic-indiscriminate words in discovered topics, improves the effectiveness of LDA.
Kai Yang 0007, Yi Cai 0001, Ho-fung Leung, Raymond Y. K. Lau
COLING5
2016 The determinants of crowdfunding success: A semantic text analytics approach
Raymond Y. K. Lau, Wei Xu 0008
Decis. Support Syst.2
2016 Extracting and reasoning about implicit behavioral evidences for detecting fraudulent online transactions in e-Commerce
Jie Zhao 0011, Raymond Y. K. Lau, Wenping Zhang, Deyu Tang
Decis. Support Syst.2
2016 Big data commerce
Raymond Y. K. Lau, J. Leon Zhao, Xunhua Guo
Inf. Manag.1
2016 Context-aware ontologies generation with basic level concepts from collaborative tags
Yi Cai 0001, Wenhao Chen 0001, Ho-fung Leung, Qing Li 0001, Haoran Xie 0001, Raymond Y. K. Lau, Huaqing Min, Fu Lee Wang
Neurocomputing6
2016 Incorporating sentiment into tag-based user profiles and resource profiles for personalized search in folksonomy
Haoran Xie 0001, Xiaodong Li 0007, Tao Wang 0036, Raymond Y. K. Lau, Tak-Lam Wong, Li Chen 0009, Fu Lee Wang, Qing Li 0001
Inf. Process. Manag.4
2016 Time series k-means: A new k-means type smooth subspace clustering for time series data
Xiaohui Huang 0003, Yunming Ye, Liyan Xiong, Raymond Y. K. Lau, Nan Jiang 0013, Shaokai Wang
Inf. Sci.4
2016 Dynamic Clustering Forest: An ensemble framework to efficiently classify textual data stream with concept drift
Yunming Ye, Haijun Zhang 0002, Xiaofei Xu 0001, Raymond Y. K. Lau, Feng Liu 0034
Inf. Sci.5
2016 Multi-opinion Ring: visualizing and predicting multiple opinion orientations in online social media
Xiaolin Du, Yunming Ye, Raymond Y. K. Lau, Yueping Li, Xiaohui Huang 0003
Multim. Tools Appl.3
2015 A Generative Model with Ensemble Manifold Regularization for Multi-view Clustering
Shaokai Wang, Yunming Ye, Raymond Y. K. Lau
ICIC (3)3
2015 OpinionRings: Inferring and visualizing the opinion tendency of socially connected users
Xiaolin Du, Yunming Ye, Raymond Y. K. Lau, Yueping Li
Decis. Support Syst.3
2015 Learning Context-Sensitive Domain Ontologies from Folksonomies: A Cognitively Motivated Method
abstract
Ontology is the backbone of the Semantic Web, helping users search for relevant resources from the Web of linked data. The existing context-free mapping approach between tags and concepts fails to address the problems of social synonymy and social polysemy when ontologies are induced from folksonomies. The novel contributions of this paper are threefold. First, grounded in the cognitively motivated category utility measure, a novel basic-level concept mining algorithm is developed to construct semantically rich concept vectors to alleviate the problem of social synonymy. Second, contextual aspects of ontology learning are exploited via probabilistic topic modeling to address the problem of social polysemy. Third, a novel context-sensitive domain ontology learning algorithm that combines link- and content-based semantic analysis is developed to identify both taxonomic and associative relations among concepts. To the best of our knowledge, this is the first successful research that exploits a cognitively motivated method to learn context-sensitive domain ontologies from folksonomies. By using the Open Directory Project ontology as a benchmark, we examined the effectiveness of the proposed algorithms based on social annotations crawled from three different folksonomy sites. Our experimental results show that the proposed ontology learning system significantly outperforms the best baseline system by 13.83% in terms of taxonomic F-measure. The practical implication of our research is that high-quality ontologies are constructed with minimal human intervention to facilitate concept-driven retrieval of linked data and the knowledge-based interoperability among enterprises.
Raymond Y. K. Lau, J. Leon Zhao, Wenping Zhang, Yi Cai 0001, Eric W. T. Ngai
INFORMS J. Comput.1
2014 Clustering tweets usingWikipedia concepts
Guoyu Tang, Yunqing Xia, Weizhi Wang, Raymond Y. K. Lau, Thomas Fang Zheng
LREC4
2014 Object typicality for effective Web of Things recommendations
Yi Cai 0001, Raymond Y. K. Lau, Stephen Shaoyi Liao, Chunping Li, Ho-fung Leung, Louis C. K. Ma
Decis. Support Syst.2
2014 Social analytics: Learning fuzzy product ontologies for aspect-oriented sentiment analysis
Raymond Y. K. Lau, Chunping Li, Stephen Shaoyi Liao
Decis. Support Syst.1
2014 An ontology-based Web mining method for unemployment rate prediction
Wei Xu 0008, Likuan Zhang, Raymond Y. K. Lau
Decis. Support Syst.4
2014 Coalition formation based on marginal contributions and the Markov process
Stephen Shaoyi Liao, Raymond Y. K. Lau, Tian-Ying Wu
Decis. Support Syst.3
2014 Effective Active Learning Strategies for the Use of Large-Margin Classifiers in Semantic Annotation: An Optimal Parameter Discovery Perspective
abstract
Classical supervised machine learning techniques have been explored for semantically annotating unstructured textual data such as consumers' comments archived at social media websites to extract business intelligence. However, these techniques often require a large number of manually labeled training examples to produce accurate annotations. Several active learning approaches that are designed based on probabilistic sequence models have been explored to minimize the number of labeled training examples for semantic annotation tasks. Recent research has shown that large-margin classifiers are viable alternatives to automated semantic annotation, given their strong generalization capabilities and the ability to process high-dimensional data. However, the existing active learning methods that are designed for probabilistic sequence models cannot be easily adapted and applied to large-margin classifiers. The main contribution of this paper is the development of novel active learning methods for large-margin classifiers to fill the aforementioned research gap. In particular, we propose an innovative perspective of taking active learning as a search of optimal parameters for large-margin classifiers. A rigorous evaluation involving two benchmark tests and an empirical test based on real-world data extracted from Amazon.com reveals that the proposed active learning methods can train effective classifiers with significantly fewer training examples while achieving similar annotation performance, compared to a typical state-of-the-art classifier that only uses several labeled training examples. More specifically, one of our proposed active learning methods can reduce the number of training examples by 19.74% at the 68% level of F1 when compared to the best baseline method, as evaluated based on the Amazon data set. Our research opens the door to the application of intelligent semantic annotation techniques to support real-world applications such as automatically analyzing consumer comments for customer relationship management.
Kaiquan Xu, Stephen Shaoyi Liao, Raymond Y. K. Lau, J. Leon Zhao
INFORMS J. Comput.3
2014 Product aspect extraction supervised with online domain knowledge
Tao Wang 0036, Yi Cai 0001, Ho-fung Leung, Raymond Y. K. Lau, Qing Li 0001, Huaqing Min
Knowl. Based Syst.4
2012 Mining fuzzy ontology for fuzzy granular IR systems
abstract
Similarity-based and popularity-based information retrieval (IR) models have been widely used by general Internet search engines. However, there are weaknesses of these models for supporting domain-specific IR. Grounded on the work in granular computing, we propose the notion of semantic information granulation estimated with respect to a fuzzy domain ontology to support domain-specific IR. The main contributions of this paper is the illustration of the design and development of a fuzzy granular IR system which can take into account two orthogonal dimensions, such as “similarity” and “granularity” to improve the effectiveness of domain-specific IR. In particular, the proposed fuzzy granular IR system is underpinned by a computational method for automated fuzzy ontology mining. Based on standard benchmark document collection, the results of our experiments confirm that the proposed fuzzy granular IR system outperforms a classical similarity-based (i.e, the vector space model) IR system for domain specific IR. Our research opens the door to the applications of granular computing and fuzzy ontology mining to enhance Internet search engines.
Raymond Y. K. Lau
FUZZ-IEEE1
2012 Unsupervised Multi-label Text Classification Using a World Knowledge Ontology
Xiaohui Tao 0001, Yuefeng Li 0001, Raymond Y. K. Lau, Hua Wang 0002
PAKDD (1)3
2012 Answering Typicality Query Based on Automatically Prototype Construction
abstract
In cognitive psychology, typicality refers to the degree of goodness of objects as exemplars in concepts. In this paper, we apply the idea of typicality analysis from cognitive psychology to query answering, and propose a novel method to answer typicality queries effectively based on theories in cognitive psychology. The proposed method adopts multi-prototype concept modeling and basic level category detection. By a systematic empirical evaluation using real data sets, we verify the accuracy and the effectiveness of our method on answering typicality queries.
Yi Cai 0001, Hong-Ke Zhao, Raymond Y. K. Lau, Ho-fung Leung, Huaqing Min
Web Intelligence4
2012 Latent Business Networks Mining: A Probabilistic Generative Model
abstract
Though numerous research has been devoted to social network discovery and analysis, relatively little research has been conducted on business network discovery. The main contribution of our research is the development of a novel probabilistic generative model for latent business networks mining. Our experimental results confirm that the proposed method outperforms the well-known vector space based model by 24% in terms of AUC value.
Wenping Zhang, Raymond Y. K. Lau, Yunqing Xia, Chunping Li, Wenjie Li 0002
Web Intelligence2
2012 An Aspect Query Language Model Based on Query Decomposition and High-Order Contextual Term Associations
abstract
In information retrieval (IR) research, more and more focus has been placed on optimizing a query language model by detecting and estimating the dependencies between the query and the observed terms occurring in the selected relevance feedback documents. In this paper, we propose a novel Aspect Language Modeling framework featuring term association acquisition, document segmentation, query decomposition, and an Aspect Model (AM) for parameter optimization. Through the proposed framework, we advance the theory and practice of applying high‐order and context‐sensitive term relationships to IR. We first decompose a query into subsets of query terms. Then we segment the relevance feedback documents into chunks using multiple sliding windows. Finally we discover the higher order term associations, that is, the terms in these chunks with high degree of association to the subsets of the query. In this process, we adopt an approach by combining the AM with the Association Rule (AR) mining. In our approach, the AM not only considers the subsets of a query as “hidden” states and estimates their prior distributions, but also evaluates the dependencies between the subsets of a query and the observed terms extracted from the chunks of feedback documents. The AR provides a reasonable initial estimation of the high‐order term associations by discovering the associated rules from the document chunks. Experimental results on various TREC collections verify the effectiveness of our approach, which significantly outperforms a baseline language model and two state‐of‐the‐art query language models namely the Relevance Model and the Information Flow model.
Dawei Song 0001, Peter Bruza, Raymond Y. K. Lau
Comput. Intell.4
2012 A two-stage decision model for information filtering
Yuefeng Li 0001, Xujuan Zhou, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau
Decis. Support Syst.5
2012 Combining social network and semantic concept analysis for personalized academic researcher recommendation
Yunhong Xu, Xitong Guo, Jin-Xing Hao, Jian Ma 0008, Raymond Y. K. Lau, Wei Xu 0008
Decis. Support Syst.5
2011 Leveraging web 2.0 data for scalable semi-supervised learning of domain-specific sentiment lexicons
abstract
Since manually constructing domain-specific sentiment lexicons is extremely time consuming and it may not even be feasible for domains where linguistic expertise is not available, research on automatic construction of domain-specific sentiment lexicons has become a hot topic in recent years. The main contribution of this paper is the illustration of a novel semi-supervised learning method which exploits both term-to-term and document-to-term relations hidden in a corpus for the construction of domain-specific sentiment lexicons. More specifically, the proposed two-pass pseudo labeling method combines shallow linguistic parsing and corpus-base statistical learning to make domain-specific sentiment extraction scalable with respect to the sheer volume of opinionated documents archived on the Internet these days. Our experiments show that the proposed method can generate high quality domain-specific sentiment lexicons according to users' evaluation.
Raymond Y. K. Lau, Chun Lam Lai, Peter Bruza, Kam-Fai Wong
CIKM1
2011 Pattern Mining for a Two-Stage Information Filtering System
Xujuan Zhou, Yuefeng Li 0001, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau
PAKDD (1)5
2011 Learning features through feedback for blog distillation
abstract
The paper is focused on blogosphere research based on the TREC blog distillation task, and aims to explore unbiased and significant features automatically and efficiently. Feedback from faceted feeds is introduced to harvest relevant features and information gain is used to select discriminative features. The evaluation result shows that the selected feedback features can greatly improve the performance and adapt well to the terabyte data.
Dehong Gao, Renxian Zhang, Wenjie Li 0002, Raymond Y. K. Lau, Kam-Fai Wong
SIGIR4
2011 Toward a semantic granularity model for domain-specific information retrieval
abstract
Both similarity-based and popularity-based document ranking functions have been successfully applied to information retrieval (IR) in general. However, the dimension of semantic granularity also should be considered for effective retrieval. In this article, we propose a semantic granularity-based IR model that takes into account the three dimensions, namely similarity, popularity, and semantic granularity, to improve domain-specific search. In particular, a concept-based computational model is developed to estimate the semantic granularity of documents with reference to a domain ontology. Semantic granularity refers to the levels of semantic detail carried by an information item. The results of our benchmark experiments confirm that the proposed semantic granularity based IR model performs significantly better than the similarity-based baseline in both a bio-medical and an agricultural domain. In addition, a series of user-oriented studies reveal that the proposed document ranking functions resemble the implicit ranking functions exercised by humans. The perceived relevance of the documents delivered by the granularity-based IR system is significantly higher than that produced by a popular search engine for a number of domain-specific search tasks. To the best of our knowledge, this is the first study regarding the application of semantic granularity to enhance domain-specific IR.
Xin Yan 0002, Raymond Y. K. Lau, Dawei Song 0001, Xue Li 0001, Jian Ma 0008
ACM Trans. Inf. Syst.2
2010 Rough sets based reasoning and pattern mining for a two-stage information filtering system
abstract
This paper presents a novel two-stage information filtering model which combines the merits of term-based and pattern- based approaches to effectively filter sheer volume of infor- mation. In particular, the first filtering stage is supported by a novel rough analysis model which efficiently removes a large number of irrelevant documents, thereby addressing the overload problem. The second filtering stage is empow- ered by a semantically rich pattern taxonomy mining model which effectively fetches incoming documents according to the specific information needs of a user, thereby addressing the mismatch problem. The experiments have been conducted to compare the proposed two-stage filtering (T-SM) model with other possible "term-based + pattern-based" or "term-based + term-based" IF models. The results based on the RCV1 corpus show that the T-SM model significantly outperforms other types of "two-stage" IF models.
Xujuan Zhou, Yuefeng Li 0001, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau
CIKM5
2010 Ontology-Based Specific and Exhaustive User Profiles for Constraint Information Fusion for Multi-agents
abstract
Intelligent agents are an advanced technology utilized in Web Intelligence. When searching information from a distributed Web environment, information is retrieved by multi-agents on the client site and fused on the broker site. The current information fusion techniques rely on cooperation of agents to provide statistics. Such techniques are computationally expensive and unrealistic in the real world. In this paper, we introduce a model that uses a world ontology constructed from the Dewey Decimal Classification to acquire user profiles. By search using specific and exhaustive user profiles, information fusion techniques no longer rely on the statistics provided by agents. The model has been successfully evaluated using the large INEX data set simulating the distributed Web environment.
Xiaohui Tao 0001, Yuefeng Li 0001, Raymond Y. K. Lau, Shlomo Geva
Web Intelligence3
2010 Predicting problem-solving performance with concept maps: An information-theoretic approach
Jin-Xing Hao, Ron Chi-Wai Kwok, Raymond Y. K. Lau, Angela Yan Yu
Decis. Support Syst.3
2009 An effective model of using negative relevance feedback for information filtering
abstract
Over the years, people have often held the hypothesis that negative feedback should be very useful for largely improving the performance of information filtering systems; however, we have not obtained very effective models to support this hypothesis. This paper, proposes an effective model that use negative relevance feedback based on a pattern mining approach to improve extracted features. This study focuses on two main issues of using negative relevance feedback: the selection of constructive negative examples to reduce the space of negative examples; and the revision of existing features based on the selected negative examples. The former selects some offender documents, where offender documents are negative documents that are most likely to be classified in the positive group. The later groups the extracted features into three groups: the positive specific category, general category and negative specific category to easily update the weight. An iterative algorithm is also proposed to implement this approach on RCV1 data collections, and substantial experiments show that the proposed approach achieves encouraging performance.
Abdulmohsen Algarni, Yuefeng Li 0001, Yue Xu 0001, Raymond Y. K. Lau
CIKM4
2009 Toward a Fuzzy Domain Ontology Extraction Method for Adaptive e-Learning
abstract
With the widespread applications of electronic learning (e-Learning) technologies to education at all levels, increasing number of online educational resources and messages are generated from the corresponding e-Learning environments. Nevertheless, it is quite difficult, if not totally impossible, for instructors to read through and analyze the online messages to predict the progress of their students on the fly. The main contribution of this paper is the illustration of a novel concept map generation mechanism which is underpinned by a fuzzy domain ontology extraction algorithm. The proposed mechanism can automatically construct concept maps based on the messages posted to online discussion forums. By browsing the concept maps, instructors can quickly identify the progress of their students and adjust the pedagogical sequence on the fly. Our initial experimental results reveal that the accuracy and the quality of the automatically generated concept maps are promising. Our research work opens the door to the development and application of intelligent software tools to enhance e-Learning.
Raymond Y. K. Lau, Dawei Song 0001, Yuefeng Li 0001, Chun-Ho Cheung, Jin-Xing Hao
IEEE Trans. Knowl. Data Eng.1
2008 A two-stage text mining model for information filtering
abstract
Mismatch and overload are the two fundamental issues regarding the effectiveness of information filtering. Both term-based and pattern (phrase) based approaches have been employed to address these issues. However, they all suffer from some limitations with regard to effectiveness. This paper proposes a novel solution that includes two stages: an initial topic filtering stage followed by a stage involving pattern taxonomy mining. The objective of the first stage is to address mismatch by quickly filtering out probable irrelevant documents. The threshold used in the first stage is motivated theoretically. The objective of the second stage is to address overload by apply pattern mining techniques to rationalize the data relevance of the reduced document set after the first stage. Substantial experiments on RCV1 show that the proposed solution achieves encouraging performance.
Yuefeng Li 0001, Xujuan Zhou, Peter Bruza, Yue Xu 0001, Raymond Y. K. Lau
CIKM5
2008 Knowledge discovery for adaptive negotiation agents in e-marketplaces
Raymond Y. K. Lau, Yuefeng Li 0001, Dawei Song 0001, Ron Chi-Wai Kwok
Decis. Support Syst.1
2008 Towards a belief-revision-based adaptive and context-sensitive information retrieval system
abstract
In an adaptive information retrieval (IR) setting, the information seekers' beliefs about which terms are relevant or nonrelevant will naturally fluctuate. This article investigates how the theory of belief revision can be used to model adaptive IR. More specifically, belief revision logic provides a rich representation scheme to formalize retrieval contexts so as to disambiguate vague user queries. In addition, belief revision theory underpins the development of an effective mechanism to revise user profiles in accordance with information seekers' changing information needs. It is argued that information retrieval contexts can be extracted by means of the information-flow text mining method so as to realize a highly autonomous adaptive IR system. The extra bonus of a belief-based IR model is that its retrieval behavior is more predictable and explanatory. Our initial experiments show that the belief-based adaptive IR system is as effective as a classical adaptive IR system. To our best knowledge, this is the first successful implementation and evaluation of a logic-based adaptive IR model which can efficiently process large IR collections.
Raymond Y. K. Lau, Peter Bruza, Dawei Song 0001
ACM Trans. Inf. Syst.1
2007 A Parallel Genetic Algorithm for Floorplan Area Optimization
abstract
Floorplanning is an important problem in Very Large- Scale Integrated-circuit (VLSI) design automation as it de- termines the performance, size, yield and reliability of VLSI chips. From the computational point of view, floorplan area minimization is an NP-hard problem. This paper presents a parallel genetic algorithm (GA) for floorplan area opti- mization. The parallel GA is based an island model with an asynchronous migration mechanism, and is implemented using Web services and multithreading technologies. The parallel GA is compared with a sequential GA that the par- allel GA is based on. Experimental results show that the parallel GA can produce better results than the sequential GA when they use the same amount of computing resources. In addition, since the number of islands and migration in- terval are two important parameters that directly affect the performance of island-based parallel GAs, the impact of the two parameters on the performance of the parallel GA are empirically studied in this paper.
Maolin Tang, Raymond Y. K. Lau
ISDA2
2007 Mining Fuzzy Domain Ontology from Textual Databases
abstract
Ontology plays an essential role in the formalization of common information (e.g., products, services, relationships of businesses) for effective human-computer interactions. However, engineering of these ontologies turns out to be very labor intensive and time consuming. Although some text mining methods have been proposed for automatic or semi-automatic discovery of crisp ontologies, the robustness, accuracy, and computational efficiency of these methods need to be improved to support large scale ontology construction for real-world applications. This paper illustrates a novel fuzzy domain ontology mining algorithm for supporting real-world ontology engineering. In particular, contextual information of the knowledge sources is exploited for the extraction of high quality domain ontologies and the uncertainty embedded in the knowledge sources is modeled based on the notion of fuzzy sets. Empirical studies have confirmed that the proposed method can discover high quality fuzzy domain ontology which leads to significant improvement in information retrieval performance.
Raymond Y. K. Lau, Yuefeng Li 0001, Yue Xu 0001
Web Intelligence1
2007 Using Information Filtering in Web Data Mining Process
abstract
The amount of Web information is growing rapidly, improving the efficiency and accuracy of Web information retrieval is uphill battle. There are two fundamental issues regarding the effectiveness of Web information gathering: information mismatch and overload. To tackle these difficult issues, an integrated information filtering and sophisticated data processing model has been presented in this paper. In the first phase of the proposed scheme, an information filter that based on user search intents was incorporated in Web search process to quickly filter out irrelevant data. In the second data processing phase, a pattern taxonomy model (PTM) was carried out using the reduced data. PTM rationalizes the data relevance by applying data mining techniques that involves more rigorous computations. Several experiments have been conducted and the results show that more effective and efficient access Web information has been achieved using the new scheme.
Xujuan Zhou, Yuefeng Li 0001, Peter Bruza, Sheng-Tang Wu, Yue Xu 0001, Raymond Y. K. Lau
Web Intelligence6
2007 An intelligent information agent for document title classification and filtering in document-intensive domains
Dawei Song 0001, Raymond Y. K. Lau, Peter Bruza, Kam-Fai Wong, Ding-Yi Chen
Decis. Support Syst.2
2007 Sequential Pattern Mining and Nonmonotonic Reasoning for Intelligent Information Agents
abstract
With the explosive growth of information available on the Internet, more effective data mining and data reasoning mechanism is required to process the sheer volume of information. Belief revision logic offers the expressive power to represent information retrieval contexts, and it also provides a sound inference mechanism to model the nonmonotonicity arising in changing retrieval contexts. Contextual knowledge for information retrieval can be extracted via efficient sequential pattern mining. We present a pattern taxonomy extraction model which efficiently performs the task of discovering descriptive frequent sequential patterns by pruning the noisy associations. This paper illustrates a novel approach of integrating the sequential data mining method into the belief revision based adaptive information agents to improve the agents' learning autonomy and prediction power. Initial experiments show that our belief revision logic and sequential pattern mining based intelligent information agents outperform the vector space model based information agents. Our work opens the door to the development of next generation of intelligent information agents to alleviate the information overload problem.
Raymond Y. K. Lau, Yuefeng Li 0001, Sheng-Tang Wu, Xujuan Zhou
Int. J. Pattern Recognit. Artif. Intell.1
2006 Utilizing Search Intent in Topic Ontology-Based User Profile for Web Mining
abstract
It is well known that taking the Web user profiles into account can enhance the effectiveness of Web mining systems. However, due to the dynamic and complex nature of Web users, automatically acquiring worthwhile user profiles was found to be very challenging. Ontology-based user profile can possess more accurate user information. This research emphasizes on acquiring search intentions information. This paper presents a new approach of developing user profile for Web searching. The model considers the user's search intentions by the process of PTM (Pattern-Taxonomy Model). Initial experiments show that the user profile based on search intention is more useful than the generic PTM user profile. Developing user profile that contains user search intentions is essential for effective Web search and retrieval.
Xujuan Zhou, Sheng-Tang Wu, Yuefeng Li 0001, Yue Xu 0001, Raymond Y. K. Lau, Peter Bruza
Web Intelligence5
2006 An evolutionary learning approach for adaptive negotiation agents
abstract
Developing effective and efficient negotiation mechanisms for real-world applications such as e-business is challenging because negotiations in such a context are characterized by combinatorially complex negotiation spaces, tough deadlines, very limited information about the opponents, and volatile negotiator preferences. Accordingly, practical negotiation systems should be empowered by effective learning mechanisms to acquire dynamic domain knowledge from the possibly changing negotiation contexts. This article illustrates our adaptive negotiation agents, which are underpinned by robust evolutionary learning mechanisms to deal with complex and dynamic negotiation contexts. Our experimental results show that GA-based adaptive negotiation agents outperform a theoretically optimal negotiation mechanism that guarantees Pareto optimal. Our research work opens the door to the development of practical negotiation systems for real-world applications. © 2006 Wiley Periodicals, Inc. Int J Int Syst 21: 41–72, 2006.
Raymond Y. K. Lau, Maolin Tang, On Wong, Stephen Milliner, Yi-Ping Phoebe Chen
Int. J. Intell. Syst.1
2005 Adaptive negotiation agents for e-business
abstract
Negotiation has been identified as one of the key steps in Business-to-Business (B2B) transaction models. However, developing effective and efficient negotiation mechanisms for e-Business is quite challenging since negotiations in such a context are characterized by combinatorial complex negotiation spaces, tough deadlines, incomplete information about the opponents, and volatile negotiator preferences. Classical negotiation models are not able to offer a satisfactory solution to address all these issues. This paper illustrates our adaptive negotiation agents which are underpinned by a robust evolutionary learning mechanism to deal with complex and dynamic negotiation situations often encountered in e-Business applications. Our experimental results show that the proposed evolutionary negotiation agents outperform a theoretically optimal negotiation mechanism which guarantees Pareto optimal. Our research work opens the door to the development of practical negotiation systems for e-Business.
Raymond Y. K. Lau
ICEC1
2004 Towards Belief Revision Logic Based Adaptive and Persuasive Negotiation Agents
Raymond Y. K. Lau, Siu Y. Chan
PRICAI1
2004 Belief revision for adaptive information retrieval
abstract
Applying Belief Revision logic to model adaptive information retrieval is appealing since it provides a rigorous theoretical foundation to model partiality and uncertainty inherent in any information retrieval (IR) processes. In particular, a retrieval context can be formalised as a belief set and the formalised context is used to disambiguate vague user queries. Belief revision logic also provides a robust computational mechanism to revise an IR system's beliefs about the users' changing information needs. In addition, information flow is proposed as a text mining method to automatically acquire the initial IR contexts. The advantage of a belief-based IRsystem is that its IR behaviour is more predictable and explanatory. However, computational efficiency is often a concern when the belief revision formalisms are applied to large real-life applications. This paper describes our belief-based adaptive IR system which is underpinned by an efficient belief revision mechanism. Our initial experiments show that the belief-based symbolic IR model is more effective than a classical quantitative IR model. To our best knowledge, this is the first successful empirical evaluation of a logic-based IR model based on large IR benchmark collections.
Raymond Y. K. Lau, Peter Bruza, Dawei Song 0001
SIGIR1
2004 Discovering Negotiation Knowledge for a Probabilistic Negotiation Web Service in e-Business
abstract
Negotiation has long been identified as one of the key processes in e-Business. Classical negotiation models have limited use in e-Business because these models often assume that complete information about the negotiation spaces is available and the computational efficiency is negligible. This paper proposes a practical negotiation system which is developed based on Bayesian learning and engineered as a Web service. The knowledge acquisition bottle-neck presented in previous Bayesian negotiation approaches is resolved by a novel method of inducing domain knowledge based on negotiation history. Our preliminary experiment shows that the proposed probabilistic negotiation method is more effective than a non-learning negotiation model and the extra computational time incurs in the probabilistic model is negligible.
Raymond Y. K. Lau, Eivind Valdal
Web Intelligence1
2003 Utilizing Domain Knowledge in Neuroevolution
James Fan, Raymond Y. K. Lau, Risto Miikkulainen
ICML2
2003 Belief Revision for Adaptive Recommender Agents in E-commerce
Raymond Y. K. Lau
IDEAL1
2003 Belief Revision and Text Mining for Adaptive Recommender Agents
Raymond Y. K. Lau, Peter van den Brand
ISMIS1
2003 Classifying Document Titles Based on Information Inference
Dawei Song 0001, Peter Bruza, Zi Huang, Raymond Y. K. Lau
ISMIS4
2003 Context Sensitive Text Mining and Belief Revision for Adaptive Information Retrieval
abstract
Autonomous information agents alleviate the information overload problem on the Internet. The AGM belief revision framework provides a rigorous foundation to develop adaptive information agents. The expressive power of the belief revision logic allow a user's information preferences and contextual knowledge of a retrieval situation to be captured and reasoned about within a single logical framework. Contextual knowledge for information retrieval can be acquired via context sensitive text mining. We illustrate a novel approach of integrating the proposed text mining method into the belief revision based adaptive information agents to improve the agents' learning autonomy and prediction power.
Raymond Y. K. Lau
Web Intelligence1
2003 Context-sensitive text mining and belief revision for intelligent information retrieval on the web
Raymond Y. K. Lau
Web Intell. Agent Syst.1
2001 Belief Revision for Adaptive Information Filtering Agents
abstract
Agent-based information filtering alleviates the problem of information overload on the Internet by proactively scanning through the incoming stream of information on behalf of the users. Nevertheless, users' information needs will change over time. Therefore, it is essential for the information filtering agents to learn and adapt to the users' changing information needs in order to maintain the accuracy of the filtering process. Applying logic-based representation and adaptation to adaptive information filtering agents is promising since the semantic relationships among information items can be captured and reasoned about during the agents' learning and adaptation processes. This opens the door to a more responsive reinforcement learning than can be obtained from a purely statistical approach. The AGM belief revision paradigm that models rational and minimal change of an agent's beliefs offers a sound theoretical foundation for constructing the learning components of adaptive information filtering agents. This paper describes a symbolic framework for representing domain knowledge in an adaptive information filtering agent, and illustrates how the AGM belief revision paradigm can be applied to develop the agent's learning mechanism.
Raymond Y. K. Lau, Arthur H. M. ter Hofstede, Peter Bruza
Int. J. Cooperative Inf. Syst.1
2000 Belief revision and possibilistic logic for adaptive information filtering agents
abstract
Prototypes of adaptive information agents have been developed to alleviate the problem of information overload on the Internet. However, the explanatory power and the learning autonomy of these agents are weak. A logic based framework for the development of information agents is appealing since semantic relationships among information objects can be captured and reasoned about. This sheds light on better explanatory power and higher learning autonomy of information agents. The paper illustrates how the AGM belief revision and possibilistic logic can be applied to develop the learning and the filtering components of adaptive information filtering agents. Their impact on the agents' learning autonomy and explanatory power is also discussed.
Raymond Y. K. Lau, Arthur H. M. ter Hofstede, Peter Bruza, Kam-Fai Wong
ICTAI1
1999 Organization, communication, and control in the GALAXY-II conversational system
Stephanie Seneff, Raymond Y. K. Lau, Joseph Polifroni
EUROSPEECH2
1998 A unified framework for sublexical and linguistic modelling supporting flexible vocabulary speech understanding
abstract
In [9], we introduced the ANGIE framework for modelling speech where morphological and phonological substructures of words are jointly characterized by a context-free grammar and represented in a multi-layered hierarchical structure. In [6], we demonstrated a competitive word-spotter based on the ANGIE framework and presented several results comparing the performance of various sublexical filler models. In the present work, completed as a part of [5], we extend the ANGIE framework to a competitive full continuous speech recognition system. Furthermore, given that ANGIE is based on a context-free framework, we have decided to combine ANGIE with TINA ([8]), a contextfree based framework for natural language understanding, into an integrated system. The integrated system led to a 21.7% reduction in word error rate compared to a baseline word bigram recognizer on ATIS. Numerous issues relating to the construction of the combined system were explored. We have also examined the addition of n...
Raymond Y. K. Lau, Stephanie Seneff
ICSLP1
1998 GALAXY-II: a reference architecture for conversational system development
abstract
GALAXY is a client-server architecture for accessing on-line in-formation using spoken dialogue that we introduced at ICSLP-94. It has served as the testbed for developing human language technologies for our group for several years. Recently, we have initiated a significant redesign of the GALAXY architecture to make it easier for many researchers to develop their own appli-cations, using either exclusively their own servers or intermixing them with servers developed by others. This redesign was done in part due to the fact that GALAXY has been designated as the first reference architecture for the new DARPA Communicator Program. The purpose of this paper is to document the changes to GALAXY that led to this first reference architecture, which makes use of a scripting language for flow control to provide flexible interaction among the servers, and a set of libraries to support rapid prototyping of new servers. We describe the new reference architecture in some detail, and report on the current status of its development. 1.
Stephanie Seneff, Edward Hurley, Raymond Y. K. Lau, Christine Pao, Philipp Schmid 0001, Victor Zue
ICSLP3
1997 Webgalaxy - integrating spoken language and hypertext navigation
abstract
The growth in the quantity of information and services offered online has been phenomenal. Nevertheless, access mechanisms have remained relatively primitive, requiring users to primarily point and click their way through a forest of Web links and to expend valuable cognitive capacities to track the geography of the Web space. Conversational systems can provide an intuitive, flexible multi-modal interface to online resources. The explosive growth of the World Wide Web, the continuing standardization of Web related technologies, and the growing penetration of Internet access enable us to embed a very thin client inside a standard Web browser, making conversational interfaces available to a much wider audience. This paper presents WebGALAXY, a conversational spoken language system for access to selected online resources from within a typical browser. A thin Java based client is employed as the front end with much of the speech and natural language processing occuring on remote servers. 1...
Raymond Y. K. Lau, Giovanni Flammia, Christine Pao, Victor Zue
EUROSPEECH1
1997 Providing sublexical constraints for word spotting within the ANGIE framework
abstract
We describe our recent work in implementing a word-spotting system based on the ANGIE framework and the effects of varying the nature of the sublexical constraints placed upon the wordspotter 's filler model. ANGIE is a framework for modelling speech where the morphological and phonological substructures of words are jointly characterized by a context-free grammar and are represented in a multi-layered hierarchical structure. In this representation, the upper layers capture syllabification, morphology, and stress, the preterminal layer represents phonemics, and the bottom terminal categories are the phones. ANGIE provides a flexible framework where we can explore the effects of sublexical constraints within a word-spotting environment. Our experiments with spotting city names in ATIS validate the intuition that increasing the constraints present in the model improves performance, from 85.3 FOM for phone bigram to 89.3 FOM for a word lexicon. They also empirically strengthens our belief...
Raymond Y. K. Lau, Stephanie Seneff
EUROSPEECH1
1997 WebGALAXY: Beyond Point and Click - a Conversational Interface to a Browser
Raymond Y. K. Lau, Giovanni Flammia, Christine Pao, Victor Zue
Comput. Networks1
1996 ANGIE: a new framework for speech analysis based on morpho-phonological modelling
abstract
This paper describes a new system for speech analysis, ANGIE, which characterizes word substructure in terms of a trainable grammar.ANGIE capture morpho-phonemic and phonological phenomena through a hierarchical framework.The terminal categories can be alternately letters or phone units, yielding a reversible letter-tosound/sound-to-letter system.In conjunction with a segment network and acoustic phone models, the system can produce phonemicto-phonetic alignments for speech waveforms.For speech recognition, ANGIE uses a one-pass bottom-up best-first search strategy.Evaluated in the ATIS domain, ANGIE achieved a phone error rate of 36%, as compared with 40% achieved with a baseline phone-bigram based recognizer under similar conditions.ANGIE potentially offers many attractive features, including dynamic vocabulary adaptation, as well as a framework for handling unknown words. £Previous experiments have yielded improved pronunciation accuracy without this layer.
Stephanie Seneff, Raymond Y. K. Lau, Helen M. Meng
ICSLP2
1993 Trigger-based language models: a maximum entropy approach
Raymond Y. K. Lau, Ronald Rosenfeld, Salim Roukos
ICASSP (2)1