Daniel S. Yeung

dblp:36/896 · DBLP profile ↗
← Back
150ranked-venue papers
29as first author
13since 2021 · last 2026
0000-0001-9397-8865ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 84 · 18 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 49 · 9 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 39 · 5 first-authorDatabases, data management, data science and information retrieval · 12 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 3Security and privacy · 2Computer networks · 1
YearPublicationVenuePosition
2026 Contrastive learning with auxiliary model
Linyi Xu, Patrick P. K. Chan, Jingwen Deng, Xuanming Liang, Daniel S. Yeung
Pattern Recognit.5
2025 Real-world nighttime image dehazing using contrastive and adversarial learning
Jingwen Deng, Patrick P. K. Chan, Daniel S. Yeung
Pattern Recognit.3
2024 Unsupervised contaminated user profile identification against shilling attack in recommender system
abstract
A recommender system is susceptible to manipulation through the injection of carefully crafted profiles. Some recent profile identification methods only perform well in specific attack scenarios. A general attack detection method is usually complicated or requires label samples. Such methods are prone to overtraining easily, and the process of annotation incurs high expenses. This study proposes an unsupervised divide-and-conquer method aiming to identify attack profiles, utilizing a specifically designed model for each kind of shilling attack. Initially, our method categorizes the profile set into two attack types, namely Standard and Obfuscated Behavior Attacks. Subsequently, profiles are separated into clusters within the extracted feature space based on the identified attack type. The selection of attack profiles is then determined through target item analysis within the suspected cluster. Notably, our method offers the advantage of requiring no prior knowledge or annotation. Furthermore, the precision is heightened as the identification method is designed to a specific attack type, employing a less complicated model. The outstanding performance of our model, validated through experimental results on MovieLens-100K and Netflix under various attack settings, demonstrates superior accuracy and reduced running time compared to current detection methods in identifying Standard and Obfuscated Behavior Attacks.
Patrick P. K. Chan, Zhi-Min He, Daniel S. Yeung
Intell. Data Anal.4
2024 Evasion on general GAN-generated image detection by disentangled representation
Patrick P. K. Chan, Chuanxin Zhang, Jingwen Deng, Daniel S. Yeung
Inf. Sci.6
2023 Multi-proxy based deep metric learning
Patrick P. K. Chan, Shute Li, Jingwen Deng, Daniel S. Yeung
Inf. Sci.4
2022 Progressive editing with stacked Generative Adversarial Network for multiple facial attribute editing
Patrick P. K. Chan, Daniel S. Yeung
Comput. Vis. Image Underst.4
2022 Unsupervised Domain Adaptation for Gesture Identification Against Electrode Shift
abstract
Surface electromyogram (sEMG)-based hand gesture recognition, which interprets commands given by humans through sEMG signals, performs well in many studies. However, its recognition accuracy drops dramatically due to electrode shift since the distributions of motion classes are changed. Although calibrating the system with newly collected samples after electrode shift maintains the accuracy, collecting labeled samples is inconvenient and time-consuming since the procedure is rigid. However, the calibration may not work properly without label, especially when the change is significant. This study proposes a user friendly and convenient calibration method for hand gesture recognition by an unsupervised domain adaptation method, which only obtains the unlabeled samples of preselected benchmark classes from users in calibration. The change of benchmark classes is captured by unlabeled samples by a clustering method. The other classes are estimated based on the benchmark classes by regression models. As a result, the information of all classes is used to calibrate the system. Linear discriminant analysis is used to demonstrate our model. A dataset with ten subjects is collected to verify the performance empirically. Experimental results confirm that our method utilizes the unlabeled benchmark class samples in calibration and achieves 75.55% average accuracy. Our method is more robust to electrode shift and improves around 8.5% accuracy consistently on all subjects compared with the methods without calibration or label information in calibration. Although the accuracy of our method is slightly less than the ones using label calibration samples, our calibration data collection is more convenient and less complicated.
Patrick P. K. Chan, Qiu Xia Li, Yinfeng Fang, Linyi Xu, Kairu Li, Honghai Liu 0001, Daniel S. Yeung
IEEE Trans. Hum. Mach. Syst.7
2022 Item Relationship Graph Neural Networks for E-Commerce
abstract
In a modern e-commerce recommender system, it is important to understand the relationships among products. Recognizing product relationships-such as complements or substitutes-accurately is an essential task for generating better recommendation results, as well as improving explainability in recommendation. Products and their associated relationships naturally form a product graph, yet existing efforts do not fully exploit the product graph's topological structure. They usually only consider the information from directly connected products. In fact, the connectivity of products a few hops away also contains rich semantics and could be utilized for improved relationship prediction. In this work, we formulate the problem as a multilabel link prediction task and propose a novel graph neural network-based framework, item relationship graph neural network (IRGNN), for discovering multiple complex relationships simultaneously. We incorporate multihop relationships of products by recursively updating node embeddings using the messages from their neighbors. An edge relational network is designed to effectively capture relational information between products. Extensive experiments are conducted on real-world product data, validating the effectiveness of IRGNN, especially on large and sparse product graphs.
Weiwen Liu, Yin Zhang 0011, Jianling Wang, Yun He 0001, James Caverlee, Patrick P. K. Chan, Daniel S. Yeung, Pheng-Ann Heng
IEEE Trans. Neural Networks Learn. Syst.7
2021 Weakly Supervised Semantic Segmentation with Patch-Based Metric Learning Enhancement
Patrick P. K. Chan, Keke Chen, Linyi Xu, Xiaoman Hu, Daniel S. Yeung
ICANN (3)5
2021 Multiple-Model Based Defense for Deep Reinforcement Learning Against Adversarial Attack
Patrick P. K. Chan, Yaxuan Wang, Natasha Kees, Daniel S. Yeung
ICANN (1)4
2021 Distribution-based Adversarial Filter Feature Selection against Evasion Attack
abstract
Feature selection plays an important role in machine learning in order to reduce model complexity and extract more meaningful information. The recent studies indicate that not only the generalization ability but also the security should be considered in selecting features in an adversarial environment in which a classifier may be misled by an adversary intentionally. However, since only the nearest legitimate sample to a malicious sample is considered, the existing adversarial filter feature selection method is sensitive to outlier samples and is not suitable for Boolean features. A distribution-based adversarial filter feature selection method is proposed against evasion attack in this study. Our method uses distribution-based measurements, such as Symmetric Uncertainty and Earth Mover's Distance, are used to quantify the generalization ability and security of a feature subset. The experiments suggest our proposed method outperforms existing methods in terms of robustness and stability in datasets with Boolean and real-valued features.
Patrick P. K. Chan, Yuanchao Liang, Fei Zhang 0003, Daniel S. Yeung
IJCNN4
2021 Class-Specific Affinity based Weakly Supervised Semantic Segmentation with Neutral Region Exploration
abstract
Image-level weakly supervised semantic segmentation (WSSS) reduces the cost of semantic segmentation significantly as only category labels are required. In the WSSS pipeline, the initially obtained seed areas are imprecise and should be refined. AffinityNet is a widely used seed area refinement method which learns pixel-level affinity between coordinates to adjust the object regions. However, the capabilities of AffinityNet are limited due to the complex simultaneous exploration of the affinities of all object classes, as the classes have different characteristics. This paper proposes a class-specific AffinityNet (CSANet) in which the affinities of each class are learnt separately by the class-specific affinity layers. Moreover, as the learning task is simple because the affinity of only one class is focused, the information contained by the uncertain neutral regions is extracted to provide additional information in the training of the class-specific affinity layers. The empirical results demonstrate the effectiveness of our method, which captures more precise boundaries and significantly improves segmentation results on the PASCAL VOC 2012 segmentation benchmark.
Keke Chen, Patrick P. K. Chan, Tianyi Xiang, Natasha Kees, Daniel S. Yeung
IJCNN5
2021 Transfer learning based countermeasure against label flipping poisoning attack
Patrick P. K. Chan, Fengzhi Luo, Zitong Chen, Ying Shu, Daniel S. Yeung
Inf. Sci.5
2020 Adversarial Attack against Deep Reinforcement Learning with Static Reward Impact Map
abstract
Security problems of deep reinforcement learning draw much attention recently. Previous works on adversary attack mainly focus on preventing the targeted agent from choosing the most desirable action at each step, which may not reduce the cumulative reward effectively. In this paper, we first investigate how changing features affect the cumulative reward achieved by an agent. The static reward impact map is introduced to quantify the influence on the reward of each feature experimentally. By focusing on tasks with the static reward impact map, an adversarial attack method against deep reinforcement learning aiming to minimize the cumulative reward is proposed. Features with the large reward impact are perturbed in crafting an adversarial sample. Deep Q-network is selected to demonstrate the performance of our attack method in the experiments. The results indicate that our proposed method achieves better performance than the existing one-time attack method and the random attack in terms of the cumulative reward and the successful attack rate under both white-box and black-box settings.
Patrick P. K. Chan, Yaxuan Wang, Daniel S. Yeung
AsiaCCS3
2019 Incremental Hash-Bit Learning for Semantic Image Retrieval in Nonstationary Environments
abstract
Images are uploaded to the Internet over time which makes concept drifting and distribution change in semantic classes unavoidable. Current hashing methods being trained using a given static database may not be suitable for nonstationary semantic image retrieval problems. Moreover, directly retraining a whole hash table to update knowledge coming from new arriving image data may not be efficient. Therefore, this paper proposes a new incremental hash-bit learning method. At the arrival of new data, hash bits are selected from both existing and newly trained hash bits by an iterative maximization of a 3-component objective function. This objective function is also used to weight selected hash bits to re-rank retrieved images for better semantic image retrieval results. The three components evaluate a hash bit in three different angles: 1) information preservation; 2) partition balancing; and 3) bit angular difference. The proposed method combines knowledge retained from previously trained hash bits and new semantic knowledge learned from the new data by training new hash bits. In comparison to table-based incremental hashing, the proposed method automatically adjusts the number of bits from old data and new data according to the concept drifting in the given data via the maximization of the objective function. Experimental results show that the proposed method outperforms existing stationary hashing methods, table-based incremental hashing, and online hashing methods in 15 different simulated nonstationary data environments.
Wing W. Y. Ng, Xing Tian, Witold Pedrycz, Xizhao Wang, Daniel S. Yeung
IEEE Trans. Cybern.5
2018 Convolutional Neural Networks based Click-Through Rate Prediction with Multiple Feature Sequences
abstract
Convolutional Neural Network (CNN) achieved satisfying performance in click-through rate (CTR) prediction in recent studies. Since features used in CTR prediction have no meaningful sequence in nature, the features can be arranged in any order. As CNN learns the local information of a sample, the feature sequence may influence its performance significantly. However, this problem has not been fully investigated. This paper firstly investigates whether and how the feature sequence affects the performance of the CNN-based CTR prediction method. As the data distribution of CTR prediction changes with time, the best current sequence may not be suitable for future data. Two multi-sequence models are proposed to learn the information provided by different sequences. The first model learns all sequences using a single feature learning module, while each sequence is learnt individually by a feature learning module in the second one. Moreover, a method of generating a set of embedding sequences which aims to consider the combined influence of all feature pairs on feature learning is also introduced. The experiments are conducted to demonstrate the effectiveness and stability of our proposed models in the offline and online environment on both the benchmark Avazu dataset and a real commercial dataset.
Patrick P. K. Chan, Xian Hu, Daniel S. Yeung
IJCAI4
2018 Shilling Attack Detection Using Rated Item Correlation for Collaborative Filtering
abstract
Collaborative filtering (CF) is vulnerable under shilling attack, which misleads recommendation of CF by injecting well-crafted profiles to a targeted system. Although a number of supervised learning based shilling attack detection methods are proposed, their features mainly measure rating values and items of a profile individually, but ignore the relation between items. This study aims to enhance the robustness of CF against shilling attack by considering the rated item correlation. Real users rate items based on their preferences, but rated items are randomly selected for malicious users profiles in most shilling attack. Therefore, the rated item correlation of real and malicious profiles is different. Three features are proposed to capture the information from different intervals of the distribution of rated item correlation in terms of Cosine Association (CA). A benchmark dataset, MovieLens 100K, is used to evaluate the proposed features. The discrimination ability of the proposed features is also illustrated. The experimental results suggest that the proposed features have significant contribution on shilling attack detection.
Keke Chen, Patrick P. K. Chan, Daniel S. Yeung
SMC3
2018 Bagging-boosting-based semi-supervised multi-hashing with query-adaptive re-ranking
Wing W. Y. Ng, Xiancheng Zhou, Xing Tian, Xizhao Wang, Daniel S. Yeung
Neurocomputing5
2018 Face Liveness Detection Using a Flash Against 2D Spoofing Attack
abstract
Face recognition technique has been widely applied to personal identification systems due to its satisfying performance. However, its security may be a crucial issue, since many studies have shown that face recognition systems may be vulnerable in an adversarial environment, in which an adversary can camouflage as a legitimate user in order to mislead the system. Although face liveness detection methods have been proposed to distinguish real and fake faces, they are either time-consuming, costly, or sensitive to noise and illumination. This paper proposes a face liveness detection method with flash against 2D spoofing attack. Flash not only can enhance the differentiation between legitimate and illegitimate users, but it also reduces the influence of environmental factors. Two images are taken from a subject, one with flash and another without flash. Four texture and 2D structure descriptors with low computational complexity are used to capture information of the two images in our model. Advantages of our method include low installation cost of flash and no user cooperation required. A data set of 50 subjects collected under different scenarios is used in the experiments to evaluate the proposed method. The experimental results indicate that the proposed model performs better than existing liveness detection methods in different environmental scenarios. This paper confirms that the use of flash successfully improves face liveness detection in terms of accuracy, robustness, and running time.
Patrick P. K. Chan, Weiwen Liu, Danni Chen, Daniel S. Yeung, Fei Zhang 0003, Xizhao Wang, Chien-Chang Hsu
IEEE Trans. Inf. Forensics Secur.4
2017 Sensitivity based robust learning for stacked autoencoder against evasion attack
Patrick P. K. Chan, Xian Hu, Eric C. C. Tsang, Daniel S. Yeung
Neurocomputing5
2017 Incremental Hashing for Semantic Image Retrieval in Nonstationary Environments
abstract
A very large volume of images is uploaded to the Internet daily. However, current hashing methods for image retrieval are designed for static databases only. They fail to consider the fact that the distribution of images can change when new images are added to the database over time. The changes in the distribution of images include both discovery of a new class and a distribution of images within a class owing to concept drift. Retraining of hash tables using all images in the database requires a large computation effort. This is also biased to old data owing to the huge volume of old images which leads to a poor retrieval performance over time. In this paper, we propose the incremental hashing (ICH) method to deal with the two aforementioned types of changes in the data distribution. The ICH uses a multihashing to retain knowledge coming from images arriving over time and a weight-based ranking to make the retrieval results adaptive to the new data environment. Experimental results show that the proposed method is effective in dealing with changes in the database.
Wing W. Y. Ng, Xing Tian, Yueming Lv, Daniel S. Yeung, Witold Pedrycz
IEEE Trans. Cybern.4
2016 Dual autoencoders features for imbalance classification problem
Wing W. Y. Ng, Guangjun Zeng, Jianjun Zhang 0004, Daniel S. Yeung, Witold Pedrycz
Pattern Recognit.4
2016 Adversarial Feature Selection Against Evasion Attacks
abstract
Pattern recognition and machine learning techniques have been increasingly adopted in adversarial settings such as spam, intrusion, and malware detection, although their security against well-crafted attacks that aim to evade detection by manipulating data at test time has not yet been thoroughly assessed. While previous work has been mainly focused on devising adversary-aware classification algorithms to counter evasion attempts, only few authors have considered the impact of using reduced feature sets on classifier security against the same attacks. An interesting, preliminary result is that classifier security to evasion may be even worsened by the application of feature selection. In this paper, we provide a more detailed investigation of this aspect, shedding some light on the security properties of feature selection against evasion attacks. Inspired by previous work on adversary-aware classifiers, we propose a novel adversary-aware feature selection model that can improve classifier security against evasion attacks, by incorporating specific assumptions on the adversary's data manipulation strategy. We focus on an efficient, wrapper-based implementation of our approach, and experimentally validate its soundness on different application examples, including spam and malware detection.
Fei Zhang 0003, Patrick P. K. Chan, Battista Biggio, Daniel S. Yeung, Fabio Roli
IEEE Trans. Cybern.4
2016 MLPNN Training via a Multiobjective Optimization of Training Error and Stochastic Sensitivity
abstract
The training of a multilayer perceptron neural network (MLPNN) concerns the selection of its architecture and the connection weights via the minimization of both the training error and a penalty term. Different penalty terms have been proposed to control the smoothness of the MLPNN for better generalization capability. However, controlling its smoothness using, for instance, the norm of weights or the Vapnik-Chervonenkis dimension cannot distinguish individual MLPNNs with the same number of free parameters or the same norm. In this paper, to enhance generalization capabilities, we propose a stochastic sensitivity measure (ST-SM) to realize a new penalty term for MLPNN training. The ST-SM determines the expectation of the squared output differences between the training samples and the unseen samples located within their Q -neighborhoods for a given MLPNN. It provides a direct measurement of the MLPNNs output fluctuations, i.e., smoothness. We adopt a two-phase Pareto-based multiobjective training algorithm for minimizing both the training error and the ST-SM as biobjective functions. Experiments on 20 UCI data sets show that the MLPNNs trained by the proposed algorithm yield better accuracies on testing data than several recent and classical MLPNN training methods.
Daniel S. Yeung, Jin-Cheng Li, Wing W. Y. Ng, Patrick P. K. Chan
IEEE Trans. Neural Networks Learn. Syst.1
2015 Spam filtering for short messages in adversarial environment
Patrick P. K. Chan, Daniel S. Yeung, Wing W. Y. Ng
Neurocomputing3
2015 Two-phase mapping hashing
Wing W. Y. Ng, Yueming Lv, Daniel S. Yeung, Patrick P. K. Chan
Neurocomputing3
2015 Diversified Sensitivity-Based Undersampling for Imbalance Classification Problems
abstract
Undersampling is a widely adopted method to deal with imbalance pattern classification problems. Current methods mainly depend on either random resampling on the majority class or resampling at the decision boundary. Random-based undersampling fails to take into consideration informative samples in the data while resampling at the decision boundary is sensitive to class overlapping. Both techniques ignore the distribution information of the training dataset. In this paper, we propose a diversified sensitivity-based undersampling method. Samples of the majority class are clustered to capture the distribution information and enhance the diversity of the resampling. A stochastic sensitivity measure is applied to select samples from both clusters of the majority class and the minority class. By iteratively clustering and sampling, a balanced set of samples yielding high classifier sensitivity is selected. The proposed method yields a good generalization capability for 14 UCI datasets.
Wing W. Y. Ng, Daniel S. Yeung, Shaohua Yin, Fabio Roli
IEEE Trans. Cybern.3
2015 Asymmetric Cyclical Hashing for Large Scale Image Retrieval
abstract
This paper addresses a problem in the hashing technique for large scale image retrieval: learn a compact hash code to reduce the storage cost with performance comparable to that of the long hash code. A longer hash code yields a better precision rate of retrieved images. However, it also requires a larger storage, which limits the number of stored images. Current hashing methods employ the same code length for both queries and stored images. We propose a new hashing scheme using two hash codes with different lengths for queries and stored images, i.e., the asymmetric cyclical hashing. A compact hash code is used to reduce the storage requirement, while a long hash code is used for the query image. The image retrieval is performed by computing the Hamming distance of the long hash code of the query and the cyclically concatenated compact hash code of the stored image to yield a high precision and recall rate. Experiments on benchmarking databases consisting up to one million images show the effectiveness of the proposed method.
Yueming Lv, Wing W. Y. Ng, Ziqian Zeng, Daniel S. Yeung, Patrick P. K. Chan
IEEE Trans. Multim.4
2014 An improved differential evolution and its application to determining feature weights in similarity-based clustering
Chunru Dong, Wing W. Y. Ng, Xizhao Wang, Patrick P. K. Chan, Daniel S. Yeung
Neurocomputing5
2014 LG-Trader: Stock trading decision support based on feature selection by weighted localized generalization error model
Wing W. Y. Ng, Xue-Ling Liang, Jin-Cheng Li, Daniel S. Yeung, Patrick P. K. Chan
Neurocomputing4
2014 Steganalysis classifier training via minimizing sensitivity for different imaging sources
Wing W. Y. Ng, Zhi-Min He, Daniel S. Yeung, Patrick P. K. Chan
Inf. Sci.3
2012 Dynamic fusion method using Localized Generalization Error Model
Patrick P. K. Chan, Daniel S. Yeung, Wing W. Y. Ng, Chih-Min Lin, James Nga-Kwok Liu
Inf. Sci.2
2011 Recent advances on machine learning and Cybernetics
Witold Pedrycz, Daniel S. Yeung, Xizhao Wang
Soft Comput.2
2010 Steganography detection using localized generalization error model
abstract
Steganography detection is a technique to tell whether there are secret messages hidden in images. The performance of a steganalysis system is mainly determined by the method of feature extraction and the architecture selection of the classifier. Selecting a proper classifier with proper parameters will improve the detection accuracy and generalization capability of the system. We propose a Radial Basis Function Neural Network (RBFNN) optimized by the Localized Generalization Error Model (L-GEM) for steganograhpy detection. In the proposed method, the discrete cosine transform (DCT) features and the Markov features are used as inputs of neural networks for detection. To enhance the generalization capability of the RBFNN and the performance of detecting steganography in future images, the architecture of the RBFNN is selected by minimizing the L-GEM. The experimental results show that the proposed method provides a better performance on testing images in comparison with the existing method in attackting Steghide, OutGuess and F5.
Zhi-Min He, Wing W. Y. Ng, Patrick P. K. Chan, Daniel S. Yeung
SMC4
2010 Adaptive filter design using recurrent cerebellar model articulation controller
abstract
A novel adaptive filter is proposed using a recurrent cerebellar-model-articulation-controller (CMAC). The proposed locally recurrent globally feedforward recurrent CMAC (RCMAC) has favorable properties of small size, good generalization, rapid learning, and dynamic response, thus it is more suitable for high-speed signal processing. To provide fast training, an efficient parameter learning algorithm based on the normalized gradient descent method is presented, in which the learning rates are on-line adapted. Then the Lyapunov function is utilized to derive the conditions of the adaptive learning rates, so the stability of the filtering error can be guaranteed. To demonstrate the performance of the proposed adaptive RCMAC filter, it is applied to a nonlinear channel equalization system and an adaptive noise cancelation system. The advantages of the proposed filter over other adaptive filters are verified through simulations.
Chih-Min Lin, Li-Yang Chen, Daniel S. Yeung
IEEE Trans. Neural Networks3
2009 Sensitivity Based Generalization Error for Supervised Learning Problems with Application in Feature Selection
Daniel S. Yeung
ADMA1
2009 IPCM Separability Ratio for Supervised Feature Selection
abstract
Collecting data is very easy now owing to fast computers and ease of Internet access. It raises the problem of the curse of dimensionality to supervised classification problems. In our previous work, an Intra-Prototype / Inter-Class Separability Ratio (IPICSR) model is proposed to select relevant features for semi-supervised classification problems. In this work, a new margin based feature selection model is proposed based on the IPICSR model for supervised classification problems. Owing to the nature of supervised classification problems, a more accurate class separating margin could be found by the classifier. We adopt this advantage in the new Intra-Prototype / Class Margin Separability Ratio (IPCMSR) model. Experimental results are promising when compared to several existing methods using 4 UCI datasets.
Wing W. Y. Ng, Jun Wang 0017, Daniel S. Yeung
SMC3
2009 A Novel Dynamic Fusion Method Using Localized Generalization Error Model
abstract
A new dynamic classifier fusion method named L-GEM fusion method (LFM) for multiple classifier systems (MCSs) is proposed. The localized generalization error upper bound for the neighborhood of a testing sample is calculated and used to estimate the local competence of base classifiers in MCSs. Different from the recent dynamic classifier selection methods, the proposed method consider not only the training error but also the sensitivity of the base classifier. Experimental results show that the MCSs using LFM has more accurate than other popular dynamic fusion methods.
Daniel S. Yeung, Patrick P. K. Chan
SMC1
2009 Radial Basis Function network learning using localized generalization error bound
Daniel S. Yeung, Patrick P. K. Chan, Wing W. Y. Ng
Inf. Sci.1
2008 Quantitative study on candlestick pattern for Shenzhen Stock Market
abstract
Shenzhen Stock Market grows rapidly yet is still a young market when compared with Hong Kong, New York and London markets. Its daily turnover reaches billions US dollars. A good prediction of stock price will bring us substantial pecuniary reward. Technical analysis is a widely adopted financial prediction tool in worldwide stock markets. Candlestick pattern is one of the most efficient methods in technical analysis. However, does candlestick pattern prediction works for Shenzhen stocks? Candlestick patterns are always defined by fuzzy terms, could we have a quantitative definition of these patterns? We perform a quantitative study on these two major research problems in this paper. We study the morning star pattern in this work and the method in this paper could be extended to other patterns easily. So, we propose adopting the Radial Basis Function Neural Networks trained with Localized Generalization Error Model to predict whether or not the stock price will increase after the appearance of it. Then, we extract the patterns from the neural network to provide a quantitative definition of the morning star pattern for a particular stock. Experimental results show that our modification to the morning star pattern prediction prevents up to 69% of false prediction of the morning star pattern. We also provide a quantitative measure of the morning star patterns for two of the Shenzhen stocks.
Huili Li, Wing W. Y. Ng, John W. T. Lee, Binbin Sun, Daniel S. Yeung
SMC5
2008 Information extraction based on information fusion from multiple news sources from the web
abstract
The traditional information extraction tools have been developed for years. But, the accuracy of extraction is not very satisfactory, especially for named entity extraction. In this work, we analyze the reasons of it and propose a novel method to improve the accuracy. Existing methods extract information from a text which is collected from a single source. This is very difficult to extract the exact information we need. From the Internet, one could easily find tens of sources for the same information (e.g. particular news). In this work, we propose to combine information extracted from multiple sources using a majority voting to find the information we needed. We use Change of CEO as an example and we extract the new CEO, original CEO and the company name for the event. A off-the-shelf named entity extraction tool is adopted and our major contribution is the fusion of extraction results. Without our work, one finds single news provides many people and company names, such that we do not know who the new CEO is. By using our method, we provide the 2 CEO names and 1 company name. Experimental results show that our method has a high accuracy in finding the exact information.
Wing W. Y. Ng, John W. T. Lee, Binbin Sun, Daniel S. Yeung
SMC5
2008 Localized generalization error based active learning for image annotation
abstract
Content-based image auto-annotation becomes a hot research topic owing to the development of image retrieval system and the storing technology of multimedia information. It is a key step in most of those image processing applications. In this work, we adopt active learning to image annotation for reducing the number of labeled images required for supervised learning procedure. Localized Generalization Error Model (L-GEM) based active learning uses localized generalization error bound as the sample selection criterion. In each turn, the most informative sample from a set of unlabeled samples is selected by the L-GEM based active learning will be labeled and added to the training dataset. A heuristic and a Q value selection improvement methods are introduced in this paper. The experimental results show that the proposed active learning efficiently reduces the number of labeled training samples. Moreover, the improvement method improve the performances in both testing accuracy and training time which are both essential in image annotation applications.
Binbin Sun, Wing W. Y. Ng, Daniel S. Yeung, Jun Wang 0017
SMC3
2008 A definition of partial derivative of random functions and its application to RBFNN sensitivity analysis
Xizhao Wang, Chun-Guo Li, Daniel S. Yeung, ShiJi Song, Hui-Min Feng
Neurocomputing3
2008 A Novel Feature Grouping Method for Ensemble Neural Network Using Localized Generalization Error Model
abstract
Multiple Classifier System (MCS) is a very popular research topic in recent years. It has been proved theoretically and empirically to be better than single classifiers in many scenarios. Creating diverse sets of classifier is one of the key issues in building MCSs. Feature grouping is one of the methods to create diverse classifiers and it has been shown to improve the accuracy of an MCS. In this paper, we propose a new feature grouping method based on Genetic Algorithm (GA) with the localized Generalization Error Model as the evaluation criterion. The combined individual classifiers using the weighted sum are examined in this paper. Moreover, several feature grouping methods are compared with the proposed method in this work. The experimental results on benchmark dataset show that the MCS trained by the proposed method is promising.
Aki P. F. Chan, Patrick P. K. Chan, Wing W. Y. Ng, Eric C. C. Tsang, Daniel S. Yeung
Int. J. Pattern Recognit. Artif. Intell.5
2008 The Localized Generalization Error Model for Single Layer Perceptron Neural Network and Sigmoid Support Vector Machine
abstract
We had developed the localized generalization error model for supervised learning with minimization of Mean Square Error. In this work, we extend the error model to Single Layer Perceptron Neural Network (SLPNN) and Support Vector Machine (SVM) with sigmoid kernel function. For a trained SLPNN or SVM and a given training dataset, the proposed error model bounds above the error for unseen samples which are similar to the training samples. As the major component of the localized generalization error model, the stochastic sensitivity measure formula for perceptron neural network derived in this work has relaxed the assumptions of same distribution for all inputs and each sample perturbed only once in previous works. These make the sensitivity measure applicable to pattern classification problems. The stochastic sensitivity measure of SVM with Sigmoid kernel is also derived in this work as a component of the localized generalization error model. At the end of this paper, we discuss the advantages of the proposed error bound over existing error bound.
Wing W. Y. Ng, Daniel S. Yeung, Eric C. C. Tsang
Int. J. Pattern Recognit. Artif. Intell.2
2008 Editorial
Xizhao Wang, Yuan Yan Tang, Daniel S. Yeung
Int. J. Pattern Recognit. Artif. Intell.3
2008 Preface: Recent advances in granular computing
Daniel S. Yeung, Xizhao Wang, Degang Chen 0002
Inf. Sci.1
2008 Feature selection using localized generalization error for supervised classification problems using RBFNN
Wing W. Y. Ng, Daniel S. Yeung, Michael Firth, Eric C. C. Tsang, Xizhao Wang
Pattern Recognit.2
2008 Attributes Reduction Using Fuzzy Rough Sets
abstract
Fuzzy rough sets are the generalization of traditional rough sets to deal with both fuzziness and vagueness in data. The existing researches on fuzzy rough sets are mainly concentrated on the construction of approximation operators. Less effort has been put on the attributes reduction of databases with fuzzy rough sets. This paper mainly focuses on the attributes reduction with fuzzy rough sets. After analyzing the previous works on attributes reduction with fuzzy rough sets, we introduce formal concepts of attributes reduction with fuzzy rough sets and completely study the structure of attributes reduction. An algorithm using discernibility matrix to compute all the attributes reductions is developed. Based on these lines of thought, we set up a solid mathematical foundation for attributes reduction with fuzzy rough sets. The experimental results show that the idea in this paper is feasible and valid.
Eric C. C. Tsang, Degang Chen 0002, Daniel S. Yeung, Xizhao Wang, John W. T. Lee
IEEE Trans. Fuzzy Syst.3
2007 Neural network ensemble pruning using sensitivity measure in web applications
abstract
Multiple Classifier Systems (MCSs) have been shown theoretically and empirically to outperform a single classifier in many applications. However, many ensemble training algorithms sometimes create a very large MCS which is combined by many individual classifiers. A large MCS not only consumes computational resources but also decreases the effectiveness. One of the solutions is the pruning method. It reduces the number of individual classifiers inside an MCS that maintains the performance well or is just slightly worse than the original one. In this paper, a new pruning method, called NNEPSM, for Neural Network ensemble based on a sensitivity measure is proposed. The classifiers which have less impact to the final output of MCS will be removed. The advantages of this method include efficient performance, low-complexity and independence on training method. NNEPSM has been applied in Web applications and other benchmark dataset. The experimental results showed that our approach performs well using different datasets.
Patrick P. K. Chan, Xiaoqin Zeng, Eric C. C. Tsang, Daniel S. Yeung, John W. T. Lee
SMC4
2007 Heuristic improvement for active learning using localized generalization error as selection criterion
abstract
Owing to the growth of Internet and computer technology, pattern recognition for large-scale datasets has become one of the hot research topics. The major challenges are to reduce the human efforts involved and to improve the efficiency. Traditional passive learning methods require labeling of all training samples may not be feasible in large-scale recognition problems because of the requirement of large-scale class labeling for the huge number of training samples. In the literatures, there are many studies on active learning methods, which does not require all training samples to be labeled and it selects training samples for labeling based on the knowledge of the current classifier. In this paper, we present an active learning method using localized generalization error of candidate sample as selection criterion. Our method uses the generalization error of candidate sample, so theoretically it should have a better performance than other methods. From the experiment results, our method outperforms other methods in both yielding higher prediction accuracy on testing dataset and selecting fewer training samples. Furthermore, we propose a heuristics improvement based on the Q-neighborhood idea of the localized generalization error model to reduce the number of samples being selected and the computational time.
Wing W. Y. Ng, Binbin Sun, Daniel S. Yeung, Xizhao Wang
SMC3
2007 Stability analysis of a discrete Hopfield neural network with delay
Eric C. C. Tsang, S. S. Qiu, Daniel S. Yeung
Neurocomputing3
2007 Internet Anomaly Detection Based on Statistical Covariance Matrix
abstract
Intrusion detection is an important part of assuring the reliability of computer systems. Different intrusion detection approaches vary with different patterns used and different intrusions addressed. However, what patterns are effective in constructing a detection system is still a challenge. This paper attempts to apply the traditional covariance matrix concept to the detection of multiple known and unknown network anomalies. With respect to the initiation of typical flood-based network intrusions, the proposed approach takes the measure of covariance matrix to reflect the changes of sequential correlativity of the network traffic when flood-based attacks happen. The differences among covariance matrices of network samples collected in temporal sequences of fixed and equal length are directly evaluated to detect multiple network anomalies. Extensive experiments on the subset of KDDCUP 1999 dataset show that the covariance matrix, as a new pattern, can be directly utilized to construct an effective detection system for flood-based attacks. It also points out that utilizing the covariance matrix in the detection of flood-based attacks can achieve higher performance over traditional approaches.
Shuyuan Jin, Daniel S. Yeung, Xizhao Wang
Int. J. Pattern Recognit. Artif. Intell.2
2007 Convergence Analysis of a Discrete Hopfield Neural Network with Delay and its Application to Knowledge Refinement
abstract
This paper investigates the convergence theorems that are associated with a Discrete Hopfield Neural Network (DHNN) with delay. We present two updating rules, one for serial mode and the other for parallel mode. The speed of convergence of these proposed updating rules is faster than all of the existing updating rules. It has been proved in this paper that a DHNN with delay will converge to a stable state when operating in a serial mode if the matrix of weights of the no-delay term is symmetric. In addition, it has been proved that they will converge to a stable state when operating in a parallel mode if the matrix of weights of the no-delay term is a symmetric and non-negative definite matrix. The condition for convergence of a DHNN without delay can been relaxed from the need to have a symmetric matrix to an even weaker condition of having a quasi-symmetric matrix. The results in this paper extend both the existing results concerning the convergence of a DHNN without delay and our previous findings. By means of the new network structure and its convergence theorems, we propose a local searching algorithm for combinatorial optimization. We also relate the maximum value of a bivariate energy function to the stable states of a DHNN with delay, which generalizes Hopfield's energy function. Moreover, for the serial model we give the relationship between the convergence of the energy function and the convergence of the corresponding network. One application is presented to demonstrate the higher rate of convergence and the accuracy of the classification using our algorithm.
Eric C. C. Tsang, S. S. Qiu, Daniel S. Yeung
Int. J. Pattern Recognit. Artif. Intell.3
2007 Learning fuzzy rules from fuzzy samples based on rough set technique
Xizhao Wang, Eric C. C. Tsang, Suyun Zhao, Degang Chen 0002, Daniel S. Yeung
Inf. Sci.5
2007 Structured large margin machines: sensitive to data distributions
Daniel S. Yeung, Defeng Wang, Wing W. Y. Ng, Eric C. C. Tsang, Xizhao Wang
Mach. Learn.1
2007 Network intrusion detection in covariance feature space
Shuyuan Jin, Daniel S. Yeung, Xizhao Wang
Pattern Recognit.2
2007 Image classification with the use of radial basis function neural networks and the minimization of the localized generalization error
Wing W. Y. Ng, Andrés Dorado, Daniel S. Yeung, Witold Pedrycz, Ebroul Izquierdo
Pattern Recognit.3
2007 Ellipsoidal support vector clustering for functional MRI analysis
Defeng Wang, Lin Shi 0001, Daniel S. Yeung, Eric C. C. Tsang, Pheng-Ann Heng
Pattern Recognit.3
2007 Machine Learning Techniques: Problems and Applications. By Guest Editors: Zhi-Qiang Liu, Daniel So Yeung, Xi-Zhao Wang, and Eric Tsang
Xizhao Wang, Daniel S. Yeung
Soft Comput.3
2007 Localized Generalization Error of Gaussian-based Classifiers and Visualization of Decision Boundaries
Wing W. Y. Ng, Daniel S. Yeung, Defeng Wang, Eric C. C. Tsang, Xizhao Wang
Soft Comput.2
2007 Weighted Mahalanobis Distance Kernels for Support Vector Machines
abstract
The support vector machine (SVM) has been demonstrated to be a very effective classifier in many applications, but its performance is still limited as the data distribution information is underutilized in determining the decision hyperplane. Most of the existing kernels employed in nonlinear SVMs measure the similarity between a pair of pattern images based on the Euclidean inner product or the Euclidean distance of corresponding input patterns, which ignores data distribution tendency and makes the SVM essentially a "local" classifier. In this paper, we provide a step toward a paradigm of kernels by incorporating data specific knowledge into existing kernels. We first find the data structure for each class adaptively in the input space via agglomerative hierarchical clustering (AHC), and then construct the weighted Mahalanobis distance (WMD) kernels using the detected data distribution information. In WMD kernels, the similarity between two pattern images is determined not only by the Mahalanobis distance (MD) between their corresponding input patterns but also by the sizes of the clusters they reside in. Although WMD kernels are not guaranteed to be positive definite (pd) or conditionally positive definite (cpd), satisfactory classification results can still be achieved because regularizers in SVMs with WMD kernels are empirically positive in pseudo-Euclidean (pE) spaces. Experimental results on both synthetic and real-world data sets show the effectiveness of "plugging" data structure into existing kernels.
Defeng Wang, Daniel S. Yeung, Eric C. C. Tsang
IEEE Trans. Neural Networks2
2007 Localized Generalization Error Model and Its Application to Architecture Selection for Radial Basis Function Neural Network
abstract
The generalization error bounds found by current error models using the number of effective parameters of a classifier and the number of training samples are usually very loose. These bounds are intended for the entire input space. However, support vector machine (SVM), radial basis function neural network (RBFNN), and multilayer perceptron neural network (MLPNN) are local learning machines for solving problems and treat unseen samples near the training samples to be more important. In this paper, we propose a localized generalization error model which bounds from above the generalization error within a neighborhood of the training samples using stochastic sensitivity measure. It is then used to develop an architecture selection technique for a classifier with maximal coverage of unseen samples by specifying a generalization error threshold. Experiments using 17 University of California at Irvine (UCI) data sets show that, in comparison with cross validation (CV), sequential learning, and two other ad hoc methods, our technique consistently yields the best testing classification accuracy with fewer hidden neurons and less training time.
Daniel S. Yeung, Wing W. Y. Ng, Defeng Wang, Eric C. C. Tsang, Xizhao Wang
IEEE Trans. Neural Networks1
2007 Covariance-Matrix Modeling and Detecting Various Flooding Attacks
abstract
This paper presents a covariance-matrix modeling and detection approach to detecting various flooding attacks. Based on the investigation of correlativity changes of monitored network features during flooding attacks, this paper employs statistical covariance matrices to build a norm profile of normal activities in information systems and directly utilizes the changes of covariance matrices to detect various flooding attacks. The classification boundary is constrained by a threshold matrix, where each element evaluates the degree to which an observed covariance matrix is different from the norm profile in terms of the changes of correlation between the monitored network features represented by this element. Based on Chebyshev inequality theory, we give a practical (heuristic) approach to determining the threshold matrix. Furthermore, the result matrix obtained in the detection serves as the second-order features to characterize the detected flooding attack. The performance of the approach is examined by detecting Neptune and Smurf attacks-two common distributed Denial-of-Service flooding attacks. The evaluation results show that the detection approach can accurately differentiate the flooding attacks from the normal traffic. Moreover, we demonstrate that the system extracts a stable set of the second-order features for these two flooding attacks
Daniel S. Yeung, Shuyuan Jin, Xizhao Wang
IEEE Trans. Syst. Man Cybern. Part A1
2006 Bankruptcy Prediction Using Multiple Classifier System with Mutual Information Feature Grouping
abstract
The prediction of bankruptcy helps an organization to choose its business partners and banks to approve or reject loan requests. So, it is essential to predict the bankruptcy of an organization. In this work, a multiple classifier system which combines decision from several different base classifiers trained using different samples with different input features is proposed for the bankruptcy prediction. The input features for each base classifier are selected using its mutual information with respect to the output. Experimental results of the proposed method using a real bankruptcy dataset from Compustat Global Dataset are promising.
Aki P. F. Chan, Wing W. Y. Ng, Daniel S. Yeung, Eric C. C. Tsang, Michael Firth
SMC3
2006 Structured Large Margin Machine Ensemble
abstract
Large margin classifiers have been widely applied in solving supervised learning problems. One representative model in large margin learning is the support vector machine (SVM). SVM is an unstructured classifier since the data structure information is underutilized and the decision hyperplane calculation relies exclusively on the support vectors. To incorporate the data covariance information into the large margin learning, structured large margin machine (SLMM) is recently proposed and show better performance than classical SVM in some applications. Instead of utilizing the data structures straightly like SLMM, SVM ensemble (SVMe) improves the generalization ability of SVM in another way by combining the outputs of a series of SVMs. Inspired by SVMe, we are going to explore the ensemble counterpart for SLMM, i.e., SLMMe, and validate the effectiveness of multiple SLMM system. Experimental results on benchmark datasets demonstrate that SLMMe improves SLMM by reducing its variance, and SLMMe outperforms SVMe in most cases in terms of both classification accuracy and variance.
Patrick P. K. Chan, Defeng Wang, Eric C. C. Tsang, Daniel S. Yeung
SMC4
2006 Learning Web Categorization with Controlled Generation of Context Features
abstract
Automatic categorization of Web pages is an important area of study due to the rapidly growing amount of Web data. Efficient and accurate classification would greatly facilitate finding what one needs in the sea of information. Context-sensitive techniques have been proven to be effective in the classification task. However, the feature space for context feature that one can explore in these techniques is enormous. To consider these features comprehensively often become prohibitive in terms of resource requirements. In this paper, we propose an approach to intelligently control generating context features for the classification learning process. We present our investigation of this approach in the context of Web page categorization using the sleeping-experts technique.
Alex K. S. Wong, John W. T. Lee, Daniel S. Yeung
SMC3
2006 Hidden neuron pruning of multilayer perceptrons using a quantified sensitivity measure
Xiaoqin Zeng, Daniel S. Yeung
Neurocomputing2
2006 Rough approximations on a complete completely distributive lattice with applications to generalized rough sets
Degang Chen 0002, Wen-Xiu Zhang, Daniel S. Yeung, Eric C. C. Tsang
Inf. Sci.3
2006 Computation of Madalines' Sensitivity to Input and Weight Perturbations
abstract
The sensitivity of a neural network's output to its input and weight perturbations is an important measure for evaluating the network's performance. In this letter, we propose an approach to quantify the sensitivity of Madalines. The sensitivity is defined as the probability of output deviation due to input and weight perturbations with respect to overall input patterns. Based on the structural characteristics of Madalines, a bottom-up strategy is followed, along which the sensitivity of single neurons, that is, Adalines, is considered first and then the sensitivity of the entire Madaline network. By means of probability theory, an analytical formula is derived for the calculation of Adalines' sensitivity, and an algorithm is designed for the computation of Madalines' sensitivity. Computer simulations are run to verify the effectiveness of the formula and algorithm. The simulation results are in good agreement with the theoretical results.
Yingfeng Wang, Xiaoqin Zeng, Daniel S. Yeung, Zhihang Peng
Neural Comput.3
2006 Rough sets and ordinal reducts
John W. T. Lee, Daniel S. Yeung, Eric C. C. Tsang
Soft Comput.2
2006 Preface for the special issue: soft computing in machine learning and cybernetics in the journal Soft Computing
Daniel S. Yeung, Xizhao Wang, Eric C. C. Tsang, John W. T. Lee
Soft Comput.2
2006 Applying undistorted neural network sensitivity analysis in iris plant classification and construction productivity prediction
Daniel S. Yeung, Wing W. Y. Ng
Soft Comput.2
2006 Structured One-Class Classification
abstract
The one-class classification problem aims to distinguish a target class from outliers. The spherical one-class classifier (SOCC) solves this problem by finding a hypersphere with minimum volume that contains the target data while keeping outlier samples outside. SOCC achieves satisfactory performance only when the target samples have the same distribution tendency in all orientations. Therefore, the performance of the SOCC is limited in the way that many superfluous outliers might be mistakenly enclosed. The authors propose to exploit target data structures obtained via unsupervised methods such as agglomerative hierarchical clustering and use them in calculating a set of hyperellipsoidal separating boundaries. This method is named the structured one-class classifier (TOCC). The optimization problem in TOCC can be formulated as a series of second-order cone programming problems that can be solved with acceptable efficiency by primal-dual interior-point methods. The experimental results on artificially generated data sets and benchmark data sets demonstrate the advantages of TOCC.
Defeng Wang, Daniel S. Yeung, Eric C. C. Tsang
IEEE Trans. Syst. Man Cybern. Part B2
2005 Multiple Classifier System with Feature Grouping for Intrusion Detection: Mutual Information Approach
Aki P. F. Chan, Wing W. Y. Ng, Daniel S. Yeung, Eric C. C. Tsang
KES (3)3
2005 Support Vector Clustering for Brain Activation Detection
Defeng Wang, Lin Shi 0001, Daniel S. Yeung, Pheng-Ann Heng, Tien-Tsin Wong, Eric C. C. Tsang
MICCAI3
2005 Active learning using localized generalization error of candidate sample as criterion
abstract
In classification problem, the learning process can be more efficient if the informative samples can be selected actively based on the knowledge of the classifier. This problem is called active learning. Most of the existing active learning methods did not directly relate to the generalization error of classifiers. Also, some of them need high computational time or are based on strict assumptions. This paper describes a new active learning strategy using the concept of localized generalization error of the candidate samples. The sample which yields the largest generalization error will be chosen for query. This method can be applied to different kinds of classifiers and its complexity is low. Experimental results demonstrate that the prediction accuracy of the classifier can be improved by using this selecting method and fewer training samples are possible for the same prediction accuracy.
Patrick P. K. Chan, Wing W. Y. Ng, Daniel S. Yeung
SMC3
2005 A feature space analysis for anomaly detection
abstract
Intrusion detection is an important part of assuring the reliability of computer systems. From the viewpoint of feature space partition of detectors, this paper investigates one of the limitations of two traditional anomaly detection technologies - NN-based anomaly detection and statistical detection approaches in detecting novel attacks. A high dimensional covariance matrix feature space and an on-line detection algorithm are proposed to detect various known and unknown attacks. An illustrative example of detecting various known and unknown probing attacks is provided.
Shuyuan Jin, Daniel S. Yeung, Xizhao Wang, Eric C. C. Tsang
SMC2
2005 Quantitative study on the generalization error of multiple classifier systems
abstract
Multiple classifier system (MCS) has been one of the hot research topics in machine learning field. A MCS merges an ensemble of different or same type of classifiers together to enhance the problem solving performance of machine learning. However, the choice of the number of classifiers and the fusion method are usually based on ad-hoc selection. In this paper, we propose a novel quantitative measure of the generalization error for MCS. The localized generalization error model bounds above the mean square error (MSE) of a MCS for unseen samples located within a neighborhood of the training samples. The relationship between the proposed model and classification accuracy is also discussed in this paper. This model quantitatively measures the goodness of the MCS in approximating the unknown input-output mapping hidden in the training dataset The localized generalization error model is applied to select a MCS, among different choices of number of classifiers and fusion methods, for a given classification problem. Experimental results on three real world datasets are performed to show promising results.
Wing W. Y. Ng, Aki P. F. Chan, Daniel S. Yeung, Eric C. C. Tsang
SMC3
2005 On attributes reduction with fuzzy rough sets
abstract
Fuzzy rough set is the generalization of Pawlak rough set to deal with both of fuzziness and vagueness in data. The existing research on fuzzy rough sets mainly concentrates on the construction of approximation operators. Less effort has been put on the attributes reduction of fuzzy rough sets. This paper mainly focuses on the attributes reduction of fuzzy rough sets. After analyzing the previous work of reduction of fuzzy rough sets, formal concepts of attributes reduction of fuzzy rough sets are introduced and the structure of reduction is studied completely. An algorithm based on discernibility matrix to compute all the attributes reductions is developed. These concepts have been demonstrated by an example.
Gloria C. Y. Tsang, Degang Chen 0002, Eric C. C. Tsang, John W. T. Lee, Daniel S. Yeung
SMC5
2005 Stability Analysis of Discrete Hopfield Neural Networks With Delay and Its Application
abstract
Discrete Hopfield neural networks (DHNNs) with delay, which can deal with temporal information, are a generalization of the DHNNs without delay. This paper investigates the convergence theorems in DHNNs with delay. We present two generalized updating rules, one for serial mode and the other for parallel mode. The convergence speed of these proposed updating rules is faster than existing updating rules. By means of the new network structure and its convergence theorems, we propose a local searching algorithm for combinatorial optimization. We also relate the maximum value of a bivariate energy function to the stable states of the DHNNs with delay. Furthermore, we describe an algorithm for the DHNNs with delay in which the delay term is regarded as noise, which has a higher convergence rate than usual algorithms in the Hopfield neural network without delay. One application is presented to demonstrate the higher rate of convergence of our algorithm
Eric C. C. Tsang, Aki P. F. Chan, Daniel S. Yeung, S. S. Qiu
SMC3
2005 A genetic algorithm for solving the inverse problem of support vector machines
Xizhao Wang, Qiang He 0003, Degang Chen 0002, Daniel S. Yeung
Neurocomputing4
2005 Combining multiple classifiers based on a statistical method for handwritten Chinese character recognition
abstract
Combining multiple classifiers is a new method that achieves a substantial gain in performance in many areas of pattern recognition. This paper demonstrates a novel method (based on statistics) of combining multiple classifiers to address the task of recognizing handwritten Chinese characters. Fusion strategies are discussed to provide a basis for the architecture of the combined classifiers. The weights of these fusion strategies are assigned via a genetic algorithm (GA). These fusion strategies are then tested using our online system for handwritten Chinese character recognition. In addition, different combinatory approaches are tested for comparison purposes. These include the conventional approach that is based on the Bayesian principle and the improved weighted combination, employing shared and distinct representations. Our experimental results demonstrate the effectiveness of these combinatory approaches.
Lei Lin 0001, Xiaolong Wang 0001, Daniel S. Yeung
Int. J. Pattern Recognit. Artif. Intell.3
2005 A Hybrid Language Model Based On Statistics And Linguistic Rules
abstract
Language modeling is a current research topic in many domains including speech recognition, optical character recognition, handwriting recognition, machine translation and spelling correction. There are two main types of language models, the mathematical and the linguistic. The most widely used mathematical language model is the n-gram model inferred from statistics. This model has three problems: long distance restriction, recursive nature and partial language understanding. Language models based on linguistics present many difficulties when applied to large scale real texts. We present here a new hybrid language model that combines the advantages of the n-gram statistical language model with those of a linguistic language model which makes use of grammatical or semantic rules. Using suitable rules, this hybrid model can solve problems such as long distance restriction, recursive nature and partial language understanding. The new language model has been effective in experiments and has been incorporated in Chinese sentence input products for Windows and Macintosh OS.
Xiaolong Wang 0001, Daniel S. Yeung, James Nga-Kwok Liu, Robert Wing Pong Luk, Xuan Wang 0002
Int. J. Pattern Recognit. Artif. Intell.2
2005 A hybrid post-processing system for offline handwritten chinese character recognition based on a statistical language model
abstract
This paper presents a post-processing system for improving the recognition rate of a Handwritten Chinese Character Recognition (HCCR) device. This three-stage hybrid post-processing system reduces the misclassification and rejection rates common in the single character recognition phase. The proposed system is novel in two respects: first, it reduces the misclassification rate by applying a dictionary-look-up strategy that bind the candidate characters into a word-lattice and appends the linguistic-prone characters into the candidate set; second, it identifies promising sentences by employing a distant Chinese word BI-Gram model with a maximum distance of three to select plausible words from the word-lattice. These sentences are then output as the upgraded result. Compared with one of our previous works in single Chinese character recognition, the proposed system improves absolute recognition rates by 12%.
Daniel S. Yeung, Daming Shi 0001
Int. J. Pattern Recognit. Artif. Intell.2
2005 FNDS: a dialogue-based system for accessing digested financial news
Kwok Cheung Lan, Edward Kei Shiu Ho, Robert Wing Pong Luk, Daniel S. Yeung
J. Syst. Softw.4
2005 Sensitivity analysis applied to the construction of radial basis function networks
Daming Shi 0001, Daniel S. Yeung, Junbin Gao
Neural Networks2
2005 A multiple classifier approach to detect Chinese character recognition errors
Kei Yuen Hung, Robert Wing Pong Luk, Daniel S. Yeung, Korris Fu-Lai Chung, Wenhao Shu
Pattern Recognit.3
2005 Hierarchical clustering based on ordinal consistency
John W. T. Lee, Daniel S. Yeung, Eric C. C. Tsang
Pattern Recognit.2
2005 On the generalization of fuzzy rough sets
abstract
Rough sets and fuzzy sets have been proved to be powerful mathematical tools to deal with uncertainty, it soon raises a natural question of whether it is possible to connect rough sets and fuzzy sets. The existing generalizations of fuzzy rough sets are all based on special fuzzy relations (fuzzy similarity relations, T-similarity relations), it is advantageous to generalize the fuzzy rough sets by means of arbitrary fuzzy relations and present a general framework for the study of fuzzy rough sets by using both constructive and axiomatic approaches. In this paper, from the viewpoint of constructive approach, we first propose some definitions of upper and lower approximation operators of fuzzy sets by means of arbitrary fuzzy relations and study the relations among them, the connections between special fuzzy relations and upper and lower approximation operators of fuzzy sets are also examined. In axiomatic approach, we characterize different classes of generalized upper and lower approximation operators of fuzzy sets by different sets of axioms. The lattice and topological structures of fuzzy rough sets are also proposed. In order to demonstrate that our proposed generalization of fuzzy rough sets have wider range of applications than the existing fuzzy rough sets, a special lower approximation operator is applied to a fuzzy reasoning system, which coincides with the Mamdani algorithm.
Daniel S. Yeung, Degang Chen 0002, Eric C. C. Tsang, John W. T. Lee, Xizhao Wang
IEEE Trans. Fuzzy Syst.1
2004 A covariance analysis model for DDoS attack detection
abstract
This paper discusses the effects of multivariate correlation analysis on the DDoS detection and proposes an example, a covariance analysis model for detecting SYN flooding attacks. The simulation results show that this method is highly accurate in detecting malicious network traffic in DDoS attacks of different intensities. This method can effectively differentiate between normal and attack traffic. Indeed, this method can detect even very subtle attacks only slightly different from the normal behaviors. The linear complexity of the method makes its real time detection practical. The covariance model in this paper to some extent verifies the effectiveness of multivariate correlation analysis for DDoS detection. Some open issues still exist in this model for further research.
Shuyuan Jin, Daniel S. Yeung
ICC2
2004 Feature Selection For Chinese Character Recognition Based On Inductive Learning
abstract
Feature selection is a difficult but important issue in the field of machine learning and pattern recognition. In this paper, features for Chinese character recognition are selected by using inductive learning algorithms. The existing inductive learning method based on extension matrix requires precise consistency between positive example and negative example sets, which is very difficult to maintain in most practical cases. The traditional decision tree algorithm ID3 considers only the performance of the discriminating power while selecting features. However, in actual practice the consideration of the associated cost of feature extraction may become a significant concern. In addressing these problems we propose a modified extension matrix approach to select feature subset from the training example set with noises. A decision tree algorithm based on information gain and cost evaluation is also proposed to facilitate cost consideration. The comparative experiments show that the proposed algorithms perform better than the existing inductive learning algorithms to a certain extent.
Guoliang Qian, Daniel S. Yeung, Eric C. C. Tsang, Wenhao Shu
Int. J. Pattern Recognit. Artif. Intell.2
2004 Refinement of generated fuzzy production rules by using a fuzzy neural network
abstract
Fuzzy production rules (FPRs) have been used for years to capture and represent fuzzy, vague, imprecise and uncertain domain knowledge in many fuzzy systems. There have been a lot of researches on how to generate or obtain FPRs. There exist two methods to obtain FPRs. One is by painstakingly, repeatedly and time-consuming interviewing domain experts to extract the domain knowledge. The other is by using some machine learning techniques to generate and extract FPRs from some training samples. These extracted rules, however, are found to be nonoptimal and sometimes redundant. Furthermore, these generated rules suffer from the problem of low accuracy of classifying or recognizing unseen examples. The reasons for having these problems are 1) the FPRs generated are not powerful enough to represent the domain knowledge, 2) the techniques used to generate FPRs are pre-matured, ad-hoc or may not be suitable for the problem, and 3) further refinement of the extracted rules has not been done. In this paper we look into the solutions of the above problems by 1) enhancing the representation power of FPRs by including local and global weights, 2) developing a fuzzy neural network (FNN) with enhanced learning algorithm, and 3) using this FNN to refine the local and global weights of FPRs. By experimenting our method with some existing benchmark examples, the proposed method is found to have high accuracy in classifying unseen samples without increasing the number of the FPRs extracted and the time required to consult with domain experts is greatly reduced.
Eric C. C. Tsang, Daniel S. Yeung, John W. T. Lee, D. M. Huang, X. Z. Wang
IEEE Trans. Syst. Man Cybern. Part B2
2004 Mining Pinyin-to-character conversion rules from large-scale corpus: a rough set approach
abstract
This paper introduces a rough set technique for solving the problem of mining Pinyin-to-character (PTC) conversion rules. It first presents a text-structuring method by constructing a language information table from a corpus for each pinyin, which it will then apply to a free-form textual corpus. Data generalization and rule extraction algorithms can then be used to eliminate redundant information and extract consistent PTC conversion rules. The design of our model also addresses a number of important issues such as the long-distance dependency problem, the storage requirements of the rule base, and the consistency of the extracted rules, while the performance of the extracted rules as well as the effects of different model parameters are evaluated experimentally. These results show that by the smoothing method, high precision conversion (0.947) and recall rates (0.84) can be achieved even for rules represented directly by pinyin rather than words. A comparison with the baseline tri-gram model also shows good complement between our method and the tri-gram language model.
Xiaolong Wang 0001, Qingcai Chen, Daniel S. Yeung
IEEE Trans. Syst. Man Cybern. Part B3
2004 Handling interaction in fuzzy production rule reasoning
abstract
When fuzzy production rules are used to approximate reasoning, interaction exists among rules that have the same consequent. Due to this interaction, the weighted average model frequently used in approximate reasoning does not work well in many real-world problems. In order to model and handle this interaction, this paper proposes to use a nonadditive nonnegative set function to replace the weights assigned to rules having the same consequent, and to draw the reasoning conclusion based on an integral with respect to the nonadditive nonnegative set function, rather than on the weighted average model. Handling interaction in fuzzy production rule reasoning in this way can lead to a good understanding of the rules base and an improvement of reasoning accuracy. This paper also investigates how to determine from data the nonadditive set function that cannot be specified by a domain expert.
Daniel S. Yeung, Xizhao Wang, Eric C. C. Tsang
IEEE Trans. Syst. Man Cybern. Part B1
2003 Monotonic decision tree for ordinal classification
abstract
While ordinal classification problems are common in many situations, induction of ordinal decision trees has not been very extensiveness studied. They are commonly treated as nominal classification problem or regression problem in tree induction. On the other hand a monotonic decision tree is often desirable to aid decision making in such situations as credit rating and student admission. This paper proposes a novel approach called MDT to monotonic decision tree induction. Experiments show that generally this new approach produces decision trees that are more succinct and more effective predictors of the original implicit ordering, apart from being monotonic.
John W. T. Lee, Daniel S. Yeung, Xizhao Wang
SMC2
2003 Input sample selection for RBF neural network classification problems using sensitivity measure
abstract
Large data sets containing irrelevant or redundant input samples reduce the performance of learning and increases storage and labeling costs. This work compares several sample selection and active learning techniques and proposes a novel sample selection method based on the stochastic radial basis function neural network sensitivity measure (SM). The experimental results for the UCI IRIS data set show that we can remove 99% of data while keeping 95% of classification accuracy when applying both sensitivity based feature and sample selection methods. We propose a single and consistent method, which is robust enough to handle both feature and sample selection for a supervised RBFNN classification system, by using the same neural network architecture for both selection and classification tasks.
Wing W. Y. Ng, Daniel S. Yeung, Ian Cloete
SMC2
2003 Determining the relevance of input features for multilayer perceptrons
abstract
This paper presents an approach to determine the relevance of individual input attributes for trained Multilayer Perceptrons (MLPs). To reflect the impact of an input attribute on the output of an MLP, the relevance is aimed at representing the output sensitivity of the MLP to the attribute variation. The sensitivity is defined as the mathematical expectation of output deviations of an MLP due to its input deviation with respect to overall input patterns. The basic idea for the introduction of such a relevance measure is that a well-trained MLP can capture salient features of the problem it deals with and thus become more sensitive to those input attributes that make more contributions to the MLP's behavior. The relevance can be employed as a relative criterion for assessing individual input attributes. The results from the experiments on two typical problems demonstrate the effectiveness of the relevance in identifying irrelevant input attribute.
Xiaoqin Zeng, Yajuan Huang, Daniel S. Yeung
SMC3
2003 An Approach To Natural Stroke Extraction For Off-Line Loosely-Constrained Handwritten Chinese Characters
abstract
This paper proposes a new approach to extracting natural strokes from the skeletons of loosely-constrained, off-line handwritten Chinese characters. It admits the output substrokes from a previously proposed fuzzy substroke extractor as its inputs. By identifying a number of expected ambiguities which include mutual similarities, unstable touches and joint/cross distortions, fuzzy stroke models are constructed and a "hit-all" fuzzy stroke matching strategy is pursued. Fuzzy partitioning technique is used to generate a ranked list of consistent stroke sets from the set of fuzzy strokes being identified. With this approach, a maximum of 20 distinct natural stroke classes can be extracted from each input character, together with an estimate on the actual count of strokes which compose the character. Our system offers a number of performance tuning capabilities such as the computation of the fuzzy scores of each extracted stroke, the adjustment on the fuzzy stroke model parameters, and the potential of incorporating one's personal writing styles into our methodology.
Daniel S. Yeung, Hank-Shun Fong, Eric C. C. Tsang, Wenhao Shu, Xiaolong Wang 0001
Int. J. Pattern Recognit. Artif. Intell.1
2003 A Quantified Sensitivity Measure for Multilayer Perceptron to Input Perturbation
abstract
The sensitivity of a neural network's output to its input perturbation is an important issue with both theoretical and practical values. In this article, we propose an approach to quantify the sensitivity of the most popular and general feedforward network: multilayer perceptron (MLP). The sensitivity measure is defined as the mathematical expectation of output deviation due to expected input deviation with respect to overall input patterns in a continuous interval. Based on the structural characteristics of the MLP, a bottom-up approach is adopted. A single neuron is considered first, and algorithms with approximately derived analytical expressions that are functions of expected input deviation are given for the computation of its sensitivity. Then another algorithm is given to compute the sensitivity of the entire MLP network. Computer simulations are used to verify the derived theoretical formulas. The agreement between theoretical and experimental results is quite good. The sensitivity measure can be used to evaluate the MLP's performance.
Xiaoqin Zeng, Daniel S. Yeung
Neural Comput.2
2003 ASAB: a Chinese screen reader
abstract
Abstract This paper describes the design and development of a computer interface for blind and visually‐impaired users, who are native speakers of Cantonese (i.e. a Chinese dialect). Apart from enabling the interface to (1) produce Chinese voice output, (2) convert Chinese characters to Braille codes, (3) facilitate Chinese Braille input, and (4) operate in a Microsoft Chinese Windows environment, the significant aspects of this paper include the following: (1) the description of an integrated architecture, which can be used for other languages; (2) a general bilingual Braille input mechanism; (3) a sentence‐based input method that can be used for contracted‐Braille‐to‐text conversion with an error rate of about 6%, operating at about 700 characters/second using a Pentium II 300 MHz PC; (4) a code‐mixed synthesis module for general bilingual and multilingual applications; (5) the potential to directly adopt the system for use with other ideographic languages (like Japanese and Korean), as well as agglutinating languages like Finnish and Turkish, which have no space between words. Copyright © 2003 John Wiley & Sons, Ltd.
Robert Wing Pong Luk, Daniel S. Yeung, Qin Lu 0001, H. L. Leung, S. Y. Li, Fred Leung
Softw. Pract. Exp.2
2003 OFFSS: optimal fuzzy-valued feature subset selection
abstract
Feature subset selection is a well-known pattern recognition problem, which aims to reduce the number of features used in classification or recognition. This reduction is expected to improve the performance of classification algorithms in terms of speed, accuracy and simplicity. Most existing feature selection investigations focus on the case that the feature values are real or nominal, very little research is found to address the fuzzy-valued feature subset selection and its computational complexity. This paper focuses on a problem called optimal fuzzy-valued feature subset selection (OFFSS), in which the quality-measure of a subset of features is defined by both the overall overlapping degree between two classes of examples and the size of feature subset. The main contributions of this paper are that: 1) the concept of fuzzy extension matrix is introduced; 2) the computational complexity of OFFSS is proved to be NP-hard; 3) a simple but powerful heuristic algorithm for OFFSS is given; and 4) the feasibility and simplicity of the proposed algorithm are demonstrated by applications of OFFSS to fuzzy decision tree induction and by comparisons with three different feature selection techniques developed recently.
Eric C. C. Tsang, Daniel S. Yeung, Xizhao Wang
IEEE Trans. Fuzzy Syst.2
2002 Improving Performance of Similarity-Based Clustering by Feature Weight Learning
abstract
Similarity-based clustering is a simple but powerful technique which usually results in a clustering graph for a partitioning of threshold values in the unit interval. The guiding principle of similarity-based clustering is "similar objects are grouped in the same cluster." To judge whether two objects are similar, a similarity measure must be given in advance. The similarity measure presented in the paper is determined in terms of the weighted distance between the features of the objects. Thus, the clustering graph and its performance (which is described by several evaluation indices defined in the paper) will depend on the feature weights. The paper shows that, by using gradient descent technique to learn the feature weights, the clustering performance can be significantly improved. It is also shown that our method helps to reduce the uncertainty (fuzziness and nonspecificity) of the similarity matrix. This enhances the quality of the similarity-based decision making.
Daniel S. Yeung, Xizhao Wang
IEEE Trans. Pattern Anal. Mach. Intell.1
2002 Ordinal fuzzy sets
abstract
Fuzzy set theory has been used as a framework for interpreting imprecise linguistic expressions. In general, a linguistic term is described by the compatibility ordering induced in some universe of discourse (UoD). A membership function in fuzzy set theory serves to reflect this ordering by assignment of values in [0, 1] for objects in UoD. When we compute the meaning of a linguistic expression such as "young and tall" using fuzzy membership functions, two implicit assumptions are made. First, we assume the membership values have quantitative meaning so that they can be quantitatively manipulated, for example, by adding or subtracting (the extensive scale assumption). Second, we assume that the scales of the membership values used in describing the different linguistic terms are comparable and the same (the common scale assumption). In many cases, these assumptions cannot be justified. Some proposals have been made to address the first issue by using ordinal scale in defining fuzzy membership functions. However, the second issue has not been properly investigated. In this paper, we propose a framework that does not depend on both of these assumptions. Such framework will facilitate our understanding and investigation of qualitative reasoning without the extensive scale and common scale assumptions.
John W. T. Lee, Daniel S. Yeung, Eric C. C. Tsang
IEEE Trans. Fuzzy Syst.2
2002 Using function approximation to analyze the sensitivity of MLP with antisymmetric squashing activation function
abstract
Sensitivity analysis on a neural network is mainly investigated after the network has been designed and trained. Very few have considered this as a critical issue prior to network design. Piche's statistical method (1992, 1995) is useful for multilayer perceptron (MLP) design, but too severe limitations are imposed on both input and weight perturbations. This paper attempts to generalize Piche's method by deriving an universal expression of MLP sensitivity for antisymmetric squashing activation functions, without any restriction on input and output perturbations. Experimental results which are based on, a three-layer MLP with 30 nodes per layer agree closely with our theoretical investigations. The effects of the network design parameters such as the number of layers, the number of neurons per layer, and the chosen activation function are analyzed, and they provide useful information for network design decision-making. Based on the sensitivity analysis of MLP, we present a network design method for a given application to determine the network structure and estimate the permitted weight range for network training.
Daniel S. Yeung, Xuequan Sun
IEEE Trans. Neural Networks1
2002 Tuning certainty factor and local weight of fuzzy production rules by using fuzzy neural network
abstract
Approximate reasoning in a fuzzy system is concerned with inferring an approximate conclusion from fuzzy and vague inputs. There are many ways in which different forms of conclusions can be drawn. Fuzzy sets are usually represented by fuzzy membership functions. These membership functions are assumed to have a clearly defined base. For other fuzzy sets such as intelligent, smart, or beautiful, etc., it would be difficult to define clearly its base because its base may consist of several other fuzzy sets or unclear nonfuzzy bases. A method to handle this kind of fuzzy set is proposed. A fuzzy neural network (FNN) is also proposed to tune knowledge representation parameters (KRPs). The contributions are that we are able to handle a broader range of fuzzy sets and build more powerful fuzzy systems so that the conclusions drawn are more meaningful, reliable, and accurate. An experiment is presented to demonstrate how our method works.
Eric C. C. Tsang, John W. T. Lee, Daniel S. Yeung
IEEE Trans. Syst. Man Cybern. Part B3
2001 A General Updating Rule for Discrete Hopfield-Type Neural Network with Delay
Shenshan Qiu, Eric C. C. Tsang, Daniel S. Yeung, Xizhao Wang
IJCAI3
2001 Sensitivity Analysis of Multilayer Perceptron
Daniel S. Yeung, Xuequan Sun, Xiaoqin Zeng
IJCAI1
2001 Learning weights of fuzzy production rules by a max-min neural network
Eric C. C. Tsang, Daniel S. Yeung, Xizhao Wang
SMC2
2001 Using neural network classifier in post-processing system for handwritten Chinese character recognition
abstract
A novel post-processing system for a handwritten Chinese character recognition system based on a neural network classifier is presented. The recognition results for input character images, namely candidate characters and their confidence scores, as the observed features of the recognizer are classified into the most probable characters. The confusing character set is established by analyzing large-scale recognition experimental results, and the statistical characteristics for a recognizer are expressed as confusing character sets. 3755 character categories in the GB2312-80 character-set are clustered into several hundreds of groups through searching the transitive closure of the similarity matrix associated with the confusing characters of each character category. A group of neural networks for these category groups is established and trained to be a classifier in the post-processing to recover the unrecognized characters and adjust confidence scores of the candidate characters when a candidate sequence for each individual character image is given. The experimental results show that an average accuracy rate improvement of 5.6% and 3.8% for an online and an offline handwritten Chinese character recognition system are achieved respectively.
Ruifeng Xu 0001, Daniel S. Yeung, Xizhao Wang
SMC2
2001 From global weight to fuzzy measure: handling interaction among fuzzy rules
abstract
Global weight is one of the knowledge representation parameters which is assigned to a set of fuzzy production rules for improving the representation accuracy and reducing the occurrence of incorrect inferences of a fuzzy production rule. Due to the inherent interaction among the rules, the fuzzy inferencing mechanism involving global weights performs unsatisfactorily. To handle this interaction, the paper proposes the use of a fuzzy measure (or in general a nonnegative and nonadditive set function) to replace global weights. Such replacement can effectively improve the reasoning results. An initial experimental result shows that, by learning the fuzzy measure, the reasoning accuracy can be improved significantly.
Daniel S. Yeung, John W. T. Lee, Minghu Ha 0001
SMC1
2001 Transferring Case Knowledge to Adaptation Knowledge: An Approach for Case-Base Maintenance
abstract
In this article we propose a case‐base maintenance methodology based on the idea of transferring knowledge between knowledge containers in a case‐based reasoning (CBR) system. A machine‐learning technique, fuzzy decision‐tree induction, is used to transform the case knowledge to adaptation knowledge. By learning the more sophisticated fuzzy adaptation knowledge, many of the redundant cases can be removed. This approach is particularly useful when the case base consists of a large number of redundant cases and the retrieval efficiency becomes a real concern of the user. The method of maintaining a case base from scratch, as proposed in this article, consists of four steps. First, an approach to learning feature weights automatically is used to evaluate the importance of different features in a given case base. Second, clustering of cases is carried out to identify different concepts in the case base using the acquired feature‐weights knowledge. Third, adaptation rules are mined for each concept using fuzzy decision trees. Fourth, a selection strategy based on the concepts of case coverage and reachability is used to select representative cases. In order to demonstrate the effectiveness of this approach as well as to examine the relationship between compactness and performance of a CBR system, experimental testing is carried out using the Traveling and the Rice Taste data sets. The results show that the testing case bases can be reduced by 36 and 39 percent, respectively, if we complement the remaining cases by the adaptation rules discovered using our approach. The overall accuracies of the two smaller case bases are 94 and 90 percent, respectively, of the originals.
Simon C. K. Shiu, Daniel S. Yeung, Cai Hung Sun, Xizhao Wang
Comput. Intell.2
2001 A new approach to fuzzy rule generation: fuzzy extension matrix
Xizhao Wang, X. F. Xu, W. D. Ling, Daniel S. Yeung
Fuzzy Sets Syst.5
2001 A General Updating Rule for Discrete Hopfield-Type Neural Network with Time-Delay and the Corresponding Search Algorithm
abstract
In this paper, the Hopfield neural network with delay (HNND) is studied from the standpoint of regarding it as an optimizing computational model. Two general updating rules for networks with delay (GURD) are given based on Hopfield-type neural networks with delay for optimization problems and characterized by dynamic thresholds. It is proved that in any sequence of updating rule modes, the GURD monotonously converges to a stable state of the network. The diagonal elements of the connection matrix are shown to have an important influence on the convergence process, and they represent the relationship of the local maximum value of the energy function to the stable states of the networks. All the ordinary discrete Hopfield neural network (DHNN) algorithms are instances of the GURD. It can be shown that the convergence conditions of the GURD may be relaxed in the context of applications, for instance, the condition of nonnegative diagonal elements of the connection matrix can be removed from the original convergence theorem. A new updating rule mode and restrictive conditions can guarantee the network to achieve a local maximum of the energy function with a step-by-step algorithm. The convergence rate improves evidently when compared with other methods. For a delay item considered as a noise disturbance item, the step-by-step algorithm demonstrates its efficiency and a high convergence rate. Experimental results support our proposed algorithm.
Daniel S. Yeung, Shenshan Qiu, Eric C. C. Tsang, Xizhao Wang
Int. J. Comput. Intell. Appl.1
2001 Sensitivity analysis of multilayer perceptron to input and weight perturbations
abstract
An important issue in the design and implementation of a neural network is the sensitivity of its output to input and weight perturbations. In this paper, we discuss the sensitivity of the most popular and general feedforward neural networks--multilayer perceptron (MLP). The sensitivity is defined as the mathematical expectation of the output errors of the MLP due to input and weight perturbations with respect to all input and weight values in a given continuous interval. The sensitivity for a single neuron is discussed first and an analytical expression that is a function of the absolute values of input and weight perturbations is approximately derived. Then an algorithm is given to compute the sensitivity for the entire MLP. As intuitively expected, the sensitivity increases with input and weight perturbations, but the increase has an upper bound that is determined by the structural configuration of the MLP, namely the number of neurons per layer and the number of layers. There exists an optimal value for the number of neurons in a layer, which yields the highest sensitivity value. The effect caused by the number of layers is quite unexpected. The sensitivity of a neural network may decrease at first and then almost keeps constant while the number increases.
Xiaoqin Zeng, Daniel S. Yeung
IEEE Trans. Neural Networks2
2001 A comparative study on heuristic algorithms for generating fuzzy decision trees
abstract
Fuzzy decision tree induction is an important way of learning from examples with fuzzy representation. Since the construction of optimal fuzzy decision tree is NP-hard, the research on heuristic algorithms is necessary. In this paper, three heuristic algorithms for generating fuzzy decision trees are analyzed and compared. One of them is proposed by the authors. The comparisons are twofold. One is the analytic comparison based on expanded attribute selection and reasoning mechanism; the other is the experimental comparison based on the size of generated trees and learning accuracy. The purpose of this study is to explore comparative strengths and weaknesses of the three heuristics and to show some useful guidelines on how to choose an appropriate heuristic for a particular problem.
Xizhao Wang, Daniel S. Yeung, Eric C. C. Tsang
IEEE Trans. Syst. Man Cybern. Part B2
2000 Detection of Language (Model) Errors
abstract
The bigram language models are popular, in much language processing applications, in both Indo-European and Asian languages. However, when the language model for Chinese is applied in a novel domain, the accuracy is reduced significantly, from 96% to 78% in our evaluation. We apply pattern recognition techniques (i.e. Bayesian, decision tree and neural network classifiers) to discover language model errors. We have examined 2 general types of features: model-based and language-specific features. In our evaluation, Bayesian classifiers produce the best recall performance of 80% but the precision is low (60%). Neural network produced good recall (75%) and precision (80%) but both Bayesian and Neural network have low skip ratio (65%). The decision tree classifier produced the best precision (81%) and skip ratio (76%) but its recall is the lowest (73%).
Kei Yuen Hung, Robert Wing Pong Luk, Daniel S. Yeung, Korris Fu-Lai Chung, Wenhao Shu
EMNLP3
2000 Stability of discrete Hopfield neural networks with time-delay
abstract
It is well-known that discrete Hopfield neural networks (DHNNs) without delay converge to a stable state. Due to this property, DHNNs without delay have wide potential applications to many fields, such as associative memory devices and combinatorial optimization. A DHNN with delay, which can deal with temporal information, is a generalization of a DHNN without delay. This paper investigates the convergence theorems of DHNNs with delay, based on new updating modes. A new bivariate energy function is constructed which represents the relationships between application problems and DHNNs with delay. It is proved that DHNNs with delay converge to a stable state. These results extend the existing results corresponding to DHNNs without delay. We also relate the maximum of this energy function to a stable state of DHNNs with delay. Furthermore, we describe algorithms for DHNNs with delay in detail.
Shenshan Qiu, Eric C. C. Tsang, Daniel S. Yeung
SMC3
2000 Mining fuzzy association rules with weighted items
abstract
In most models of mining fuzzy association rules, the items are considered to have equal importance. Due to diverse human interest and preference for items, such models do not work well in many situations. To improve such models, we propose a method to mine fuzzy association rules with weighted items. One of the major problems in data mining research is the development of good measures of interest of discovered rules. The weighted support and weighted confidence for fuzzy association rules are defined. Kohonen self-organized mapping is used to fuzzify the numerical attributes into linguistic terms. A new fuzzy association rule mining algorithm, which generalizes the popular Apriori Gen large itemset based algorithm, is developed. The advantages of the new algorithm are shown by testing it on a census database with 5000 transaction records.
Yue Joyce Shu, Eric C. C. Tsang, Daniel S. Yeung, Daming Shi 0001
SMC3
2000 Fuzzy weighted classification rules induction from data
abstract
One popular approach for automatic generation of fuzzy classification rules is decision tree induction, but almost all of the existing decision tree induction methods have not considered the importance of each proposition in the antecedent (i.e. the weight) contributing to the consequent. Unfortunately, this weight plays an important role in many real world problems. We present an effective approach for learning fuzzy weighted classification rules from data. The weights for each rule antecedent propositions will be assigned based on a relative weight matrix. Some experiments are conducted and the results show that this approach usually can obtain a compact set of fuzzy rules and considerable classification accuracy, especially, the learning accuracy can be improved by incorporating the weight.
Eric C. C. Tsang, Hongbing Li, Daniel S. Yeung, John W. T. Lee
SMC3
2000 Refinement of fuzzy production rules by neuro-fuzzy networks
abstract
The knowledge acquisition bottleneck is well-known in the development of fuzzy knowledge based systems (i.e. FKBSs), and knowledge maintenance and refinement are important issues. The paper improves fuzzy production rule (FPR) representation power by exploiting prior knowledge and develops refinement tools which assist in debugging a FKBS's knowledge, thus easing the knowledge acquisition and maintenance bottlenecks. We focus on knowledge refinement where the FKBS's knowledge is debugged or updated in reaction to evidence that the FKBS is faulty or out-of-date. Some of the applied methods are presented. To select a feasible fuzzy rule set for classification, the most difficult task is finding a set of rules pertaining to the specific classification by choosing adaptive knowledge representation parameters such as local and global weights in fuzzy rules. We map the weighted fuzzy rules to a new neural network (five-layer-based knowledge neural network) so the knowledge representation parameters can be refined and fuzzy rule representation power can be improved. The dynamic assigning neuron method, gradient-descent method with penalizing functions and evolving strategy are considered. We show that this refinement method can maintain the accuracy and improve the comprehensibility and representation power of FPRs. Experiments on a special domain indicate that the refinement method and evolving strategy are able to significantly increase an FPR's representation power when compared with standard fuzzy knowledge-based networks.
Eric C. C. Tsang, Shenshan Qiu, Daniel S. Yeung
SMC3
2000 Using fuzzy integral to modeling case based reasoning with feature interaction
abstract
The guiding principle of case-based reasoning (CBR) is the CBR-hypothesis which assumes that "similar problems have similar solutions". This principle requires a model to compute the problem-similarity in terms of individual features. One frequently used model is to consider the weighted average of feature-similarities as an overall similarity measure. Due to some inherent interaction among diverse features, the weighted average model does not work well in many real-world problems. This paper proposes using a non-linear integral tool to address such a problem. Five fuzzy integrals with respect to a fuzzy measure or a nonadditive set function are discussed in this paper. The interaction among the features is considered to be reflected in the non-additive set function, and the overall similarity is computed by using the integral model instead of using the weighted average model. Because the weighted average can be regarded as a special case of nonlinear integral, this paper to some extent generalizes the application scope of traditional CBR techniques based on similarity.
Xizhao Wang, Daniel S. Yeung
SMC2
2000 Using a neuro-fuzzy technique to improve the clustering based on similarity
abstract
Although there have been many approaches to fuzzy clustering, the clustering based on a similarity matrix is still a popular technique which performs by means of transforming the similarity matrix into its transitive closure. The clustering performance depends strongly on the similarity matrix in which elements are determined according to a distance metric in many situations. For a given case library in which diverse similarity measures can be defined, different similarity matrixes result in different clustering results. This paper introduces the concept of feature weight and then incorporates this concept into the process of computing similarity between two cases, such that the similarity matrix relies on these feature weights. The purpose of this paper is to improve the clustering performance by adjusting these weights in terms of a neural-fuzzy technique. To learn the feature weights, a neural network is designed for minimizing an objective function. For achieving a local minimum of the objective function, the gradient-descent technique is used to train this network. Several indexes for measuring the quality of a clustering result are defined in this paper to compare the performance.
Daniel S. Yeung, Xizhao Wang
SMC1
2000 Sensitivity analysis of multilayer perceptron to input perturbation
abstract
An important issue in the design and implementation of neural networks is the sensitivity of neural network output to parameter perturbations. Past research in this area has focused on network sensitivity analysis after training. Very few research projects have considered sensitivity analysis as a design issue prior to network implementation. The authors discuss the sensitivity of the most popular and general feedforward networks (multilayer perceptron (MLP)) to its input perturbation. The sensitivity is defined as the mathematical expectation of output errors of the MLP arising from input error with respect to all input and weight values in a given continuous interval. The sensitivity for a single neuron is discussed first, and an analytical expression that is a function of the input error is approximately derived. Then an algorithm is given to compute the sensitivity for an entire MLP network. The theoretical results of the derived formula were shown to agree with experimental results. By analyzing the derived analytical expression and implementing the given algorithm on a number of representative MLP networks, some significant observations on the behavior of sensitivity are discovered, which could be useful for network design consideration.
Xiaoqin Zeng, Daniel S. Yeung, Xuequan Sun
SMC2
2000 Improving learning accuracy of fuzzy decision trees by hybrid neural networks
abstract
Although the induction of fuzzy decision tree (FDT) has been a very popular learning methodology due to its advantage of comprehensibility, it is often criticized to result in poor learning accuracy. Thus, one fundamental problem is how to improve the learning accuracy while the comprehensibility is kept. This paper focuses on this problem and proposes using a hybrid neural network (HNN) to refine the FDT. This HNN, designed according to the generated FDT and trained by an algorithm derived in this paper, results in a FDT with parameters, called weighted FDT. The weighted FDT is equivalent to a set of fuzzy production rules with local weights (LW) and global weights (GW) introduced in our previous work (1998). Moreover, the weighted FDT, in which the reasoning mechanism incorporates the trained LW and GW, significantly improves the FDTs' learning accuracy while keeping the FDT comprehensibility. The improvements are verified on several selected databases. Furthermore, a brief comparison of our method with two benchmark learning algorithms, namely, fuzzy ID3 and traditional backpropagation, is made. The synergy between FDT induction and HNN training offers new insight into the construction of hybrid intelligent systems with higher learning accuracy.
Eric C. C. Tsang, Xizhao Wang, Daniel S. Yeung
IEEE Trans. Fuzzy Syst.3
1999 Neocognitron's Parameter Tuning by Genetic Algorithms
abstract
The further study on the sensitivity analysis of Neocognitron is discussed in this paper. Fukushima's Neocognitron is capable of recognizing distorted patterns as well as tolerating positional shift. Supervised learning of the Neocognitron is fulfilled by training patterns layer by layer. However, many parameters, such as selectivity and receptive fields are set manually. Furthermore, in Fukushima's original Neocognitron, all the training patterns are designed empirically. In this paper, we use Genetic Algorithms (GAs) to tune the parameters of Neocognitron and search its reasonable training pattern sets. Four contributions are claimed: first, by analyzing the learning mechanism of Fukushima's original Neocognitron, the correlations amongst the training patterns are claimed to affect the performance of Neocognitron, tuning the Neocognitron's number of planes is equivalent to searching reasonable training patterns for its supervised learning; second, a GA-based supervised learning of the Neocognitron is carried out in this way, searching the parameters and training patterns by GAs but specifying the connection weights by training the Neocognitron; third, other than traditional GAs which are unsuitable for the large searching space of training patterns set, the cooperative coevolution is incorporated to play this role; fourth, an effective fitness function is given out when applying the above methodology into numeral recognition. The evolutionary computation in our initial experiments is implemented based on the original training pattern set, e.g. the individuals of the population are generated from Fukushima's original training patterns during initialization of GAs. The results prove that our correlation analysis is reasonable, and show that the performance of a Neocognitron is sensitive to its training patterns, selectivity and receptive fields, especially, the performance is not monotonically increasing with respect to the number of training patterns, and this GA-based supervised learning is able to improve Neocognitron's performance.
Daming Shi 0001, Chunlei Dong, Daniel S. Yeung
Int. J. Neural Syst.3
1999 Sensitivity analysis of neocognitron
abstract
Fukushima's (1988; 1989; 1992; 1993) neocognitron model is well-known for its performance in visual pattern recognition. Through a training process, the visual pattern information is stored in a form of numerical weights in memory. When the model is actually implemented in hardware, weight errors and input noises caused by hardware imprecision and imperfect input devices respectively cannot be avoided and consequently the recognition performance usually degrades substantially from the theoretical result. In this paper, the effects of weight imprecision and input noise to the recognition performance of neocognitron are studied through a sensitivity analysis of the model. The sensitivity of an S-cell to weight and input perturbations is first derived, as a function of the weight and input perturbation ratios. An algorithm is proposed to combine the sensitivities of the S-cells in different layers to form the overall sensitivity of the model. The established sensitivity measure is then demonstrated to be a useful tool for hardware design. In addition, it has been found that the decision (recognition) error of neocognitron increases with the weight perturbation, the input perturbation, and the number of weights per neuron. This is similar to the result obtained for the Madaline model. Another important result is that decision error increases with the threshold/selectivity parameter. This supports the functional description of the threshold reported by Fukushina.
A. Y. Cheng, Daniel S. Yeung
IEEE Trans. Syst. Man Cybern. Part C2
1998 An extension of a fuzzy substroke extractor
abstract
In employing a structural approach to recognize handwritten Chinese characters (HCC), substroke is usually chosen as a feature set. Various types of ambiguities may arise during the substroke extraction. A fuzzy substroke extractor was previously (1990) proposed by the authors to handle a number of ambiguities caused by substroke touchings. The extractor then made use of the scoring information on the extracted items to produce a set of "consistent" outputs, from which no edge-sharing is found among the extracted items. The major advantage gained from this approach is that it is possible to reconstruct the skeleton in a "just-fit" mode. Thus, for the candidate set with most of the desirable extracted substrokes, the noise rate is much lower than those reported by using a conventional "explore-all-possibility" approach. To achieve this goal, the segmentation of character skeleton is transformed into a fuzzy set partitioning task. This paper extends our current technique to handle another type of ambiguity problem in substroke extraction, i.e. broken substrokes. Two cases of broken substrokes are addressed. Under certain conditions, virtual edges are introduced to "complete" the "supposed" broken skeleton graph. By considering this new skeleton as the actual input, suspected broken substrokes are detectable as well. With this proposed extension, most ambiguities encountered during substroke extraction can now be successfully treated in a unified framework.
Hank-Shun Fong, Daniel S. Yeung
SMC2
1998 Evaluation of printed circuit board assembly manufacturing systems using fuzzy colored Petri nets
abstract
In this paper, implementation and evaluation of a printed circuit board assembly (PCBA) manufacturing system model based on fuzzy colored Petri nets (FCPN) modeling was performed. From the Petri nets simulation results, two contributions were claimed: (1) The successful application of our approach to evaluate systems that consist of complicated concurrent processes with embedded data structures and uncertainty reasoning. (2) An approach in automatic determination of the threshold values in the fuzzy production rules using FCPN. The details of the assembling processes and various manufacturing data were gathered from computer manufacturers located in China. Different resource allocation strategies were introduced in the experiments, and simulation results were obtained. The resources included screen printers, different pick and place machines, infrared soldering machines and component insertion robots. All the experiments were built by using the Petri net tool DESIGN/CPN. A number of hierarchy pages, color sets, places, arcs and transitions representing the PCBA and fuzzy production rules were constructed in the experiments. Simulation runs of the Petri nets model were carried out. The result shows that our approach could be used to guide decision-makers in the design and selection of a suitable PCBA manufacturing strategy. Future extension of our work is described.
Simon C. K. Shiu, Eric C. C. Tsang, Daniel S. Yeung, Martin B. Lam
SMC3
1998 Refining local weights and certainty factors using a fuzzy neural network
abstract
In this paper a novel approach of tuning knowledge representation parameters (KRP) in a fuzzy production rule (FPR) using a fuzzy neural network (FNN) is proposed. Two KRP will be considered, i.e., the local weight of each proposition in a conjunctive and a disjunctive fuzzy production rules and the certainty factor of the whole rule. These parameters will be used in FPR whose fuzzy terms could not be represented with a distinct membership function or a discrete fuzzy set due to the fact that the bases of these fuzzy sets cannot be clearly or easily defined or identified. The significance of this research is that the refined parameters will result in a more accurate drawn conclusion of a multilevel reasoning system by adjusting the degree of truth of the drawn consequent with the local weight. Furthermore, the time required to consult with domain experts to tune or refine these parameters is greatly reduced as a FNN could help knowledge engineers solve this refinement problem.
Eric C. C. Tsang, Daniel S. Yeung
SMC2
1998 Refinement of knowledge representation parameters in fuzzy production rules by genetic algorithms
abstract
Focuses on using a genetic algorithm (GA) optimization technique to help knowledge engineers refine knowledge representation parameters (KRPs) which have been identified to be important and necessary to enhance the representation power of fuzzy production rules in the applications of fuzzy expert systems. These parameters include certainty factor, threshold value, fuzzy membership value, local and global weights. The gradient descent method in a multilayer perceptron neural network (NN) is replaced by a GA. The significance of such a refinement is that the refined or tuned parameters will enable the fuzzy expert system (FES) to draw more accurate and reasonable conclusions and reduce the chance of leading to a wrong goal. Furthermore, the time required to repeatedly consult with the domain experts will be reduced. The main concerns using a GA to solve this kind of problem are how to represent the problem by a GA and how to measure the fitness of each solution being considered. In the paper the representation method and fitness function are provided together with an experiment to demonstrate the GA approach to solve the refinement problem.
Eric C. C. Tsang, Daniel S. Yeung, John W. T. Lee
SMC2
1998 A proposed model for developing distributed programs by prototyping
abstract
The model for prototyping has three basic elements that can function as independent tools. These three elements are: programming environment, system, and program visualization. The programming environment helps the programmer write high level distributed programs with no more effort than sequential programming. The system element executes distributed programs and records their behaviour in dedicated descriptors that act as the model's unifying paradigm. Programming visualization lets users monitor program behaviour and rectify problems interactively by program reversion and partial program replacement. Test results show that the model can shorten software development time and enhance software quality by uncovering and rectifying errors quickly.
Allan K. Y. Wong, Daniel S. Yeung
SMC2
1998 Experiments on the use of corpus-based word BI-gram in Chinese word segmentation
abstract
The first step of Chinese language processing is to segment a Chinese sentence into a sequence of words due to the fact that there is no original separation between adjacent words. An efficient corpus-based statistical method is adopted here to address such a problem. In this paper, some word BI-gram statistical measures derived from corpus are employed to remove the segmentation ambiguities. To segment a Chinese sentence, a bidirectional maximum matching method is firstly used to do pre-matching in order to get segmentation candidates and locate possible ambiguities. The statistical measures based on word BI-gram information and word frequency will be used to construct a discriminate function, which is applied to ambiguity strings in order to get an utmost correct segmentation. Experimental results are analyzed to describe the features and limitations of this approach, and preliminary results indicate that our approach is compared favorably to other existing techniques.
Daniel S. Yeung
SMC2
1998 Neocognitron based handwriting recognition system performance tuning using genetic algorithm
abstract
Neural networks have been used to recognize handwritten characters such as Chinese, English or numerals. But their performance, i.e., the recognition rate, depends on a number of factors which may include the network architecture, feature selection, network parameter setting, learning strategy, learning sample selection, test pattern preprocessing, etc. These factors are important to network engineer in designing a network for a particular application problem, but unfortunately there is a lack of systematic way to guide their decision-making regarding the selection of these parameters. This paper presents a parameter tuning (namely the selectivity parameter) methodology based on a sensitivity analysis of the neocognitron model, and the off-line handwritten numeral recognition with supervised learning is chosen to be the demonstrated application problem. Genetic algorithm (GA) is used to select parameters leading to improved recognition results. We used a set of training pattern provided by Fukushima (1988) as our training patterns which involved no preprocessing, and our experimental results show a significant improvement in performance. A brief discussion on alternate hybrid architecture involving neural network and genetic algorithm, and different fitting functions for the GA will be presented.
Daniel S. Yeung, Y. T. Cheng, Hank-Shun Fong, Korris Fu-Lai Chung
SMC1
1998 Ambiguity handling of similar categories in handwritten Chinese character recognition
abstract
Chinese characters consist of thousands of categories, some of which have very similar structural characteristics. Ambiguity in recognition may thus arise. The authors previously (1990) proposed a technique for recognizing characters off-line based on their structural characteristics. Each input character is subject to various alternatives in stroke segmentation, and for each resulting stroke set, the strokes are matched against the category templates maintained in a knowledge base. This knowledge base is so devised to offer certain degree of tolerance to handwriting ambiguities, with respect to individual character categories. This paper aims at extending our current technique to distinguish characters within a similar category as well. Basically, Chinese characters composed of the same stroke set, but with different geometric attributes, are grouped into a similar category. A secondary knowledge base is created to store these similar groups. Our previous methodology will be modified so that once a candidate output is identified (together with its computed score of matching result), all characters belonging to the same similar category will have their matching scores recalculated and possibly new ranking information may ultimately lead to new output candidates. This may eventually improve the system's recognition performance. The complete construction of the secondary knowledge will not be possible since it is application domain specific. But a number of similar categories will be presented to demonstrate how the method works.
Daniel S. Yeung, Hank-Shun Fong
SMC1
1998 A multilevel weighted fuzzy reasoning algorithm for expert systems
abstract
The applications of fuzzy production rules (FPR) are rather limited if the relative degree of importance of each proposition in the antecedent contributing to the consequent (i.e., the weight) is ignored or assumed to be equal. Unfortunately, this is the case for many existing FPR and most existing fuzzy expert system development shells or environments offer no such functionality for users to incorporate different weights in the antecedent of FPR. This paper proposes to assign a weight parameter to each proposition in the antecedent of a FPR and a new fuzzy production rule evaluation method (FPREM) which generalizes the traditional method by taking the weight factors into consideration is devised. Furthermore, a multilevel weighted fuzzy reasoning algorithm (MLWFRA) incorporating this new FPREM, which is based on the reachability and adjacent place characteristics of a fuzzy Petri net, is developed. The MLWFRA has the advantages that i) it offers multilevel reasoning capability; ii) it allows multiple conclusions to be drawn if they exist; iii) it offers a new fuzzy production rule evaluation method; and iv) it is capable of detecting cycle rules.
Daniel S. Yeung, Eric C. C. Tsang
IEEE Trans. Syst. Man Cybern. Part A1
1997 Offline Handwritten Chinese Character Recognition viaRadical Extraction and Recognition
abstract
Despite the fact that Chinese characters are composed of radicals and that Chinese people usually formulate their knowledge of Chinese characters as a combination of radicals, very few studies have focused on a character decomposition approach to recognition, i.e., recognizing a character by first extracting and recognizing its radicals. Such an approach is adopted and the problem of how to extract radical sub-images from character images is particularly addressed. A radical extraction algorithm based on deformable templates (DTs) has been developed. The advantage of the character decomposition approach is demonstrated by feeding the extracted radical images to an adopted structural based Chinese character recognizer whose outputs are then combined to produce the class label of the input character. Simulation results show that the performance of the adopted Chinese character recognition system can be improved significantly when the character decomposition approach is used.
Wilson W. S. Ip, Korris Fu-Lai Chung, Daniel S. Yeung
ICDAR3
1997 Formal verification of the correctness in hybrid expert systems
abstract
It has been increasingly recognized over recent years that expert systems which combine one or more techniques greatly increase the problem solving capability and help overcome some of the shortcomings associated with any single technique. The verification of these expert systems requires methods which could tackle the multiple knowledge representation paradigms and integrated inference mechanisms used. The paper provides a formal description technique for verifying the correctness of hybrid expert systems (HES) that emphasizes an integration of object hierarchy, property inheritance and production rules. The main idea is to convert the HES into a state controlled coloured Petri net (SCCPN) where the object hierarchy, property inheritance and production rules are modelled as separated components in the same SCCPN. The detection and analysis of the anomalies in the system are done by constructing and examining the reachability tree spanned by the knowledge inference. This provides a formal basis for automating the deduction process and a means of verifying HES. A set of propositions is formulated to verify errors and anomalies in HES. Lastly, future extension of the approach is discussed.
Simon C. K. Shiu, James Nga-Kwok Liu, Daniel S. Yeung
KES (2)3
1997 Acquiring and tuning knowledge representation parameters of fuzzy production rules using fuzzy expert networks
abstract
Fuzzy production rules (FPRs) have been used and proved to be a very useful knowledge representation method to capture and represent fuzzy, uncertain, incomplete and vague domain expert knowledge. The knowledge representation capability of these FPRs could be enhanced if parameters like local weights, certainty factors or threshold values are incorporated. These parameters, together with the membership values of fuzzy sets, are, however, difficult to acquire or extract from domain experts during the knowledge acquisition phases and to fine-tune during the system upgrade and maintenance phase. In this paper, the fuzzy expert networks (FENs) proposed by the authors in the World Congress on Neural Networks, pp. 500-3 (1996) are extended so that they can acquire and fine-tune more knowledge representation parameters (KRPs). Local weight is added to the KRPs and incorporated into the antecedent part of a conjunctive FPR. The knowledge acquisition and refinement problems of these parameters and the membership values of fuzzy sets can be solved by using FENs which not only have the reasoning mechanism of a fuzzy expert system (FES) but also the learning capability of a neural network. An experiment is presented to illustrate the workability of our proposed method.
Eric C. C. Tsang, Daniel S. Yeung
KES (2)2
1997 Weighted fuzzy production rules
Daniel S. Yeung, Eric C. C. Tsang
Fuzzy Sets Syst.1
1997 A comparative study on similarity-based fuzzy reasoning methods
abstract
If the given fact for an antecedent in a fuzzy production rule (FPR) does not match exactly with the antecedent of the rule, the consequent can still be drawn by technique such as fuzzy reasoning. Many existing fuzzy reasoning methods are based on Zadeh's Compositional Rule of Inference (CRI) which requires setting up a fuzzy relation between the antecedent and the consequent part. There are some other fuzzy reasoning methods which do not use Zadeh's CRI. Among them, the similarity-based fuzzy reasoning methods, which make use of the degree of similarity between a given fact and the antecedent of the rule to draw the conclusion, are well known. In this paper, six similarity-based fuzzy reasoning methods are compared and analyzed. Two of them are newly proposed by the authors. The comparisons are two-fold. One is to compare the six reasoning methods in drawing appropriate conclusions for a given set of FPRs. The other is to compare them based on five issues: 1) types of FPR handled by these methods; 2) the complexity of the methods; 3) the accuracy of the conclusion drawn; 4) the accuracy of the similarity measure; and 5) the multi-level reasoning capability. The results have shed some lights on how to select an appropriate fuzzy reasoning method under different environments.
Daniel S. Yeung, Eric C. C. Tsang
IEEE Trans. Syst. Man Cybern. Part B1
1996 Behavioural modelling in object-oriented methodology
Kai-On Chow, Daniel S. Yeung
Inf. Softw. Technol.2
1996 A fuzzy substroke extractor for handwritten Chinese characters
Daniel S. Yeung, Hank-Shun Fong
Pattern Recognit.1
1995 Modelling Hybrid Rule/Frame-Based Expert Systems Using Coloured Petri Nets
Simon C. K. Shiu, James Nga-Kwok Liu, Daniel S. Yeung
IEA/AIE3
1995 A knowledge matrix representation for a rule-mapped neural network
Daniel S. Yeung, Hank-Shun Fong
Neurocomputing1
1994 Knowledge Matrix - An Explanation & Knowledge Refinement Facility for a Rule Induced Neural Network
Daniel S. Yeung, Hank-Shun Fong
AAAI1
1994 Improved fuzzy knowledge representation and rule evaluation using fuzzy petri nets and degree of subsethood
abstract
In this article a variation of fuzzy Petri net (FPN) model is proposed to accommodate for the possibility of mapping fuzzy production rule (FPR) having different threshold values in their propositions into FPN. the purpose of assigning a different threshold value for each proposition in the FPR and of using the rule checking and evaluation method proposed here is to prevent misfiring of the rule, which can result with other methods; the purpose of having variation of FPN model is to capture and represent more information of FPR in the FPN model. the rule checking and evaluation method is an enhancement of the approach proposed by Yeung (D. S. Yeung et al., Proceedings of the 6th International Conference on System Research Informatics and Cybernetics, Germany, 1992). As mentioned by the authors, the degree of subsethood between two vectors is the basis of the method. the subsethood method will first be used to make certain that each input value for the proposition in the antecedent is greater than or equal to its corresponding threshold value. When such condition holds, the subsethood method is used to infer the degree of truth of the consequent of the rule. an enhanced fuzzy reasoning algorithm is included. Comparison of this method with other methods is presented. Future research work in determining acceptable threshold values and certainty factors is addressed. © 1994 John Wiley & Sons, Inc.
Daniel S. Yeung, Eric C. C. Tsang
Int. J. Intell. Syst.1
1994 A Hybrid Cognitive System Using Production Rules to Synthesize Neocognitrons
abstract
A hybrid cognitive system is proposed where a working neocognitron is synthesized with a set of production rules. The knowledge base of a neocognitron is constructed through incorporating production rules into its interlayer connections. Training for prototype patterns is not required. The semantic of interlayer connections is established. The resulting network can now be analyzed according to the rule structure and problematic portions can be corrected. Neocognitrons constructed using this hybrid approach have been tested on the same set of handwritten numerals initiated by Fukushima with scaling and skewing distortions, and with noise contamination. It is found that the performance is comparable to that of Fukushima's network obtained by supervised training.
Daniel S. Yeung, Hing-Yip Chan
Int. J. Neural Syst.1
1994 Incorporating Production Rules with Spatial Information Onto a Neocognitron Neural Network
abstract
Rule-embedded neocognitron (REN) is proposed where the knowledge base of a neocognitron is constructed through incorporating production rules into its interlayer connections. Prototype patterns training is not required. The semantic of interlayer connections is established. The resulting network can now be analyzed according to the rule structure and problematic portions can be corrected. We demonstrate the ease with which performance can be improved by applying REN on handwritten numeral recognition. The same set of handwritten numerals initiated by Fukushima is used to test this methodology. It is found that the performance is comparable with that of Fukushima's neocognitron with supervised training.
Daniel S. Yeung, Hing-Yip Chan, Kwan-Fai Cheung
Int. J. Neural Syst.1
1994 Handwritten Chinese Character Recognition by Rule-Embedded Neocognitron
Daniel S. Yeung, Hank-Shun Fong
Neural Comput. Appl.1