VLDB 2026 Research / reviewers in the wild / expert
Wing W. Y. Ng
dblp:67/6060
· DBLP profile ↗
123ranked-venue papers
31as first author
74since 2021 · last 2026
0000-0003-0783-3585ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 16 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 8 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 25 · 7 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 16 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 5 since 2021Computer networks · 4 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VERIMLA: A Unified Framework for Verilog Generation via Multi-Level Alignment and Formal Equivalence Verification
Wing W. Y. Ng, Wen Li 0015, Meihua Liu |
ISCAS | 1 |
| 2026 | Dual self-distillation: Employing progressive and masked reconstruction to enhance adversarial defense
Lei Zhao 0029, Wing W. Y. Ng, Jianjun Zhang 0004, Han Zhou 0014 |
Expert Syst. Appl. | 2 |
| 2026 | Incremental sampling hashing for image retrieval with concept drift
Wing W. Y. Ng, Linfei Wang, Qihua Li, Xing Tian |
Pattern Recognit. | 1 |
| 2026 | Localized Intra- and Inter-Tumoral Heterogeneity for Predicting Treatment Response to Neoadjuvant Chemotherapy in Breast CancerabstractThis study proposes a novel method for extracting breast cancer tumor heterogeneity descriptors to non-invasively predict whether pathological complete response (pCR) can be achieved after neoadjuvant chemotherapy (NAC). These localized descriptors extract corresponding heterogeneity features for different radiomic features and are able to capture tumor characteristics at various localization levels. These descriptors also capture tumor heterogeneity both at the individual tumor level and across the whole dataset, providing decision-making models with features that are both more effective and interpretable. We validated the effectiveness of the proposed features with the Kolmogorov-Arnold network (KAN) across multiple centers, yielding an AUC of 0.92 when combined with pathological features and demonstrating good performance in external datasets (AUCs of 0.84 and 0.81). Additionally, we transform the best model into a symbolic formula to intuitively explain the machine learning model's prediction process, showing how factors such as age, HER2, Ki-67 and heterogeneity influence the prediction. The symbolized model is consistent with the experience of clinical experts, which enhances users' confidence in deep models. The experimental results show that our proposed features and method outperform classical heterogeneity features and end-to-end neural networks with a small additional computational cost. Yinhao Liang, Qingcong Kong, Ting Wang 0015, Jianjun Zhang 0004, Wing W. Y. Ng |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Efficient Robustness for Small Models via Dual Adversarial Distillation With Hybrid SupervisionabstractSmall models, despite their computational efficiency for real-time and edge applications, remain vulnerable to adversarial attacks. Adversarial Distillation (AD) has proven effective in enhancing model robustness, yet current approaches predominantly rely on single-strategy adversarial samples with either fixed or adaptive teacher supervision. Fixed supervision often leads to over-smoothing due to static guidance, while adaptive supervision incurs higher computational costs and convergence instability. To address these limitations, we propose Dual Adversarial Distillation (DAD), a novel framework that synergistically combines fixed and adaptive supervision through multi-intensity adversarial samples, with customized distillation strategies designed to enhance knowledge diversity and feature transfer. For fixed supervision, a Mixup-based strategy is employed to diversify teacher knowledge, ensuring robust feature representations by aligning the student model with the teacher's rich representations through cross-strength feature consistency. For adaptive supervision, learning diversity is augmented by integrating transformed features from multiple projection modules, effectively minimizing the feature distribution gap between teacher and student models. Computational efficiency is optimized through variable iteration strategies. Extensive experiments demonstrate the effectiveness of our method in improving the robustness of small models, achieving state-of-the-art baselines. Lei Zhao 0029, Wing W. Y. Ng, Jianjun Zhang 0004, Han Zhou 0014, Sam Kwong |
IEEE Trans. Multim. | 2 |
| 2025 | Updatable Verifiable Credential Sharing with Selective Disclosure
Jui-Yung Lin, Xingfu Yan, Wing W. Y. Ng, Gong Zheng, Ying Gao 0004 |
ICA3PP (1) | 3 |
| 2025 | EPOMTA: Efficient and Privacy-Preserving Online Multi-task Allocation in Mobile Crowdsourcing
Jiashuang Xu, Xingfu Yan, Wing W. Y. Ng |
ICA3PP (4) | 3 |
| 2025 | Incremental Hashing with Asymmetric Distance for Image Retrieval in Non-stationary Environments
Xing Tian, Zihao Zhan, Dezhong Zhu, Wing W. Y. Ng, Chunlin Xu |
PRICAI (5) | 4 |
| 2025 | ACIL-SED: An Acoustic Clustering and Imbalance Learning Method for Sound Event DetectionabstractSound Event Detection (SED) aims to identify and locate specific sound events within audio streams. Despite recent advancements, current methodologies exhibit two fundamental limitations: (1) Insufficient modeling of discriminative temporal-spectral characteristics induces compromised differentiation for acoustically similar events. (2) Failing to address both inter-class and intra-class imbalance problems induces biased classifications toward dominant classes and inactive frames. To tackle these challenges, we propose Acoustic Clustering and Imbalance Learning-based Sound Event Detection (ACIL-SED), which consists of Cluster-Specialized Convolutional Recurrent Neural Network (CS-CRNN) and Dual-Objective Adaptive Balance Loss (DOABLoss). The CS-CRNN employs acoustic clustering-guided specialized sub-models, where each cluster’s sub-model focuses on discriminative feature learning in specific temporal-spectral characteristics, thereby resolving the compromised feature representation inherent in a shared single-model architecture. The DOABLoss adaptively combines mean squared error (MSE) and weighted binary cross-entropy (WBCE) to simultaneously address the issues of inter-class duration imbalance and intra-class active-inactive frame imbalance. Experimental results show that ACIL-SED outperforms state-of-the-art methods in complex acoustic environments. Cankun Zhong, Yun Liang 0003, Tang Luo, Wing W. Y. Ng |
SMC | 6 |
| 2025 | A Multimodal Gait-Based Depression Recognition Model with Consistency TrainingabstractDepression is one of the most common psychological disorders. In recent years, gait data-based depression recognition methods have drawn a lot of attention. However, in existing gait-based depression recognition models, the extensive use of dropout leads to non-negligible inconsistencies between the training and inference stages, which may limit further improvements of model performance. In this paper, we propose a multimodal gait-based depression recognition model, which integrates skeleton and silhouette modalities of gait and incorporates a consistency training method to address the inconsistencies caused by dropout. Experiments conducted on the D-Gait depression dataset demonstrate that the proposed model with consistency training yields better and more stable performance. Additionally, through extended experiments, we provide substantial insights into the trade-offs between advantages and disadvantages of the consistency training method. Zhuoyong Huang, Wing W. Y. Ng, Zimin Guo, Huakang Li |
SMC | 2 |
| 2025 | ConMH-Based Multi-Modal Video Retrieval with Contrastive Hashing and FusionabstractWith the rapid progress of urbanization, city governance faces growing challenges such as traffic violations and environmental pollution. Traditional manual monitoring methods are inefficient and costly. To enhance the efficiency of monitoring and managing uncivil behaviors in urban environments, we propose a self-supervised video hashing retrieval framework for uncivil behavior recognition. Leveraging deep learning techniques, our method generates compact binary hash codes for both video and text modalities via a contrastive masked autoencoder (ConMH), enabling efficient large-scale retrieval. We further improve ConMH by introducing cross-attention mechanisms in the text hashing branch to better handle context dependencies. To optimize retrieval results, we integrate five multimodal fusion and ranking strategies, including a novel Hybrid Distance-Rank Fusion method that balances similarity scores and rank information. Experiments conducted on MSRVTT and MSVD datasets demonstrate that our approach achieves superior performance in mAP@K and NDCG metrics. The framework significantly enhances cross-modal semantic coverage, ensures high retrieval precision, and maintains low computational and storage overhead through binary encoding. Rongye Ling, Jingrou Li, Wing W. Y. Ng, Qihua Li, Xing Tian, Xingfu Yan |
SMC | 3 |
| 2025 | DCED: Deformable Convolutional Encoder-Decoder Network for Inflamed Appendix Segmentation and Classification from CT ImagesabstractAcute appendicitis (AA) is one of the most prevalent surgical acute abdominal condition diseases. The recognition and segmentation of the inflamed appendix are important for AA diagnosis. However, it is a challenging task to find and segment the inflamed appendix from computed tomography (CT) images due to the varying sizes and shapes of different appendices and blurred borders with nearby tissues. To the best of our knowledge, the general expert segmentation model suffers due to the characterization of the inflamed appendix. Thus, we propose a deformable convolutional encoder-decoder network (DCED) for better recognition and segmentation of the inflamed appendix. The network consists of an encoder, a bottleneck, and a decoder. The encoder is composed of several convolutional neural network (CNN) layers to capture the local structural information. The bottleneck based on a vision transformer (ViT) focuses on the region of interest (ROI) using the global attention mechanism. The encoder and bottleneck modules effectively combine the local and global information of input data to locate the inflamed appendix. The decoder based on a deformable convolutional network (DCN) learns the varied boundary information, which helps to improve the accuracy of boundary segmentation. Extensive experimental results on a real-world AA dataset show that the proposed method yields the best average Dice similarity coefficient (DSC) of 71.29% and average Hausdorff Distance 95% (HD95) of 12.38 mm in comparison to state-of-the-art segmentation methods. Wing W. Y. Ng, Peixin Zheng, Yinhao Liang, Ting Wang 0015, Jianjun Zhang 0004, Xinhua Wei |
SMC | 1 |
| 2025 | GaitBranch: A multi-branch refinement model combined with frame-channel attention mechanism for gait recognition
Huakang Li, Yidan Qiu, Huimin Zhao 0001, Jin Zhan, Rongjun Chen 0001, Jinchang Ren, Ying Gao 0004, Wing W. Y. Ng |
Comput. Vis. Image Underst. | 8 |
| 2025 | Robust graph representation learning with asymmetric debiased contrasts
Wen Li 0015, Wing W. Y. Ng, Hengyou Wang, Jianjun Zhang 0004, Cankun Zhong |
Expert Syst. Appl. | 2 |
| 2025 | Broad hashing for image retrieval
Wing W. Y. Ng, Xuyu Liu, Xing Tian, Ting Wang 0015, Jianjun Zhang 0004, C. L. Philip Chen |
Neurocomputing | 1 |
| 2025 | ERMAV: Efficient and Robust Graph Contrastive Learning via Multiadversarial Views TrainingabstractGraph contrastive learning (GCL) is emerging as a pivotal technique in graph representation learning. However, recent research indicates that GCL is vulnerable to adversarial attacks, while existing robust GCL methods against adversarial attacks are inefficient and lack scalability due to the significant computational expenses of explicit adversarial attacks on the graph structure. To address the shortcomings of existing approaches, we propose an efficient and robust GCL via multiadversarial views training framework, called ERMAV. Specifically, the ERMAV generates two adversarial views by attacking both node attributes and latent representations on randomly sampled subgraphs. The method conducts explicit adversarial attacks on node attributes by attacking node attributes and implicit adversarial attacks on the graph structure by attacking latent representations, which avoids the costly computation of explicit graph structure attacks. Moreover, two efficient attack methods are developed to construct adversarial perturbations, which can dynamically generate different adversarial views to enhance sample diversity in the training phase. Furthermore, to validate the effectiveness and robustness of the proposed framework, extensive experiments of node classification on seven real-world datasets are conducted. Experimental results show that our ERMAV outperforms state-of-the-art GCL methods on the original graphs and is consistently more robust than existing robust GCL methods on a variety of attacked graphs. This demonstrates the strong robustness and great potential of our ERMAV in real-world applications. Wen Li 0015, Wing W. Y. Ng, Hengyou Wang, Jianjun Zhang 0004, Cankun Zhong, Liang Yang 0002 |
IEEE Trans. Cybern. | 2 |
| 2025 | Scope: On Detecting Constrained Backdoor Attacks in Federated LearningabstractFederated learning (FL) allows multiple clients to train an efficient deep-learning model collaboratively but is susceptible to backdoor attacks. Traditional detection-based defenses depend on specific metrics to distinguish client gradients. Defense-aware attackers exploit this by constraining attack gradients on these metrics to evade detection, leading to metric-constrained attacks. This paper concretely instantiates such threats and introduces cosine-constrained attacks, which successfully compromise advanced defenses based on cosine distance. To address the aforementioned challenge, we propose Scope, a novel defense that detects cosine-constrained attacks using cosine distance by exposing the constrained backdoor dimensions of attack gradients. Scope employs dimension-wise normalization and differential scaling to amplify the distinction between backdoor dimensions and benign or unused ones, countering sophisticated attackers’ attempts to obscure them. Moreover, we develop a novel clustering approach, namely Dominant Gradient Clustering (DGC), to isolate and eliminate backdoor gradients. Extensive experiments across various datasets, models, FL settings, and adversary scenarios demonstrate that Scope consistently outperforms existing defenses by a significant margin, especially against the cosine-constrained attack. Additionally, we present a Scope-tailored attack designed to evade Scope, but it remains ineffective even when maximizing stealthiness, further underscoring the robustness of Scope. We release our source code at:https://github.com/siquanhuang/Scope. Siquan Huang, Yijiang Li, Xingfu Yan, Ying Gao 0004, Chong Chen 0011, Leyu Shi, Wing W. Y. Ng |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2025 | Fairness-Aware Client Selection and Payment Determination for Differentially Private Federated LearningabstractFederated Learning (FL) mitigates data leakage by sharing only local machine learning models instead of raw data. However, it remains vulnerable to differential attacks. Differential Privacy (DP) addresses this concern by introducing noise to make it challenging for adversaries to reconstruct training samples. Nonetheless, clients often have varying attitudes toward data privacy, quantified by their privacy budgets. Low privacy budgets indicate the stringent privacy requirements of clients, requiring high compensations to incentivize their participation. Focusing solely on privacy budgets, however, can introduce selection bias, potentially compromising model generalization. Therefore, it is essential to emphasizes the fairness of client participation, ensuring that clients with lower privacy budgets also have opportunities to contribute to the training process. To tackle the above challenges, this paper formulates a novel DP-based incentive problem in FL, aiming to optimize the utilities of both the server and the clients. Specifically, we propose an auction mechanism that jointly selects participants based on their heterogeneous privacy budgets and determines appropriate payments. The proposed auction mechanism is proven to achieve several desirable properties, including computational efficiency, individual rationality, budget balance, truthfulness, and guaranteed optimization performance. Finally, simulation results validate the effectiveness of the proposed mechanism. Xiumin Wang 0005, Weiwei Lin 0001, Wing W. Y. Ng, Kai Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | XRadNet: A Radiomics-Guided Breast Cancer Molecular Subtype Prediction Network With a Radiomics ExplanationabstractIn this work, we propose a radiomics-guided neural network, XRadNet, for breast cancer molecular subtype prediction. XRadNet is a two-head neural network, with one for predicting molecular subtypes and the other for approximating radiomic features. In addition, a training scheme with radiomics guidance is proposed to improve performance. First, we conduct a series of experiments to test the radiomic feature learning capacity of different neural networks, which determines the backbone of XRadNet. Moreover, significant radiomic features are also determined according to radiomics and prior knowledge. XRadNet is subsequently pretrained in a self-supervised manner. The pretraining uses synthetic samples to train the backbone and radiomic feature regression head. This mitigates the impact of an insufficient number of samples. Finally, XRadNet is fine-tuned with a downstream real-world dataset by enabling all heads. Furthermore, a logistic regression is built with radiomic features and learned features, which provides a new way to interpreting the trained model with concepts familiar to radiologists. The experimental results show that XRadNet effectively predicts the four molecular subtypes of breast cancer. These results also demonstrate that the proposed training scheme yields better or competitive performance than those models pretrained on ImageNet or medical datasets. Yinhao Liang, Jianjun Zhang 0004, Ting Wang 0015, Wing W. Y. Ng, Kuiming Jiang, Xinhua Wei, Xinqing Jiang |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | An Innovative Multisource Teacher Collaborative Framework for Self-Knowledge DistillationabstractSelf-knowledge distillation, abbreviated as SKD, exhibits greater computational efficiency than traditional knowledge distillation (KD) because it learns from its own predictions rather than from a pretrained teacher. Existing SKD methods diversify knowledge through auxiliary branches, data augmentation, historical models, and label smoothing. However, previous methods primarily extract knowledge from a single-source teacher, overlooking the diversity and complementarity of various types of teacher knowledge in model learning, thereby limiting performance improvements. In response to this challenge, we propose a pioneering paradigm termed multisource teacher collaboration for self-knowledge distillation (MSTCS-KD), which integrates knowledge from diverse types of teachers to complementarily enhance the model's learning capability. We start by adding lightweight auxiliary branches with different structures in the shallow layers to build the student network, while also incorporating a teacher-guided attention mechanism to support adaptive learning. Then, we perform collaborative distillation by combining "heterogeneous knowledge" from the primary network's deepest layers with "homogeneous knowledge" from the student's outputs on augmented samples. This complementary distillation approach improves the model's ability to learn features, generalize, and enhance trainability. Extensive experiments demonstrate that our method outperforms other state-of-the-art SKD methods across various network architectures and datasets. Lei Zhao 0029, Wing W. Y. Ng, Jianjun Zhang 0004, Xiguang Wu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Fastmandarin: Efficient Local Modeling for Natural Mandarin Speech SynthesisabstractAttention-based speech synthesis methods often suffer from dispersed attention across the entire input sequence, resulting in poor local modeling and unnatural Mandarin synthesized speech. To address these issues, we present FastMandarin, a rapid and natural Mandarin speech synthesis framework that employs two explicit methods to enhance local modeling and improve pronunciation representation. Firstly, we tag Chinese characters to delineate phrase boundaries within a sentence, and these tags are integrated into the network’s hidden layer features at each time step, effectively bolstering local contributions in latent representations. Secondly, we introduce a multi-scale context feature extractor network that employs parallel convolution with various filters. Additionally, we optimize duration alignment and Mel-spectrogram reconstruction to enhance overall performance. Experimental results demonstrate that FastMandarin excels in local modeling, delivering robust Mandarin speech synthesis results. Chenglong Jiang, Ying Gao 0004, Linrong Pan, Wing W. Y. Ng |
ICASSP | 5 |
| 2024 | MVRMLM 2024: Multimodal Video Retrieval and Multimodal Language ModellingabstractAs the proliferation of video content continues, and many video archives lack suitable metadata, therefore, video retrieval, particularly through example-based search, has become increasingly crucial. Existing metadata often fails to meet the needs of specific types of searches, especially when videos contain elements from different modalities, such as visual and audio. Consequently, developing video retrieval methods that can handle multi-modal content is essential. In designing our novel video retrieval framework named Multi-modal Video Search by Examples (MVSE)1, we focused on accuracy (precision and recall), efficiency (retrieval time in seconds), interactivity, and extensibility, with key components including advanced data processing and a user-friendly interface aimed at enhancing search effectiveness and user experience. With the advent of Large Language Models (LLMs), the interaction between multimodal data, including image and audio has been transformed with a significant leap forward towards a bigger goal of artificial general intelligence. This workshop aims to bring together experts from diverse domains to explore the possibilities of developing novel ways of multimodal data search, understanding and interaction. Hui Wang 0001, Josef Kittler, Mark J. F. Gales, Rob Cooper, Maurice D. Mulvenna, Wing W. Y. Ng, Yang Hua 0001, Richard Gault, Abbas Haider, Guanfeng Wu |
ICMR | 6 |
| 2024 | A Mathematics Framework of Artificial Shifted Population Risk and Its Further Understanding Related to Consistency Regularization
Xiliang Yang, Shenyang Deng, Shicong Liu, Yuanchi Suo, Wing W. Y. Ng, Jianjun Zhang 0004 |
ECML/PKDD (1) | 5 |
| 2024 | On the Adversarial Robustness of Hierarchical ClassificationabstractDeep neural networks (DNNs) have demonstrated remarkable success on various learning problems, but they face a formidable challenge in the form of adversarial attacks. Especially, when dealing with complex classification tasks for numerous classes with a hierarchical structure, the adversarial robustness of a DNN model may drop seriously. In this paper, we investigate the adversarial robustness of DNN models on such complex classification tasks. In response, we propose a two-stage hierarchical classification framework, which is composed of a coarse-grained classifier and a series of fine-grained classifiers. A data correction sampling module is designed between the two stages, in order to mitigate the influence of misclassification caused by the coarse-grained classifier; and a discriminative filter learning module is employed in the fine-grained classification, in order to gain better distinguish abilities among fine-grained categories. Experiments on the well-known dataset CIFAR-100 and a newly-constructed hierarchical dataset mini-ImageNet76 demonstrate that employing a hierarchical framework can effectively improve the model robustness on such complex classification tasks. Ran Wang 0001, Simeng Zeng, Wenhui Wu 0001, Yuheng Jia, Wing W. Y. Ng, Xizhao Wang |
SMC | 5 |
| 2024 | AOCN: Appendix Object Correction Network Utilizing Relationships Across CT SlicesabstractWhen analyzing CT images of patients with suspected appendicitis, radiologists need to observe and examine consecutive 2D CT slices. Computer-assisted detection of the appendix in 2D CT slices significantly improve the diagnostic efficiency of radiologists. However, existing 2D medical image object detection methods primarily focus on spatial features within a single CT slice, which overlook spatial relationships between consecutive slices. We propose an Appendix Object Correction Network (AOCN) to refine predictions of universal object detectors. Although AOCN is a 2D network, it effectively leverages spatial relationships across consecutive CT slices. AOCN requires only a few training epochs to improve the accuracy of bounding boxes significantly, which offers advantages such as high scalability, low cost, and reduced training time. It consists of a global case feature learning module for extracting global feature map from the CT case and an object feature relation module for modeling the relationships between objects across slices. Experimental results demonstrate the effectiveness and efficiency of AOCN in correcting the output bounding boxes of several mainstream object detection networks, with a 6% to 14% improvement in Recall while requiring only a few training epochs. Wing W. Y. Ng, Yinhao Liang, Ting Wang 0015, Jianjun Zhang 0004, Xinhua Wei |
SMC | 1 |
| 2024 | MCD: Defense Against Query-Based Black-Box Surrogate AttacksabstractDeep neural networks (DNNs) is susceptible to surrogate attacks, where adversaries use surrogate data and corresponding outputs from the target model to build their own stolen model. Model stealing attacks jeopardize model privacy and model owners' commercial benefits. To address this issue, this paper proposes a hybrid protection approach-Maximize the confidence differences between benign samples and adversarial samples (MCD), to protect models from theft. Firstly, the LogitNorm approach is used to overcome the overconfidence problem in adversary query classification. Then, samples are divided into four groups according to ES and RS. Different groups are poisoned by different degrees. In addition to enhancing defensive performance and accounting for model integrity, the MCD uses a trigger to confirm the cloned model's owner. Experimental results show that the MCD defends against a variety of original models and attack techniques well. Against KnockoffNets and DFME attacks, the MCD yields an average defense performance of 54.58 % on five datasets, which is a great improvement over other defenses. Compared to other poisoning techniques, the Strong Poisoning (SP) module reduces the adversary's accuracy by 48.23 % on average. Additionally, the MCD overcomes the issue of OOD overconfidence while safeguarding the model accuracy in OOD detection and reduces the misclassification rate of ID samples for multiple OOD datasets. Yiwen Zou, Wing W. Y. Ng, Xueli Zhang, Brick Loo, Xingfu Yan, Ran Wang 0001 |
SMC | 2 |
| 2024 | Lightweight multimodal Cycle-Attention Transformer towards cancer diagnosis
Shicong Liu, Xin Ma 0023, Shenyang Deng, Yuanchi Suo, Jianjun Zhang 0004, Wing W. Y. Ng |
Expert Syst. Appl. | 6 |
| 2024 | Multi-modal video search by examples - A video quality impact analysisabstractAbstract As the proliferation of video content continues, and many video archives lack suitable metadata, therefore, video retrieval, particularly through example‐based search, has become increasingly crucial. Existing metadata often fails to meet the needs of specific types of searches, especially when videos contain elements from different modalities, such as visual and audio. Consequently, developing video retrieval methods that can handle multi‐modal content is essential. An innovative Multi‐modal Video Search by Examples (MVSE) framework is introduced, employing state‐of‐the‐art techniques in its various components. In designing MVSE, the authors focused on accuracy, efficiency, interactivity, and extensibility, with key components including advanced data processing and a user‐friendly interface aimed at enhancing search effectiveness and user experience. Furthermore, the framework was comprehensively evaluated, assessing individual components, data quality issues, and overall retrieval performance using high‐quality and low‐quality BBC archive videos. The evaluation reveals that: (1) multi‐modal search yields better results than single‐modal search; (2) the quality of video, both visual and audio, has an impact on the query precision. Compared with image query results, audio quality has a greater impact on the query precision (3) a two‐stage search process (i.e. searching by Hamming distance based on hashing, followed by searching by Cosine similarity based on embedding); is effective but increases time overhead; (4) large‐scale video retrieval is not only feasible but also expected to emerge shortly. Guanfeng Wu, Abbas Haider, Xing Tian, Erfan Loweimi, Chi-Ho Chan, Mengjie Qian 0001, Muhammad Junaid Awan, Ivor T. A. Spence, Rob Cooper, Wing W. Y. Ng, Josef Kittler, Mark J. F. Gales, Hui Wang 0001 |
IET Comput. Vis. | 10 |
| 2024 | Semantic dependency and local convolution for enhancing naturalness and tone in text-to-speech synthesis
Chenglong Jiang, Ying Gao 0004, Wing W. Y. Ng, Jiyong Zhou, Jinghui Zhong, Hongzhong Zhen, Xiping Hu |
Neurocomputing | 3 |
| 2024 | Low-rank matrix recovery with total generalized variation for defending adversarial examplesabstractLow-rank matrix decomposition with first-order total variation (TV) regularization exhibits excellent performance in exploration of image structure. Taking advantage of its excellent performance in image denoising, we apply it to improve the robustness of deep neural networks. However, although TV regularization can improve the robustness of the model, it reduces the accuracy of normal samples due to its over-smoothing. In our work, we develop a new low-rank matrix recovery model, called LRTGV, which incorporates total generalized variation (TGV) regularization into the reweighted low-rank matrix recovery model. In the proposed model, TGV is used to better reconstruct texture information without over-smoothing. The reweighted nuclear norm and L 1 -norm can enhance the global structure information. Thus, the proposed LRTGV can destroy the structure of adversarial noise while re-enhancing the global structure and local texture of the image. To solve the challenging optimal model issue, we propose an algorithm based on the alternating direction method of multipliers. Experimental results show that the proposed algorithm has a certain defense capability against black-box attacks, and outperforms state-of-the-art low-rank matrix recovery methods in image restoration. Wen Li 0015, Hengyou Wang, Qiang He 0003, Zhiquan He, Wing W. Y. Ng |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2024 | Deep supervised fused similarity hashing for cross-modal retrieval
Wing W. Y. Ng, Yongzhi Xu, Xing Tian, Hui Wang 0001 |
Multim. Tools Appl. | 1 |
| 2024 | Improving domain generalization by hybrid domain attention and localized maximum sensitivity
Wing W. Y. Ng, Cankun Zhong, Jianjun Zhang 0004 |
Neural Networks | 1 |
| 2024 | SBHA: Sensitive Binary Hashing Autoencoder for Image RetrievalabstractBinary hashing is an effective approach for content-based image retrieval, and learning binary codes with neural networks has attracted increasing attention in recent years. However, the training of hashing neural networks is difficult due to the binary constraint on hash codes. In addition, neural networks are easily affected by input data with small perturbations. Therefore, a sensitive binary hashing autoencoder (SBHA) is proposed to handle these challenges by introducing stochastic sensitivity for image retrieval. SBHA extracts meaningful features from original inputs and maps them onto a binary space to obtain binary hash codes directly. Different from ordinary autoencoders, SBHA is trained by minimizing the reconstruction error, the stochastic sensitive error, and the binary constraint error simultaneously. SBHA reduces output sensitivity to unseen samples with small perturbations from training samples by minimizing the stochastic sensitive error, which helps to learn more robust features. Moreover, SBHA is trained with a binary constraint and outputs binary codes directly. To tackle the difficulty of optimization with the binary constraint, we train the SBHA with alternating optimization. Experimental results on three benchmark datasets show that SBHA is competitive and significantly outperforms state-of-the-art methods for binary hashing. Ting Wang 0015, Su Lu, Jianjun Zhang 0004, Xuyu Liu, Xing Tian, Wing W. Y. Ng, Weineng Chen |
IEEE Trans. Cybern. | 6 |
| 2024 | Fog-Enabled Privacy-Preserving Multi-Task Data Aggregation for Mobile CrowdsensingabstractPrivacy-preserving data aggregation in mobile crowdsensing (MCS) focuses on mining information from massive sensing data while protecting users' privacy. The existence of multiple concurrent tasks is common in urban environments, so privacy-preserving multi-task data aggregation is essential and useful to a large-scale crowdsensing server. However, existing privacy-preserving data aggregation schemes in MCS mainly focus on the single-task data aggregation and the privacy protection of user's data. Little attention is paid to the privacy of user's decision of accepting tasks. Therefore, we propose a privacy-preserving and server-oriented efficient multi-task data aggregation scheme for MCS based fog computing. The proposed scheme can aggregate multiple concurrent tasks from multiple requesters (e.g., for 9 tasks, the proposed scheme completes all tasks in one round as opposed to existing schemes, which finish 9 tasks in nine rounds). Our scheme protects the privacy of user's decision, user's data, and aggregation result of each requester under collusion attacks. Through formal security analyses, our scheme is proved to be secure and privacy-preserving. Both theoretical analyses and experiments show our scheme is efficient. Xingfu Yan, Wing W. Y. Ng, Bowen Zhao 0001, Yuxian Liu, Ying Gao 0004, Xiumin Wang 0005 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | HRadNet: A Hierarchical Radiomics-Based Network for Multicenter Breast Cancer Molecular Subtypes PredictionabstractBreast cancer is a heterogeneous disease, where molecular subtypes of breast cancer are closely related to the treatment and prognosis. Therefore, the goal of this work is to differentiate between luminal and non-luminal subtypes of breast cancer. The hierarchical radiomics network (HRadNet) is proposed for breast cancer molecular subtypes prediction based on dynamic contrast-enhanced magnetic resonance imaging. HRadNet fuses multilayer features with the metadata of images to take advantage of conventional radiomics methods and general convolutional neural networks. A two-stage training mechanism is adopted to improve the generalization capability of the network for multicenter breast cancer data. The ablation study shows the effectiveness of each component of HRadNet. Furthermore, the influence of features from different layers and metadata fusion are also analyzed. It reveals that selecting certain layers of features for a specified domain can make further performance improvements. Experimental results on three data sets from different devices demonstrate the effectiveness of the proposed network. HRadNet also has good performance when transferring to other domains without fine-tuning. Yinhao Liang, Ting Wang 0015, Wing W. Y. Ng, Kuiming Jiang, Xinhua Wei, Xinqing Jiang |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Self-Supervised Temporal Sensitive Hashing for Video RetrievalabstractSelf-supervised video hashing methods retrieve large-scale video data without labels by making full use of visual and temporal information in original videos. Existing methods are not robust enough to handle small temporal differences between similar videos, because of the ignoring of future unseen samples on temporal which leads to large generalization errors. At the same time, existing self-supervised methods cannot preserve pairwise similarity information between large-scale unlabeled data efficiently and effectively. Thus, a self-supervised temporal sensitive video hashing (TSVH) is proposed in the paper for video retrieval. The TSVH uses a transformer-based autoencoder network with temporal sensitivity regularization to achieve low sensitivity of local temporal perturbations and preserve information of global temporal sequence. The pairwise similarity between video samples is effectively preserved by applying a hashing-based affinity matrix in the method. Experiments on realistic datasets show that the TSVH outperforms several state-of-the-art methods and classic methods. Qihua Li, Xing Tian, Wing W. Y. Ng |
IEEE Trans. Multim. | 3 |
| 2024 | BASS: Broad Network Based on Localized Stochastic SensitivityabstractThe training of the standard broad learning system (BLS) concerns the optimization of its output weights via the minimization of both training mean square error (MSE) and a penalty term. However, it degrades the generalization capability and robustness of BLS when facing complex and noisy environments, especially when small perturbations or noise appear in input data. Therefore, this work proposes a broad network based on localized stochastic sensitivity (BASS) algorithm to tackle the issue of noise or input perturbations from a local perturbation perspective. The localized stochastic sensitivity (LSS) prompts an increase in the network's noise robustness by considering unseen samples located within a Q -neighborhood of training samples, which enhances the generalization capability of BASS with respect to noisy and perturbed data. Then, three incremental learning algorithms are derived to update BASS quickly when new samples arrive or the network is deemed to be expanded, without retraining the entire model. Due to the inherent superiorities of the LSS, extensive experimental results on 13 benchmark datasets show that BASS yields better accuracies on various regression and classification problems. For instance, BASS uses fewer parameters (12.6 million) to yield 1% higher Top-1 accuracy in comparison to AlexNet (60 million) on the large-scale ImageNet (ILSVRC2012) dataset. Ting Wang 0015, Jianjun Zhang 0004, Wing W. Y. Ng, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | SeDepTTS: Enhancing the Naturalness via Semantic Dependency and Local Convolution for Text-to-Speech SynthesisabstractSelf-attention-based networks have obtained impressive performance in parallel training and global context modeling. However, it is weak in local dependency capturing, especially for data with strong local correlations such as utterances. Therefore, we will mine linguistic information of the original text based on a semantic dependency and the semantic relationship between nodes is regarded as prior knowledge to revise the distribution of self-attention. On the other hand, given the strong correlation between input characters, we introduce a one-dimensional (1-D) convolution neural network (CNN) producing query(Q) and value(V) in the self-attention mechanism for a better fusion of local contextual information. Then, we migrate this variant of the self-attention networks to speech synthesis tasks and propose a non-autoregressive (NAR) neural Text-to-Speech (TTS): SeDepTTS. Experimental results show that our model yields good performance in speech synthesis. Specifically, the proposed method yields significant improvement for the processing of pause, stress, and intonation in speech. Chenglong Jiang, Ying Gao 0004, Wing W. Y. Ng, Jiyong Zhou, Jinghui Zhong, Hongzhong Zhen |
AAAI | 3 |
| 2023 | LSSED: A Robust Segmentation Network for Inflamed Appendix from CT ImagesabstractAcute appendicitis (AA) is one of the most prevalent surgical acute abdominal condition diseases. The treatment management of A A is highly dependent on the CT image diagnosis. However, the in-flamed appendix exhibits blurred boundaries with nearby tissue, varying shapes, and sizes. These properties require high robustness and generalization capability of inflamed appendix segmentation networks. In this paper, we propose a CNN-Transformer-based encoder-decoder segmentation network (LSSED) equipped with localized stochastic sensitivity (LSS) loss function and residual dilated paths (RD-Paths) to solve above problems. The proposed method effectively learns robust features of the input data by reducing the LSS of unseen samples. In addition, the RD-Paths capture multiscale feature information and reduce the semantic gap between the encoder and decoder, which improves the accuracy of the segmentation. Empirical studies on a real-world AA dataset show that our method yields the best performance in terms of average Dice similarity coefficient (DSC) and Hausdorff Distance of 95% (HD95) compared to several state-of-the-art segmentation networks. Wing W. Y. Ng, Peixin Zheng, Ting Wang 0015, Jianjun Zhang 0004, Yinhao Liang, Xinhua Wei |
ICASSP | 1 |
| 2023 | A Lightweight and Efficient Model for Audio Anti-SpoofingabstractWith the rapid development of speech conversion and speech synthesis algorithms, automatic speaker verification (ASV) systems are vulnerable to spoofing attacks. In recent years, researchers had proposed anti-spoofing systems based on hand-crafted features. However, using hand-crafted features rather than raw waveform will lose implicit information for audio anti-spoofing. Inspired by the promising performance of ConvNeXt in classification tasks, we reference the network architecture design of ConvNeXt and propose a Lightweight and Efficient Model for Audio Anti-Spoofing (LEMAAS). With no preceding feature extraction process, we employ raw waveforms as direct inputs to our proposed model. By integrating with the channel attention module and using the focal loss function, the proposed model can focus on the most informative features representation of speech and the difficult samples that are hard to classify. Experimental results show that our proposed system could achieve an equal error rate of 0.64% and min-tDCF of 0.0187 for the ASVspoof 2019 LA evaluation dataset, which outperforms the state-of-the-art systems. Moreover, even when trained only on the ASVspoof 2019 LA dataset, the model still achieved equal error rates of 0.86% and 1.18% on the ASVspoof 2015 development dataset and evaluation dataset, respectively. This demonstrates that our model has achieved promising generalization performance during cross-dataset testing. Qiaowei Ma, Jinghui Zhong, Weiheng Liu, Ying Gao 0004, Wing W. Y. Ng |
MMAsia | 6 |
| 2023 | Moisture Content Prediction of Sugi Wood Drying Using Deep LSTM AE Minimizing Perturbed ErrorabstractWood drying technology plays a key role in extending service lifetime of wood as the moisture content has a great influence on the wood quality. This paper presents a moisture content prediction model based on the deep long short-term memory (LSTM) autoencoders with stochastic sensitivity (DLASS) to extract a hidden representation of input data. The DLASS uses multiple LSTM encoders to learn more informative hidden representations from unseen samples, which are then decoded using multiple LSTM decoders. The DLASS is trained by minimizing the perturbed error from historical moisture content data. Furthermore, a nonlinear fully connected feedforward neural network as a regression layer is applied to predict moisture content using hidden representations learned by the DLASS. The DLASS is applied to real-world industrial data of Sugi wood processed by a drying kiln made by SECEA from August 4 to 19, 2008. Multiple test cases and comparisons with existing classical and state-of-the-art models show that the DLASS model yields more accurate moisture content prediction results and has high generalization capability. To be specific, the DLASS yields the lowest MAE (0.026), MAPE (0.280), and RMSE (0.058) for predicting moisture content during the wood drying process. Ting Wang 0015, Wing W. Y. Ng, Xueli Zhang, Jianjun Zhang 0004, Mingcong Deng |
SMC | 2 |
| 2023 | AERF: Adaptive ensemble random fuzzy algorithm for anomaly detection in cloud computing
Jun Jiang 0003, Fagui Liu, Wing W. Y. Ng, Quan Tang 0001, Guoxiang Zhong, Xuhao Tang 0001, Bin Wang 0048 |
Comput. Commun. | 3 |
| 2023 | Multiscale echo self-attention memory network for multivariate time series classification
Huizi Lyu, Desen Huang, Sen Li 0001, Wing W. Y. Ng, Qianli Ma 0001 |
Neurocomputing | 4 |
| 2023 | Robust recurrent neural networks for time series forecasting
Xueli Zhang, Cankun Zhong, Jianjun Zhang 0004, Ting Wang 0015, Wing W. Y. Ng |
Neurocomputing | 5 |
| 2023 | Multi-object tracking for horse racing
Wing W. Y. Ng, Xuyu Liu, Xuli Yan, Xing Tian, Cankun Zhong, Sam Kwong |
Inf. Sci. | 1 |
| 2023 | Length adaptive hashing for semi-supervised semantic image retrieval
Si-chao Lei, Xing Tian, Wing W. Y. Ng, Yue-Jiao Gong |
Multim. Tools Appl. | 3 |
| 2023 | Deep Incremental Hashing for Semantic Image Retrieval With Concept DriftabstractHashing methods are widely used for content-based image retrieval due to their attractive time and space efficiencies. Several dynamic hashing methods have been proposed for image retrieval tasks in non-stationary environments. However, concept drift problems in non-stationary environment are seldomly considered which lead to significant deterioration of performance. Therefore, we propose Deep Incremental Hashing (DIH). For the learning part, similarity-preserving object codes of each newly arriving data chunk are computed using the product of its label matrix and a random Gaussian matrix generated offline. A point-wise loss function is then devised to guide the learning of a deep hash neural network. To retain the learned knowledge of former chunks, a weighting-based method is utilized to combine different hash tables trained at different time steps to form a multi-table hashing system. Experimental results on 13 simulated concept drift environments show that DIH adapts to non-stationary data environments well and yields better retrieval performance than existing dynamic hashing methods. Xing Tian, Wing W. Y. Ng |
IEEE Trans. Big Data | 2 |
| 2023 | Knowledge Distillation Hashing for Occluded Face RetrievalabstractDeep hashing has proven to be efficient and effective for large-scale face retrieval. However, existing hashing methods are designed for normal face images only. They fail to consider the fact that face images may be occluded because of wearing masks, hats, glasses, etc. Retrieval performance of existing face retrieval methods is much worse when dealing with occluded face images. In this work, we propose the knowledge distillation hashing (KDH) to deal with occluded face images. The KDH is a two-stage learning approach with teacher-student model distillation. We first train a teacher hashing network using normal face images and then the knowledge from teacher model is used to guide the optimization of the student model using occluded face images as input only. With knowledge distillation, we build a connection between imperfect face information and the optimal hash codes. Experimental results show that the KDH yields significant improvements and better retrieval performance in comparison to existing state-of-the-art deep hashing retrieval methods under six different face occlusion situations. Xing Tian, Wing W. Y. Ng, Ying Gao 0004 |
IEEE Trans. Multim. | 3 |
| 2023 | A Robust Frequency-Domain-Based Graph Adaptive Network for Parkinson's Disease Detection From Gait DataabstractParkinson's disease (PD) is a neurodegenerative disease with a high incidence rate. Effective early diagnosis of PD is critical to prevent further deterioration of a patient's condition, where gait abnormalities are important factors for doctors to diagnose PD. Deep learning (DL)-based methods for PD detection using gait information recorded by non-invasive sensors have emerged to assist doctors in accurate and efficient disease diagnosis. However, most existing DL-based PD detection models neglect information in the frequency domain and do not adaptively model the correlation of signals among sensors. Moreover, different people have different gait patterns. Therefore, the generalization capabilities of PD detection models on diversities of individuals' gaits are essential. This work proposes a novel robust frequency-domain-based graph adaptive network (RFdGAD) for PD detection from gait information (i.e., vertical ground reaction force signals recorded by foot sensors). Specifically, the RFdGAD first learns the frequency-domain features of signals from each foot sensor by a frequency representation learning block. Then, the RFdGAD utilizes a graph adaptive network block taking frequency-domain features as input to adaptively learn and exploit the interconnection between different sensor signals for accurate PD detection. Moreover, the RFdGAD is trained by minimizing the proposed Jensen-Shannon divergence-based localized generalization error to improve the generalization performance of RFdGAD on unseen subjects. Experimental results show that the RFdGAD outperforms existing DL-based models for PD detection on three widely used datasets in terms of three metrics, including accuracy, F1-score, and geometric mean. Cankun Zhong, Wing W. Y. Ng |
IEEE Trans. Multim. | 2 |
| 2023 | Population-Based Hyperparameter Tuning With Multitask CollaborationabstractPopulation-based optimization methods are widely used for hyperparameter (HP) tuning for a given specific task. In this work, we propose the population-based hyperparameter tuning with multitask collaboration (PHTMC), which is a general multitask collaborative framework with parallel and sequential phases for population-based HP tuning methods. In the parallel HP tuning phase, a shared population for all tasks is kept and the intertask relatedness is considered to both yield a better generalization ability and avoid data bias to a single task. In the sequential HP tuning phase, a surrogate model is built for each new-added task so that the metainformation from the existing tasks can be extracted and used to help the initialization for the new task. Experimental results show significant improvements in generalization abilities yielded by neural networks trained using the PHTMC and better performances achieved by multitask metalearning. Moreover, a visualization of the solution distribution and the autoencoder's reconstruction of both the PHTMC and a single-task population-based HP tuning method is compared to analyze the property with the multitask collaboration. Wendi Li, Ting Wang 0015, Wing W. Y. Ng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | KNNENS: A k-Nearest Neighbor Ensemble-Based Method for Incremental Learning Under Data Stream With Emerging New ClassesabstractIn this brief, we investigate the problem of incremental learning under data stream with emerging new classes (SENC). In the literature, existing approaches encounter the following problems: 1) yielding high false positive for the new class; i) having long prediction time; and 3) having access to true labels for all instances, which is unrealistic and unacceptable in real-life streaming tasks. Therefore, we propose the k -Nearest Neighbor ENSemble-based method (KNNENS) to handle these problems. The KNNENS is effective to detect the new class and maintains high classification performance for known classes. It is also efficient in terms of run time and does not require true labels of new class instances for model update, which is desired in real-life streaming classification tasks. Experimental results show that the KNNENS achieves the best performance on four benchmark datasets and three real-world data streams in terms of accuracy and F1-measure and has a relatively fast run time compared to four reference methods. Codes are available at https://github.com/Ntriver/KNNENS. Jianjun Zhang 0004, Ting Wang 0015, Wing W. Y. Ng, Witold Pedrycz |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | A Sensitivity-based Pruning Method for Convolutional Neural NetworksabstractThe application of convolutional neural networks (CNNs) is sometimes limited by a large number of parameters and floating-point operations. Pruning methods have been proved to be effective to solve this problem. These methods improve the efficiency and storage occupancy of CNNs by removing weights connected with certain neurons/channels. The key issue is the selection of suitable neurons/channels to be pruned. Then, fine-tuning is usually applied to restore the performance of a pruned model to that before the pruning. However, existing neurons/channels selection methods do not explicitly consider the impact of the pruning on the model output. Moreover, the performance of a fine-tuned model may suffer from the information loss problem caused by the pruned neurons/channels. In this work, a stochastic sensitivity measure-based neurons/channels selection criterion is proposed to choose and prune insensitive neurons/channels, which effectively reduces the degradation of model performance. Moreover, a compensation operation followed by fine-tuning is proposed to relieve the information loss problem and restore model performance. Experimental results show that our method yields comparable compression and acceleration rates with less accuracy degradation compared with existing pruning methods for CNNs. For instance, the proposed method achieves more 6.8% FLOPs reduction and 0.25% accuracy improvement on VGG-16 compared with a recently proposed pruning method. Cankun Zhong, Yang He 0002, Yifei An, Wing W. Y. Ng, Ting Wang 0015 |
SMC | 4 |
| 2022 | Hashing-based affinity matrix for dominant set clustering
Qihua Li, Xing Tian, Wing W. Y. Ng, Marcello Pelillo |
Neurocomputing | 3 |
| 2022 | Ensembling perturbation-based oversamplers for imbalanced datasets
Jianjun Zhang 0004, Ting Wang 0015, Wing W. Y. Ng, Witold Pedrycz |
Neurocomputing | 3 |
| 2022 | P2SIM: Privacy-Preserving and Source-Reliable Incentive Mechanism for Mobile CrowdsensingabstractIn mobile crowdsensing (MCS), providing appropriate rewards is a common and efficient way to motivate participants to participate in sensing tasks. However, the privacy of task participants is not protected well in most quality-aware incentive schemes. Moreover, these schemes are designed for general MCS application scenarios where data are collected by internal sensors embedded in participants’ smartphones, and not suitable for scenarios where additional sensors (ASs) except internal sensors are to collect data (e.g., household medical devices). In scenarios with ASs, malicious participants can fabricate sensing data instead of collecting data from ASs, i.e., the source reliability of sensing data cannot be ensured. To address these issues, we propose P2SIM, a privacy-preserving and the source-reliable incentive mechanism scheme for MCS with ASs. We combine redactable signature with private hash function to achieve the source reliability verification of sensing data without revealing the privacy of participants. Moreover, rewards are divided into two parts: 1) fixed rewards and 2) floating rewards, to enhance the flexibility of rewards distribution. Both formal theoretical analysis and extensive experimental evaluations on a real data set show that the proposed P2SIM is secure and efficient. Xingfu Yan, Wing W. Y. Ng, Bowen Zhao 0001, Ying Gao 0004 |
IEEE Internet Things J. | 2 |
| 2022 | Bit-wise attention deep complementary supervised hashing for image retrieval
Wing W. Y. Ng, Jiayong Li, Xing Tian, Hui Wang 0001 |
Multim. Tools Appl. | 1 |
| 2022 | A Deep Clustering via Automatic Feature Embedded Learning for Human Activity RecognitionabstractTraditional clustering algorithms are widely used for building bag-of-words (BOW) models to aggregate spatio-temporal feature points extracted from a video for human activity recognition problems. Their performances are restricted by the computational complexity which limits the number of feature points being used. In contrast, deep clustering yields good clustering performance without the limit of the number of feature points. Therefore, this work proposes a dual stacked autoencoders features embedded clustering (DSAFEC) and a BOW construction method based on the DSAFEC (B-DSAFEC) to reduce the computational complexity and to remove the selection restriction. The DSAFEC first transforms feature points extracted from a video to a learned feature space and then probabilities of cluster assignment of feature points are predicted to build BOWs for human activity recognition. A soft clustering is used by assigning each feature point to multiple clusters yielding the largest probabilities instead of only one in hard clustering. Experimental results on three benchmark human activity datasets show that the B-DSAFEC yields better performance compared to five reference methods which are developed based on either traditional clustering methods or deep clustering methods. Ting Wang 0015, Wing W. Y. Ng, Jinde Li, Qiuxia Wu, Shuai Zhang 0001, Chris D. Nugent, Colin Shewell |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Difference-Guided Representation Learning Network for Multivariate Time-Series ClassificationabstractMultivariate time series (MTSs) are widely found in many important application fields, for example, medicine, multimedia, manufacturing, action recognition, and speech recognition. The accurate classification of MTS has become an important research topic. Traditional MTS classification methods do not explicitly model the temporal difference information of time series, which is, in fact, important and reflects the dynamic evolution information. In this article, the difference-guided representation learning network (DGRL-Net) is proposed to guide the representation learning of time series by dynamic evolution information. The DGRL-Net consists of a difference-guided layer and a multiscale convolutional layer. First, in the difference-guided layer, we propose a difference gating LSTM to model the time dependency and dynamic evolution of the time series to obtain feature representations of both raw and difference series. Then, these two representations are used as two input channels of the multiscale convolutional layer to extract multiscale information. Extensive experiments demonstrate that the proposed model outperforms state-of-the-art methods on 18 MTS benchmark datasets and achieves competitive results on two skeleton-based action recognition datasets. Furthermore, the ablation study and visualized analysis are designed to verify the effectiveness of the proposed model. Qianli Ma 0001, Shuai Tian, Wing W. Y. Ng |
IEEE Trans. Cybern. | 4 |
| 2022 | Hashing-Based Undersampling Ensemble for Imbalanced Pattern Classification ProblemsabstractUndersampling is a popular method to solve imbalanced classification problems. However, sometimes it may remove too many majority samples which may lead to loss of informative samples. In this article, the hashing-based undersampling ensemble (HUE) is proposed to deal with this problem by constructing diversified training subspaces for undersampling. Samples in the majority class are divided into many subspaces by a hashing method. Each subspace corresponds to a training subset which consists of most of the samples from this subspace and a few samples from surrounding subspaces. These training subsets are used to train an ensemble of classification and regression tree classifiers with all minority class samples. The proposed method is tested on 25 UCI datasets against state-of-the-art methods. Experimental results show that the HUE outperforms other methods and yields good results on highly imbalanced datasets. Wing W. Y. Ng, Shichao Xu, Jianjun Zhang 0004, Xing Tian, Tongwen Rong, Sam Kwong |
IEEE Trans. Cybern. | 1 |
| 2022 | An Intelligent Interaction Framework for Teleoperation Based on Human-Machine CooperationabstractMost existing teleoperation technologies cannot guide robots to complete tasks quickly and accurately in unstructured environments. Therefore, this article proposes an intelligent interaction framework for teleoperation based on human-machine cooperation. The framework is divided into three layers: 1) perception, 2) decision-making, and 3) execution layers. In the perception layer, machines require high-precision environmental information, and human interaction requires rapid feedback on environmental changes. Therefore, a fast three-dimensional reconstruction method combining rough reconstruction and fine reconstruction is proposed. In the decision-making layer, a bare-hand interaction method based on natural interaction and an alignment assistance method based on constraint recognition are proposed. The combination of these two methods realizes the coordinated control of robot movement by humans and computers. In the execution layer, a vision-based error compensation method for assistance is proposed to decrease measurement errors and delays and reduce differences between virtual and real scenes. Finally, experimental results obtained by 15 nonprofessional volunteers show that the proposed method is efficient and user friendly. Guanglong Du, Yongda Deng, Wing W. Y. Ng, Di Li 0001 |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2022 | Two-Echelon Dispatching Problem With Mobile Satellites in City LogisticsabstractAt present, city logistics mostly adopts a two-echelon dispatching model which combines distribution centers located in suburbs and fixed satellites located in urban areas for distribution. However, both expensive rental fees and daily changes of customer demand in metropolitan areas make dispatching route generated by fixed satellites inefficient. Moreover, the existing mobile depot model needs a large investment for facilities. In this paper, we propose a two-echelon city dispatching model with mobile satellites (2ECD-MS) which locations of mobile satellites change according to demands of customers to ensure the efficiency of delivery routes in every day. A cluster-based variable neighborhood search scheduling algorithm is proposed to determine locations of mobile satellites and dispatching routes of trucks and tricycles. Then, the 2ECD-MS is extended to 2ECD-MS-TDD to allow trucks dispatching directly (TDD) for further cost reduction. Experimental results show that the 2ECD-MS significantly reduces the total cost against the model using fixed satellites mode by 3.5% while the 2ECD-MS-TDD further reduces the total cost against the 2ECD-MS significantly by 3.25% in 54 cases with different customer scales, geographical scopes, and distribution types. These show the superiority of the proposed methods in cost reduction for city logistics in comparison to the traditional fixed model. Yulin Lan, Fagui Liu, Zhixing Huang, Wing W. Y. Ng, Jinghui Zhong |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Multi-Objective Two-Echelon City Dispatching Problem With Mobile Satellites and Crowd-ShippingabstractRecently, a two-echelon city dispatching model with mobile satellites (2ECD-MS) has been proposed to reduce costs effectively. However, in addition to costs, speeds of delivery to customers are increasingly demanding in urban dispatching. This work extends 2ECD-MS to 2ECD-MS-CS by adopting the crowd-shipping model in the second-echelon dispatching, which uses occasional drivers of private vehicles to deliver parcels to improve the delivery speed. Furthermore, existing works generally consider the optimization from a single aspect, e.g., the delivery company. However, the sustainable development of a logistics company must also focus on other subjects in logistics activities, such as customers and delivery employees. So, we define a multi-objective model considering company cost, customer satisfaction, and income satisfaction of crowd-shippers simultaneously. The multi-objective optimization problem of 2ECD-MS-CS is solved by a multi-directional evolutionary algorithm (MDEA). In MDEA, multiple neighborhood operators are designed and combined with the multi-directional search strategy to fully explore the Pareto Front. Finally, we generate 40 new 2ECD-MS-CS instances based on existing common vehicle routing datasets. Experimental results show that 2ECD-MS-CS reduces the average cost by 3.4% and improves the delivery speed by 42% against 2ECD-MS in 40 instances with different customer scales, numbers of mobile satellites, and geographic scopes. The proposed MDEA outperforms several popular multi-objective optimization algorithms in both convergence and diversity. These illustrate the advantages of 2ECD-MS-CS especially in terms of delivery speed and the effectiveness of the proposed MDEA. Yulin Lan, Fagui Liu, Wing W. Y. Ng, Mengke Gui, Chengqi Lai |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Multi-Localized Sensitive Autoencoder-Attention-LSTM For Skeleton-based Action RecognitionabstractOne of key challenges of skeleton-based action recognition (SAR) tasks is the complex nature of human motion patterns. Variations such as performers and viewpoints may impose negative effects to the action recognition accuracy. In this work, we propose the Multi-Localized Sensitive Autoencoder-Attention-LSTM (Multi-LiSAAL) for SAR. The Localized Stochastic Sensitive Autoencoder (LiSSA) encodes both spatial and temporal information, and extracts meaningful features from different parts (four limbs and a trunk) from the skeleton. The LiSSA is trained by minimizing the localized generalization error to enhance the robustness of autoencoders via reducing its sensitivity with respect to small variations in inputs. We apply an attention mechanism to assign different weights to different skeleton parts and focus more on informative sections. Then, a backbone classifier network takes weighted features as inputs to differentiates actions. Experimental results on five public benchmarking datasets show that the Multi-LiSAAL outperforms state-of-the-art methods. Wing W. Y. Ng, Ting Wang 0015 |
IEEE Trans. Multim. | 1 |
| 2021 | Learning Efficient Rotation Representation for Point Cloud via Local-Global AggregationabstractRecently, there have been attempted to solve the problem of rotation perturbation in point cloud analysis. However, most of them fail to exploit the long-distance context and lose global location information. To address this issue, we propose a novel rotation-invariant network called LGANet, which is assembled with two key modules: local representation learning module and global alignment module. The local representation learning module is to capture local geometric features from K-nearest neighbors in both 3D Cartesian space and a latent space, while the global alignment module focuses on supplementing global location information with the adaptive selection mechanism. Extensive experiments on widely used datasets have demonstrated that our LGANet is superior to other state-of-the-art methods on the premise of ensuring rotation invariance in both classification and part segmentation. Ruibin Gu, Qiuxia Wu, Wing W. Y. Ng, Zhiyong Wang 0001 |
ICME | 4 |
| 2021 | A deep learning based hybrid method for hourly solar radiation forecasting
Chun Sing Lai, Cankun Zhong, Keda Pan, Wing W. Y. Ng, Loi Lei Lai |
Expert Syst. Appl. | 4 |
| 2021 | HELP: An LSTM-based approach to hyperparameter exploration in neural network learning
Wendi Li, Wing W. Y. Ng, Ting Wang 0015, Marcello Pelillo, Sam Kwong |
Neurocomputing | 2 |
| 2021 | Verifiable, Reliable, and Privacy-Preserving Data Aggregation in Fog-Assisted Mobile CrowdsensingabstractFog-assisted mobile crowdsensing (FA-MCS) alleviates challenges with respect to computation, communication, and storage from the traditional model of mobile crowdsensing (MCS) “requester-server-users.” Data aggregation, as a specific MCS task, has attracted a lot of attentions in mining the potential value of the massive crowdsensing data. However, the process of data aggregation in FA-MCS may threaten the privacies of both users' data and aggregation results. The untrusted server and fog nodes (FNs) may damage the correctness of aggregation results. Moreover, bad FNs, which do not upload data to server or fail to verify successfully, can endanger the reliability of FA-MCS and the accuracy of aggregation results. To tackle these problems, we propose a verifiable, reliable, and privacy-preserving data aggregation scheme for FA-MCS. Specifically, the proposed scheme preserves privacies of both users' data and aggregation results, enables requester to verify the correctness of aggregation result, and is able to tolerate several bad FNs without affecting the data aggregation result. Through formal security analysis, the proposed scheme is shown to be secure and privacy preserving. Extensive experiments also show the proposed scheme is efficient and reliable. Xingfu Yan, Wing W. Y. Ng, Changlu Lin, Yuxian Liu, Lu Lu 0011, Ying Gao 0004 |
IEEE Internet Things J. | 2 |
| 2021 | D4Net: De-deformation defect detection network for non-rigid products with large patterns
Xuemiao Xu, Huaidong Zhang, Wing W. Y. Ng |
Inf. Sci. | 4 |
| 2021 | ERINet: Enhanced rotation-invariant network for point cloud classification
Ruibin Gu, Qiuxia Wu, Wing W. Y. Ng, Zhiyong Wang 0001 |
Pattern Recognit. Lett. | 3 |
| 2021 | Convolutional Multitimescale Echo State NetworkabstractAs efficient recurrent neural network (RNN) models, echo state networks (ESNs) have attracted widespread attention and been applied in many application domains in the last decade. Although they have achieved great success in modeling time series, a single ESN may have difficulty in capturing the multitimescale structures that naturally exist in temporal data. In this paper, we propose the convolutional multitimescale ESN (ConvMESN), which is a novel training-efficient model for capturing multitimescale structures and multiscale temporal dependencies of temporal data. In particular, a multitimescale memory encoder is constructed with a multireservoir structure, in which different reservoirs have recurrent connections with different skip lengths (or time spans). By collecting all past echo states in each reservoir, this multireservoir structure encodes the history of a time series as nonlinear multitimescale echo state representations (MESRs). Our visualization analysis verifies that the MESRs provide better discriminative features for time series. Finally, multiscale temporal dependencies of MESRs are learned by a convolutional layer. By leveraging the multitimescale reservoirs followed by a convolutional learner, the ConvMESN has not only efficient memory encoding ability for temporal data with multitimescale structures but also strong learning ability for complex temporal dependencies. Furthermore, the training-free reservoirs and the single convolutional layer provide high-computational efficiency for the ConvMESN to model complex temporal data. Extensive experiments on 18 multivariate time series (MTS) benchmark datasets and 3 skeleton-based action recognition datasets demonstrate that the ConvMESN captures multitimescale dynamics and outperforms existing methods. Qianli Ma 0001, Enhuan Chen, Zhenxi Lin, Jiangyue Yan, Zhiwen Yu 0002, Wing W. Y. Ng |
IEEE Trans. Cybern. | 6 |
| 2021 | Concept Preserving Hashing for Semantic Image Retrieval With Concept DriftabstractCurrent hashing-based image retrieval methods mostly assume that the database of images is static. However, this assumption is not true in cases where the databases are constantly updated (e.g., on the Internet) and there exists the problem of concept drift. The online (also known as incremental) hashing methods have been proposed recently for image retrieval where the database is not static. However, they have not considered the concept drift problem. Moreover, they update hash functions dynamically by generating new hash codes for all accumulated data over time which is clearly uneconomical. In order to solve these two problems, concept preserving hashing (CPH) is proposed. In contrast to the existing methods, CPH preserves the original concept, that is, the set of hash codes representing a concept is preserved over time, by learning a new set of hash functions to yield the same set of hash codes for images (old and new) of a concept. The objective function of CPH learning consists of three components: 1) isomorphic similarity; 2) hash codes partition balancing; and 3) heterogeneous similarity fitness. The experimental results on 11 concept drift scenarios show that CPH yields better retrieval precisions than the existing methods and does not need to update hash codes of previously stored images. Xing Tian, Wing W. Y. Ng, Hui Wang 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | LiSSA: Localized Stochastic Sensitive AutoencodersabstractThe training of autoencoder (AE) focuses on the selection of connection weights via a minimization of both the training error and a regularized term. However, the ultimate goal of AE training is to autoencode future unseen samples correctly (i.e., good generalization). Minimizing the training error with different regularized terms only indirectly minimizes the generalization error. Moreover, the trained model may not be robust to small perturbations of inputs which may lead to a poor generalization capability. In this paper, we propose a localized stochastic sensitive AE (LiSSA) to enhance the robustness of AE with respect to input perturbations. With the local stochastic sensitivity regularization, LiSSA reduces sensitivity to unseen samples with small differences (perturbations) from training samples. Meanwhile, LiSSA preserves the local connectivity from the original input space to the representation space that learns a more robustness features (intermediate representation) for unseen samples. The classifier using these learned features yields a better generalization capability. Extensive experimental results on 36 benchmarking datasets indicate that LiSSA outperforms several classical and recent AE training methods significantly on classification tasks. Ting Wang 0015, Wing W. Y. Ng, Marcello Pelillo, Sam Kwong |
IEEE Trans. Cybern. | 2 |
| 2021 | SALMNet: A Structure-Aware Lane Marking Detection NetworkabstractLane marking detection is a fundamental task, which serves as an important prerequisite for automatic driving or driver-assistance systems. However, the complex and uncontrollable driving road environment as well as the discontinuous lane marking appearance make this task challenging. In this work, a novel deep neural network architecture is presented to detect lane markings in a complex environment by analyzing their structure information. There are two contributions to the network design. Firstly, a semantic-guided channel attention (SGCA) module is developed to select the low-level features of a deep convolutional neural network by taking the high-level features as the guidance. Secondly, a pyramid deformable convolution (PDC) module is formulated to enlarge the receptive fields and to capture the complex structures of lane markings by applying deformable convolutions on multiple feature maps with different scales. Hence, our network can better reduce false detection and enhance lane marking structures simultaneously. The experimental results on three benchmark datasets for lane marking detection show that our method outperforms other methods on all the benchmark datasets. Xuemiao Xu, Tianfei Yu, Xiaowei Hu 0001, Wing W. Y. Ng, Pheng-Ann Heng |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Complementary Incremental Hashing With Query-Adaptive Re-Ranking for Image RetrievalabstractConcept drift is prevalent in non-stationary data environments but is rarely researched in image retrieval. Therefore, more research is needed on image retrieval in non-stationary data environments so that highly relevant images can still be retrieved when concept drifts happen. Hashing is a key technique to allow efficient image retrieval, so incremental hashing technique emerges in recent years for image retrieval in non-stationary environments. A state-of-the-art method isIncremental Hashing(ICH). ICH trains new hash tables on new data without considering the performance of previous hash tables, so the dependency of successive hash tables is ignored. To make use of this dependency in order to improve the performance of image retrieval in non-stationary environments,Complementary Incremental Hashing with query-adaptive Re-ranking(CIHR) is proposed in this paper. CIHR trains multiple hash tables incrementally, one for each data chunk of images. A new hash table is trained on a new data chunk of images as well as those images badly hashed by previous hash tables, thus the new hash table is complementary to the previous hash tables. To use the hash tables more effectively, a query-adaptive re-ranking method is used to weight all hash functions in each hash table according to their retrieval performance with respect to a given query. Weighted Hamming distance is finally used to evaluate the similarity between the query and the images in the database, as the basis of image retrieval. Experimental results on simulated non-stationary scenarios show that the proposed CIHR method achieves higher retrieval accuracy than all methods being compared, thus setting a new state of the art in image retrieval in non-stationary data environments. Xing Tian, Wing W. Y. Ng, Hui Wang 0001, Sam Kwong |
IEEE Trans. Multim. | 2 |
| 2020 | Minority Oversampling Using SensitivityabstractThe Synthetic Minority Oversampling Technique (SMOTE) is effective to handle imbalance classification problems. However, the random candidate selection of SMOTE may lead to severe overlap between classes and introduce new noise factors. Many variants of SMOTE have been proposed to relieve these problems by generating new examples in safe regions. Most of these methods generate new examples with existing minority examples without considering the negative impact that class imbalance have brought on these examples. In this paper, we handle the imbalance classification using Bayes' decision rule and propose a novel oversampling method, the Minority Oversampling using Sensitivity (MOSS). Candidates for new example generations are selected considering their sensitivity with respect to class imbalance. New examples are then generated by interpolating the candidate and one of its adjacent examples. Experiments on 30 datasets confirm the superiority of the MOSS against one baseline method and seven oversampling methods. Jianjun Zhang 0004, Ting Wang 0015, Wing W. Y. Ng, Witold Pedrycz, Shuai Zhang 0001, Chris D. Nugent |
IJCNN | 3 |
| 2020 | Concurrent optimization of multiple base learners in neural network ensembles: An adaptive niching differential evolution approach
Ting Huang 0001, Danting Duan, Yue-Jiao Gong, Long Ye, Wing W. Y. Ng, Jun Zhang 0003 |
Neurocomputing | 5 |
| 2020 | Multi-level supervised hashing with deep features for efficient image retrieval
Wing W. Y. Ng, Jiayong Li, Xing Tian, Hui Wang 0001, Sam Kwong, Jonathan Wallace |
Neurocomputing | 1 |
| 2020 | Bootstrap dual complementary hashing with semi-supervised re-ranking for image retrieval
Xing Tian, Xiancheng Zhou, Wing W. Y. Ng, Jiayong Li, Hui Wang 0001 |
Neurocomputing | 3 |
| 2020 | Within-class multimodal classificationabstractAbstract In many real-world classification problems there exist multiple subclasses (or clusters) within a class; in other words, the underlying data distribution is within-class multimodal. One example is face recognition where a face (i.e. a class) may be presented in frontal view or side view, corresponding to different modalities. This issue has been largely ignored in the literature or at least under studied. How to address the within-class multimodality issue is still an unsolved problem. In this paper, we present an extensive study of within-class multimodality classification. This study is guided by a number of research questions, and conducted through experimentation on artificial data and real data. In addition, we establish a case for within-class multimodal classification that is characterised by the concurrent maximisation of between-class separation, between-subclass separation and within-class compactness. Extensive experimental results show that within-class multimodal classification consistently leads to significant performance gains when within-class multimodality is present in data. Furthermore, it has been found that within-class multimodal classification offers a competitive solution to face recognition under different lighting and face pose conditions. It is our opinion that the case for within-class multimodal classification is established, therefore there is a milestone to be achieved in some machine learning algorithms (e.g. Gaussian mixture model) when within-class multimodal classification, or part of it, is pursued. Huan Wan, Hui Wang 0001, Bryan W. Scotney, Jun Liu 0001, Wing W. Y. Ng |
Multim. Tools Appl. | 5 |
| 2019 | Incremental Hashing with UndersamplingabstractMost of current hashing methods are proposed based on the assumption that the database is stationary. However, this assumption is not always true as the data environment is sometimes non-stationary. When new images being added to the database, data distributions of existing classes may change and new classes may also appear which result in concept drifts. The problem of concept drifts is unavoidable in non-stationary data environments. Incremental Hashing (ICH) is an effective method for image retrieval in non-stationary data environments with concept drifts using multiple hash tables. In ICH, new concept is adapted by training new hash table using the most updated data chunks. However, images in the new data chunk may not be all informative for updating. To enhance the efficiency of ICH, ICH with Undersampling (ICHUS) is proposed to select informative samples in the new data chunk for the training of new hash table to adapt to the non-stationary data environment. Experimental results show that ICHUS yields a better retrieval performance than ICH and state-of-art non-stationary hashing methods. Xiaoxia Jiang, Wing W. Y. Ng, Xing Tian, Sam Kwong, Hui Wang 0001 |
SMC | 2 |
| 2019 | A robust correlation analysis framework for imbalanced and dichotomous data with uncertainty
Chun Sing Lai, Yingshan Tao, Wing W. Y. Ng, Youwei Jia, Chao Huang 0002, Loi Lei Lai, Zhao Xu 0002, Giorgio Locatelli |
Inf. Sci. | 4 |
| 2019 | Attention-based spatio-temporal dependence learning network
Qianli Ma 0001, Shuai Tian, Jia Wei 0003, Jiabing Wang, Wing W. Y. Ng |
Inf. Sci. | 5 |
| 2019 | Global-Local Mutual Attention Model for Text ClassificationabstractText classification is a central field of inquiry in natural language processing (NLP). Although some models learn local semantic features and global long-term dependencies simultaneously, they simply combine them through concatenation either in a cascade way or in parallel while mutual effects between them are ignored. In this paper, we propose the Global-Local Mutual Attention (GLMA) model for text classification problems, which introduces a mutual attention mechanism for mutual learning between local semantic features and global long-term dependencies. The mutual attention mechanism consists of a Local-Guided Global-Attention (LGGA) and a Global-Guided Local-Attention (GGLA). The LGGA allows to assign weights and combine global long-term dependencies of word positions that are semantic related. It captures combined semantics and alleviates the gradient vanishing problem. The GGLA automatically assigns more weights to relevant local semantic features, which captures key local semantic information and filters both noises and irrelevant words/phrases. Furthermore, a weighted-over-time pooling operation is developed to aggregate the most informative and discriminative features for classification. Extensive experiments demonstrate that our model obtains the state-of-the-art performance on seven benchmark datasets and sixteen Amazon product reviews datasets. Both the result analysis and the mutual attention weights visualization further demonstrate the effectiveness of the proposed model. Qianli Ma 0001, Liuhong Yu, Shuai Tian, Enhuan Chen, Wing W. Y. Ng |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | Incremental Hash-Bit Learning for Semantic Image Retrieval in Nonstationary EnvironmentsabstractImages are uploaded to the Internet over time which makes concept drifting and distribution change in semantic classes unavoidable. Current hashing methods being trained using a given static database may not be suitable for nonstationary semantic image retrieval problems. Moreover, directly retraining a whole hash table to update knowledge coming from new arriving image data may not be efficient. Therefore, this paper proposes a new incremental hash-bit learning method. At the arrival of new data, hash bits are selected from both existing and newly trained hash bits by an iterative maximization of a 3-component objective function. This objective function is also used to weight selected hash bits to re-rank retrieved images for better semantic image retrieval results. The three components evaluate a hash bit in three different angles: 1) information preservation; 2) partition balancing; and 3) bit angular difference. The proposed method combines knowledge retained from previously trained hash bits and new semantic knowledge learned from the new data by training new hash bits. In comparison to table-based incremental hashing, the proposed method automatically adjusts the number of bits from old data and new data according to the concept drifting in the given data via the maximization of the objective function. Experimental results show that the proposed method outperforms existing stationary hashing methods, table-based incremental hashing, and online hashing methods in 15 different simulated nonstationary data environments. Wing W. Y. Ng, Xing Tian, Witold Pedrycz, Xizhao Wang, Daniel S. Yeung |
IEEE Trans. Cybern. | 1 |
| 2019 | Cost-Sensitive Weighting and Imbalance-Reversed Bagging for Streaming Imbalanced and Concept Drifting in Electricity Pricing ClassificationabstractIn data streaming environments such as a smart grid, it is impossible to restrict each data chunk to have the same number of samples in each class. Hence, in addition to the concept drift, classification problems in streaming data environments are inherently imbalanced. However, streaming imbalanced and concept drifting problems in the power system and smart grid have rarely been studied. Incremental learning aims to learn the correct classification for the future unseen samples from the given streaming data. In this paper, we propose a new incremental ensemble learning method to handle both concept drift and class imbalance issues. The class imbalance issue is tackled by an imbalance-reversed bagging method that improves the true positive rate while maintains a low false positive rate. The adaptation to concept drift is achieved by a dynamic cost-sensitive weighting scheme for component classifiers according to their classification performances and stochastic sensitivities. The proposed method is applied to a case study for the electricity pricing in Australia to predict whether the price of New South Wales will be higher or lower than that of Victorias in a 24-h period. Experimental results show the effectiveness of the proposed algorithm with statistical significance in comparison to the state-of-the-art incremental learning methods. Wing W. Y. Ng, Jianjun Zhang 0004, Chun Sing Lai, Witold Pedrycz, Loi Lei Lai, Xizhao Wang |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | New Appliance Detection for Nonintrusive Load MonitoringabstractCurrent methods for nonintrusive load monitoring (NILM) problems assume that the number of appliances in the target location is known, however, this may not be realistic. In real-world situations, the initial setup of the site can be known but new appliances may be added by users after a period of time, especially in a household or nonrestrictive scenarios. In this sense, current methods without detecting new appliances may not accurately monitor loads of different appliances and scenarios. In this paper, a novel new appliance detection method is proposed for NILM with imbalance classification for appliances switching ON or OFF. The prediction of appliances being switched ON or OFF is an important step in load monitoring and the switching on frequencies for coffee machine and air conditioning in a household are different, making the problem inherently imbalanced. Experimental results show that the proposed method yields outstanding performance against the well-known oversampling method, synthetic minority oversampling technique, on real NILM applications in scenarios with new appliances emerging. Jianjun Zhang 0004, Xuanqun Chen, Wing W. Y. Ng, Chun Sing Lai, Loi Lei Lai |
IEEE Trans. Ind. Informatics | 3 |
| 2018 | Incremental Hashing with Dynamic Semantic PoolabstractMost of the existing hashing methods for image retrieval are based on the assumption the image database is stationary. However, in the real world data environments are always changing or non-stationary, therefore the underlying data distribution may change from time to time which will result in the problem of concept drift. Incremental Hashing (ICH) is the only existing method to handle image retrieval with concept drift in non-stationary data environments. It builds hash codes for the database through increments. At each increment, a set of new hash functions is built with the new chunk of data, which is utilized to update the multi-hashing system to generate multiple sets of hash codes for all data. However, only the newest data chunk is used to train individual hash functions, while the semantic similarity information of previous data is missed. In this paper, we present a new hashing method based on ICH for image retrieval with concept drift, Incremental Hashing with Dynamic Semantic Pool (ICH-DSP). It builds a semantic pool to collect representative labeled data for each existing class. The semantic pool is updated incrementally and is used as the supervisory information for the training of hash functions. Experimental results on three real world image databases show that ICH-DSP outperforms the original ICH and other state-of-the-art hashing methods. Xing Tian, Wing W. Y. Ng, Hui Wang 0001 |
SMC | 2 |
| 2018 | Stochastic Sensitivity Measure-Based Noise Filtering and Oversampling Method for Imbalanced Classification ProblemsabstractClass imbalance problems occur in many real-world applications. Oversampling methods are effective to handle class imbalance issues by replicating or generating new minority samples to rebalance the class distribution. However, current methods directly using all minority samples will also use noisy samples to generate new samples which may lead to more severe class overlapping and introduce more noisy samples. In this work, we propose a stochastic sensitivity measure-based noise filtering and oversampling method, i.e. the SSMNFOS, to improve the robustness of oversampling method with respect to noisy samples. Samples yielding high stochastic sensitivities are identified as noises by a neural network ensemble and will not participate in the oversampling method for rebalancing the class distribution. Comprehensive experimental studies are carried out on ten datasets with five different noise levels to analyze the effectiveness of the proposed method. Experimental results show that the SSMNFOS outperforms state-of-the-art methods with 95% statistical significance. Jianjun Zhang 0004, Wing W. Y. Ng |
SMC | 2 |
| 2018 | Bagging-boosting-based semi-supervised multi-hashing with query-adaptive re-ranking
Wing W. Y. Ng, Xiancheng Zhou, Xing Tian, Xizhao Wang, Daniel S. Yeung |
Neurocomputing | 1 |
| 2017 | Bsmboost for imbalanced pattern classification problemsabstractNumbers of samples in different classes are in nature imbalanced in many machine learning problems. Single classifier-based methods are subject to high variance. Therefore, ensemble-based methods are more suitable for dealing with imbalanced pattern classification problems. In this work, we propose a boosting-based method: BSMBoost which creates an ensemble of classifiers using samples selected by both the stochastic sensitivity measure (SSM) and the AdaBoost algorithm to yield higher and more robust performances. Experimental results show that the BSMBoost yields better and more robust performances in comparison to other state-of-the-art boosting-based imbalanced classification methods. Wing W. Y. Ng, Yuda Zhang, Jianjun Zhang 0004 |
SMC | 1 |
| 2017 | Incremental Hashing for Semantic Image Retrieval in Nonstationary EnvironmentsabstractA very large volume of images is uploaded to the Internet daily. However, current hashing methods for image retrieval are designed for static databases only. They fail to consider the fact that the distribution of images can change when new images are added to the database over time. The changes in the distribution of images include both discovery of a new class and a distribution of images within a class owing to concept drift. Retraining of hash tables using all images in the database requires a large computation effort. This is also biased to old data owing to the huge volume of old images which leads to a poor retrieval performance over time. In this paper, we propose the incremental hashing (ICH) method to deal with the two aforementioned types of changes in the data distribution. The ICH uses a multihashing to retain knowledge coming from images arriving over time and a weight-based ranking to make the retrieval results adaptive to the new data environment. Experimental results show that the proposed method is effective in dealing with changes in the database. Wing W. Y. Ng, Xing Tian, Yueming Lv, Daniel S. Yeung, Witold Pedrycz |
IEEE Trans. Cybern. | 1 |
| 2017 | Color Distribution Pattern Metric for Person ReidentificationabstractAccompanying the growth of surveillance infrastructures, surveillance IP cameras mount up rapidly, crowding Internet of Things (IoT) with countless surveillance frames and increasing the need of person reidentification (Re-ID) in video searching for surveillance and forensic fields. In real scenarios, performance of current proposed Re-ID methods suffers from pose and viewpoint variations due to feature extraction containing background pixels and fixed feature selection strategy for pose and viewpoint variations. To deal with pose and viewpoint variations, we propose the color distribution pattern metric ( CDPM ) method, employing color distribution pattern ( CDP ) for feature representation and SVM for classification. Different from other methods, CDP does not extract features over a certain number of dense blocks and is free from varied pedestrian image resolutions and resizing distortion. Moreover, it provides more precise features with less background influences under different body types, severe pose variations, and viewpoint variations. Experimental results show that our CDPM method achieves state-of-the-art performance on both 3DPeS dataset and ImageLab Pedestrian Recognition dataset with 68.8% and 79.8% rank 1 accuracy, respectively, under the single-shot experimental setting. Yingsheng Ye, Wing W. Y. Ng |
Wirel. Commun. Mob. Comput. | 3 |
| 2016 | Dual autoencoders features for imbalance classification problem
Wing W. Y. Ng, Guangjun Zeng, Jianjun Zhang 0004, Daniel S. Yeung, Witold Pedrycz |
Pattern Recognit. | 1 |
| 2016 | MLPNN Training via a Multiobjective Optimization of Training Error and Stochastic SensitivityabstractThe training of a multilayer perceptron neural network (MLPNN) concerns the selection of its architecture and the connection weights via the minimization of both the training error and a penalty term. Different penalty terms have been proposed to control the smoothness of the MLPNN for better generalization capability. However, controlling its smoothness using, for instance, the norm of weights or the Vapnik-Chervonenkis dimension cannot distinguish individual MLPNNs with the same number of free parameters or the same norm. In this paper, to enhance generalization capabilities, we propose a stochastic sensitivity measure (ST-SM) to realize a new penalty term for MLPNN training. The ST-SM determines the expectation of the squared output differences between the training samples and the unseen samples located within their Q -neighborhoods for a given MLPNN. It provides a direct measurement of the MLPNNs output fluctuations, i.e., smoothness. We adopt a two-phase Pareto-based multiobjective training algorithm for minimizing both the training error and the ST-SM as biobjective functions. Experiments on 20 UCI data sets show that the MLPNNs trained by the proposed algorithm yield better accuracies on testing data than several recent and classical MLPNN training methods. Daniel S. Yeung, Jin-Cheng Li, Wing W. Y. Ng, Patrick P. K. Chan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Spam filtering for short messages in adversarial environment
Patrick P. K. Chan, Daniel S. Yeung, Wing W. Y. Ng |
Neurocomputing | 4 |
| 2015 | Two-phase mapping hashing
Wing W. Y. Ng, Yueming Lv, Daniel S. Yeung, Patrick P. K. Chan |
Neurocomputing | 1 |
| 2015 | Diversified Sensitivity-Based Undersampling for Imbalance Classification ProblemsabstractUndersampling is a widely adopted method to deal with imbalance pattern classification problems. Current methods mainly depend on either random resampling on the majority class or resampling at the decision boundary. Random-based undersampling fails to take into consideration informative samples in the data while resampling at the decision boundary is sensitive to class overlapping. Both techniques ignore the distribution information of the training dataset. In this paper, we propose a diversified sensitivity-based undersampling method. Samples of the majority class are clustered to capture the distribution information and enhance the diversity of the resampling. A stochastic sensitivity measure is applied to select samples from both clusters of the majority class and the minority class. By iteratively clustering and sampling, a balanced set of samples yielding high classifier sensitivity is selected. The proposed method yields a good generalization capability for 14 UCI datasets. Wing W. Y. Ng, Daniel S. Yeung, Shaohua Yin, Fabio Roli |
IEEE Trans. Cybern. | 1 |
| 2015 | Asymmetric Cyclical Hashing for Large Scale Image RetrievalabstractThis paper addresses a problem in the hashing technique for large scale image retrieval: learn a compact hash code to reduce the storage cost with performance comparable to that of the long hash code. A longer hash code yields a better precision rate of retrieved images. However, it also requires a larger storage, which limits the number of stored images. Current hashing methods employ the same code length for both queries and stored images. We propose a new hashing scheme using two hash codes with different lengths for queries and stored images, i.e., the asymmetric cyclical hashing. A compact hash code is used to reduce the storage requirement, while a long hash code is used for the query image. The image retrieval is performed by computing the Hamming distance of the long hash code of the query and the cyclically concatenated compact hash code of the stored image to yield a high precision and recall rate. Experiments on benchmarking databases consisting up to one million images show the effectiveness of the proposed method. Yueming Lv, Wing W. Y. Ng, Ziqian Zeng, Daniel S. Yeung, Patrick P. K. Chan |
IEEE Trans. Multim. | 2 |
| 2014 | An improved differential evolution and its application to determining feature weights in similarity-based clustering
Chunru Dong, Wing W. Y. Ng, Xizhao Wang, Patrick P. K. Chan, Daniel S. Yeung |
Neurocomputing | 2 |
| 2014 | LG-Trader: Stock trading decision support based on feature selection by weighted localized generalization error model
Wing W. Y. Ng, Xue-Ling Liang, Jin-Cheng Li, Daniel S. Yeung, Patrick P. K. Chan |
Neurocomputing | 1 |
| 2014 | Steganalysis classifier training via minimizing sensitivity for different imaging sources
Wing W. Y. Ng, Zhi-Min He, Daniel S. Yeung, Patrick P. K. Chan |
Inf. Sci. | 1 |
| 2012 | Dynamic fusion method using Localized Generalization Error Model
Patrick P. K. Chan, Daniel S. Yeung, Wing W. Y. Ng, Chih-Min Lin, James Nga-Kwok Liu |
Inf. Sci. | 3 |
| 2010 | Steganography detection using localized generalization error modelabstractSteganography detection is a technique to tell whether there are secret messages hidden in images. The performance of a steganalysis system is mainly determined by the method of feature extraction and the architecture selection of the classifier. Selecting a proper classifier with proper parameters will improve the detection accuracy and generalization capability of the system. We propose a Radial Basis Function Neural Network (RBFNN) optimized by the Localized Generalization Error Model (L-GEM) for steganograhpy detection. In the proposed method, the discrete cosine transform (DCT) features and the Markov features are used as inputs of neural networks for detection. To enhance the generalization capability of the RBFNN and the performance of detecting steganography in future images, the architecture of the RBFNN is selected by minimizing the L-GEM. The experimental results show that the proposed method provides a better performance on testing images in comparison with the existing method in attackting Steghide, OutGuess and F5. Zhi-Min He, Wing W. Y. Ng, Patrick P. K. Chan, Daniel S. Yeung |
SMC | 2 |
| 2009 | IPCM Separability Ratio for Supervised Feature SelectionabstractCollecting data is very easy now owing to fast computers and ease of Internet access. It raises the problem of the curse of dimensionality to supervised classification problems. In our previous work, an Intra-Prototype / Inter-Class Separability Ratio (IPICSR) model is proposed to select relevant features for semi-supervised classification problems. In this work, a new margin based feature selection model is proposed based on the IPICSR model for supervised classification problems. Owing to the nature of supervised classification problems, a more accurate class separating margin could be found by the classifier. We adopt this advantage in the new Intra-Prototype / Class Margin Separability Ratio (IPCMSR) model. Experimental results are promising when compared to several existing methods using 4 UCI datasets. Wing W. Y. Ng, Jun Wang 0017, Daniel S. Yeung |
SMC | 1 |
| 2009 | Radial Basis Function network learning using localized generalization error bound
Daniel S. Yeung, Patrick P. K. Chan, Wing W. Y. Ng |
Inf. Sci. | 3 |
| 2008 | Semantic Chunk Annotation for questions using Maximum Entropyabstractwe present a ME (Maximum Entropy) model for Semantic Chunk Annotation in a Chinese Question and Answer (Q&A) system. The model was derived from a corpus of real world questions, which are collected from some discussion groups on the Internet. The questions are supposed to be answered by other people, so the questions are very complex. The semantic chunks were introduced. Feature for the model was described and MI (Mutual Information) was adopted for feature selection. The training data consists of 14000 sentences and the test data consists of 4000 sentences. The result: F-score is 90.68%. Shixi Fan, Yaoyun Zhang, Wing W. Y. Ng, Xuan Wang 0002, Xiaolong Wang 0001 |
SMC | 3 |
| 2008 | Quantitative study on candlestick pattern for Shenzhen Stock MarketabstractShenzhen Stock Market grows rapidly yet is still a young market when compared with Hong Kong, New York and London markets. Its daily turnover reaches billions US dollars. A good prediction of stock price will bring us substantial pecuniary reward. Technical analysis is a widely adopted financial prediction tool in worldwide stock markets. Candlestick pattern is one of the most efficient methods in technical analysis. However, does candlestick pattern prediction works for Shenzhen stocks? Candlestick patterns are always defined by fuzzy terms, could we have a quantitative definition of these patterns? We perform a quantitative study on these two major research problems in this paper. We study the morning star pattern in this work and the method in this paper could be extended to other patterns easily. So, we propose adopting the Radial Basis Function Neural Networks trained with Localized Generalization Error Model to predict whether or not the stock price will increase after the appearance of it. Then, we extract the patterns from the neural network to provide a quantitative definition of the morning star pattern for a particular stock. Experimental results show that our modification to the morning star pattern prediction prevents up to 69% of false prediction of the morning star pattern. We also provide a quantitative measure of the morning star patterns for two of the Shenzhen stocks. Huili Li, Wing W. Y. Ng, John W. T. Lee, Binbin Sun, Daniel S. Yeung |
SMC | 2 |
| 2008 | Information extraction based on information fusion from multiple news sources from the webabstractThe traditional information extraction tools have been developed for years. But, the accuracy of extraction is not very satisfactory, especially for named entity extraction. In this work, we analyze the reasons of it and propose a novel method to improve the accuracy. Existing methods extract information from a text which is collected from a single source. This is very difficult to extract the exact information we need. From the Internet, one could easily find tens of sources for the same information (e.g. particular news). In this work, we propose to combine information extracted from multiple sources using a majority voting to find the information we needed. We use Change of CEO as an example and we extract the new CEO, original CEO and the company name for the event. A off-the-shelf named entity extraction tool is adopted and our major contribution is the fusion of extraction results. Without our work, one finds single news provides many people and company names, such that we do not know who the new CEO is. By using our method, we provide the 2 CEO names and 1 company name. Experimental results show that our method has a high accuracy in finding the exact information. Wing W. Y. Ng, John W. T. Lee, Binbin Sun, Daniel S. Yeung |
SMC | 2 |
| 2008 | Localized generalization error based active learning for image annotationabstractContent-based image auto-annotation becomes a hot research topic owing to the development of image retrieval system and the storing technology of multimedia information. It is a key step in most of those image processing applications. In this work, we adopt active learning to image annotation for reducing the number of labeled images required for supervised learning procedure. Localized Generalization Error Model (L-GEM) based active learning uses localized generalization error bound as the sample selection criterion. In each turn, the most informative sample from a set of unlabeled samples is selected by the L-GEM based active learning will be labeled and added to the training dataset. A heuristic and a Q value selection improvement methods are introduced in this paper. The experimental results show that the proposed active learning efficiently reduces the number of labeled training samples. Moreover, the improvement method improve the performances in both testing accuracy and training time which are both essential in image annotation applications. Binbin Sun, Wing W. Y. Ng, Daniel S. Yeung, Jun Wang 0017 |
SMC | 2 |
| 2008 | A Novel Feature Grouping Method for Ensemble Neural Network Using Localized Generalization Error ModelabstractMultiple Classifier System (MCS) is a very popular research topic in recent years. It has been proved theoretically and empirically to be better than single classifiers in many scenarios. Creating diverse sets of classifier is one of the key issues in building MCSs. Feature grouping is one of the methods to create diverse classifiers and it has been shown to improve the accuracy of an MCS. In this paper, we propose a new feature grouping method based on Genetic Algorithm (GA) with the localized Generalization Error Model as the evaluation criterion. The combined individual classifiers using the weighted sum are examined in this paper. Moreover, several feature grouping methods are compared with the proposed method in this work. The experimental results on benchmark dataset show that the MCS trained by the proposed method is promising. Aki P. F. Chan, Patrick P. K. Chan, Wing W. Y. Ng, Eric C. C. Tsang, Daniel S. Yeung |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2008 | The Localized Generalization Error Model for Single Layer Perceptron Neural Network and Sigmoid Support Vector MachineabstractWe had developed the localized generalization error model for supervised learning with minimization of Mean Square Error. In this work, we extend the error model to Single Layer Perceptron Neural Network (SLPNN) and Support Vector Machine (SVM) with sigmoid kernel function. For a trained SLPNN or SVM and a given training dataset, the proposed error model bounds above the error for unseen samples which are similar to the training samples. As the major component of the localized generalization error model, the stochastic sensitivity measure formula for perceptron neural network derived in this work has relaxed the assumptions of same distribution for all inputs and each sample perturbed only once in previous works. These make the sensitivity measure applicable to pattern classification problems. The stochastic sensitivity measure of SVM with Sigmoid kernel is also derived in this work as a component of the localized generalization error model. At the end of this paper, we discuss the advantages of the proposed error bound over existing error bound. Wing W. Y. Ng, Daniel S. Yeung, Eric C. C. Tsang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2008 | Feature selection using localized generalization error for supervised classification problems using RBFNN
Wing W. Y. Ng, Daniel S. Yeung, Michael Firth, Eric C. C. Tsang, Xizhao Wang |
Pattern Recognit. | 1 |
| 2007 | Heuristic improvement for active learning using localized generalization error as selection criterionabstractOwing to the growth of Internet and computer technology, pattern recognition for large-scale datasets has become one of the hot research topics. The major challenges are to reduce the human efforts involved and to improve the efficiency. Traditional passive learning methods require labeling of all training samples may not be feasible in large-scale recognition problems because of the requirement of large-scale class labeling for the huge number of training samples. In the literatures, there are many studies on active learning methods, which does not require all training samples to be labeled and it selects training samples for labeling based on the knowledge of the current classifier. In this paper, we present an active learning method using localized generalization error of candidate sample as selection criterion. Our method uses the generalization error of candidate sample, so theoretically it should have a better performance than other methods. From the experiment results, our method outperforms other methods in both yielding higher prediction accuracy on testing dataset and selecting fewer training samples. Furthermore, we propose a heuristics improvement based on the Q-neighborhood idea of the localized generalization error model to reduce the number of samples being selected and the computational time. Wing W. Y. Ng, Binbin Sun, Daniel S. Yeung, Xizhao Wang |
SMC | 1 |
| 2007 | Structured large margin machines: sensitive to data distributions
Daniel S. Yeung, Defeng Wang, Wing W. Y. Ng, Eric C. C. Tsang, Xizhao Wang |
Mach. Learn. | 3 |
| 2007 | Image classification with the use of radial basis function neural networks and the minimization of the localized generalization error
Wing W. Y. Ng, Andrés Dorado, Daniel S. Yeung, Witold Pedrycz, Ebroul Izquierdo |
Pattern Recognit. | 1 |
| 2007 | Localized Generalization Error of Gaussian-based Classifiers and Visualization of Decision Boundaries
Wing W. Y. Ng, Daniel S. Yeung, Defeng Wang, Eric C. C. Tsang, Xizhao Wang |
Soft Comput. | 1 |
| 2007 | Localized Generalization Error Model and Its Application to Architecture Selection for Radial Basis Function Neural NetworkabstractThe generalization error bounds found by current error models using the number of effective parameters of a classifier and the number of training samples are usually very loose. These bounds are intended for the entire input space. However, support vector machine (SVM), radial basis function neural network (RBFNN), and multilayer perceptron neural network (MLPNN) are local learning machines for solving problems and treat unseen samples near the training samples to be more important. In this paper, we propose a localized generalization error model which bounds from above the generalization error within a neighborhood of the training samples using stochastic sensitivity measure. It is then used to develop an architecture selection technique for a classifier with maximal coverage of unseen samples by specifying a generalization error threshold. Experiments using 17 University of California at Irvine (UCI) data sets show that, in comparison with cross validation (CV), sequential learning, and two other ad hoc methods, our technique consistently yields the best testing classification accuracy with fewer hidden neurons and less training time. Daniel S. Yeung, Wing W. Y. Ng, Defeng Wang, Eric C. C. Tsang, Xizhao Wang |
IEEE Trans. Neural Networks | 2 |
| 2006 | Bankruptcy Prediction Using Multiple Classifier System with Mutual Information Feature GroupingabstractThe prediction of bankruptcy helps an organization to choose its business partners and banks to approve or reject loan requests. So, it is essential to predict the bankruptcy of an organization. In this work, a multiple classifier system which combines decision from several different base classifiers trained using different samples with different input features is proposed for the bankruptcy prediction. The input features for each base classifier are selected using its mutual information with respect to the output. Experimental results of the proposed method using a real bankruptcy dataset from Compustat Global Dataset are promising. Aki P. F. Chan, Wing W. Y. Ng, Daniel S. Yeung, Eric C. C. Tsang, Michael Firth |
SMC | 2 |
| 2006 | Applying undistorted neural network sensitivity analysis in iris plant classification and construction productivity prediction
Daniel S. Yeung, Wing W. Y. Ng |
Soft Comput. | 3 |
| 2005 | Multiple Classifier System with Feature Grouping for Intrusion Detection: Mutual Information Approach
Aki P. F. Chan, Wing W. Y. Ng, Daniel S. Yeung, Eric C. C. Tsang |
KES (3) | 2 |
| 2005 | Active learning using localized generalization error of candidate sample as criterionabstractIn classification problem, the learning process can be more efficient if the informative samples can be selected actively based on the knowledge of the classifier. This problem is called active learning. Most of the existing active learning methods did not directly relate to the generalization error of classifiers. Also, some of them need high computational time or are based on strict assumptions. This paper describes a new active learning strategy using the concept of localized generalization error of the candidate samples. The sample which yields the largest generalization error will be chosen for query. This method can be applied to different kinds of classifiers and its complexity is low. Experimental results demonstrate that the prediction accuracy of the classifier can be improved by using this selecting method and fewer training samples are possible for the same prediction accuracy. Patrick P. K. Chan, Wing W. Y. Ng, Daniel S. Yeung |
SMC | 2 |
| 2005 | Quantitative study on the generalization error of multiple classifier systemsabstractMultiple classifier system (MCS) has been one of the hot research topics in machine learning field. A MCS merges an ensemble of different or same type of classifiers together to enhance the problem solving performance of machine learning. However, the choice of the number of classifiers and the fusion method are usually based on ad-hoc selection. In this paper, we propose a novel quantitative measure of the generalization error for MCS. The localized generalization error model bounds above the mean square error (MSE) of a MCS for unseen samples located within a neighborhood of the training samples. The relationship between the proposed model and classification accuracy is also discussed in this paper. This model quantitatively measures the goodness of the MCS in approximating the unknown input-output mapping hidden in the training dataset The localized generalization error model is applied to select a MCS, among different choices of number of classifiers and fusion methods, for a given classification problem. Experimental results on three real world datasets are performed to show promising results. Wing W. Y. Ng, Aki P. F. Chan, Daniel S. Yeung, Eric C. C. Tsang |
SMC | 1 |
| 2003 | Input sample selection for RBF neural network classification problems using sensitivity measureabstractLarge data sets containing irrelevant or redundant input samples reduce the performance of learning and increases storage and labeling costs. This work compares several sample selection and active learning techniques and proposes a novel sample selection method based on the stochastic radial basis function neural network sensitivity measure (SM). The experimental results for the UCI IRIS data set show that we can remove 99% of data while keeping 95% of classification accuracy when applying both sensitivity based feature and sample selection methods. We propose a single and consistent method, which is robust enough to handle both feature and sample selection for a supervised RBFNN classification system, by using the same neural network architecture for both selection and classification tasks. Wing W. Y. Ng, Daniel S. Yeung, Ian Cloete |
SMC | 1 |