Chi-Man Vong

dblp:68/5768 · also Chi Man Vong, Matthew Chi-Man Vong · DBLP profile ↗
← Back
118ranked-venue papers
9as first author
71since 2021 · last 2026
0000-0001-7997-8279ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 84 · 9 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Online Conformal Selection with Accept-to-Reject Changes
abstract
Selecting a subset of promising candidates from a large pool is crucial across various scientific and real-world applications. Conformal selection offers a distribution-free and model-agnostic framework for candidate selection with uncertainty quantification. While effective in offline settings, its application to online scenarios, where data arrives sequentially, poses challenges. Notably, conformal selection permits the deselection of previously selected candidates, which is incompatible with applications requiring irreversible selection decisions. This limitation is particularly evident in resource-intensive sequential processes, such as drug discovery, where advancing a compound to subsequent stages renders reversal impractical. To address this issue, we extend conformal selection to an online Accept-to-Reject Changes (ARC) procedure: non-selected data points can be reconsidered for selection later, and once a candidate is selected, the decision is irreversible. Specifically, we propose a novel conformal selection method, Online Conformal Selection with Accept-to-Reject Changes (dubbed OCS-ARC), which incorporates online Benjamini–Hochberg procedure into the candidate selection process. We provide theoretical guarantees that OCS-ARC controls the false discovery rate (FDR) at or below the nominal level at any timestep under both i.i.d. and exchangeable data assumptions. Additionally, we theoretically show that our approach naturally extends to multivariate response settings. Extensive experiments on synthetic and real-world datasets demonstrate that OCS-ARC significantly improves selection power over the baseline while maintaining valid FDR control across all examined timesteps.
Kangdao Liu, Huajun Xi, Chi-Man Vong, Hongxin Wei
AAAI3
2026 Anchoring the Affective Manifold: Learning Canonical and Disentangled Representations via Generative Cross-Modal Alignment
abstract
Dominant multimodal emotion recognition paradigms often neglect the intrinsic geometric structure of affect, resulting in representations heavily entangled with non-affective factors.To address this, we propose a Canonical Disentangled Multimodal Generative Framework aimed at recovering the canonical affective manifold from raw data.We explicitly decompose the latent space into a canonical Shared Affective Subspace (z vad ) and a Private Modality Subspace (z priv ).We facilitate this factorization through Supervised Manifold Anchoring and Cross-Modal Manifold Alignment.Experiments demonstrate that our model effectively disentangles affect from private attributes (e.g., identity), achieving superior robustness in zero-shot cross-domain transfer compared to fully supervised baselines, while enabling controllable emotion generation.
Jintao Cheng, Chi-Man Vong
ACL (1)4
2026 A novel interpretable fault diagnosis approach for chillers based on self-adaptive oversampling technique and tree-based multi-meta stacking ensemble method
Guoyu Yao, Zhengdong Wang, Chi-Man Vong, Zhenbao Liu
Eng. Appl. Artif. Intell.5
2026 Interactive multiple instance learning network for whole slide image analysis
Qi Lai, Chi-Man Vong, Tao Yan 0006, Xiaokun Liang
Expert Syst. Appl.2
2026 ICL: In-loop continual learning framework for language model pre-training for E-commerce
abstract
Pre-trained language models have become a critical natural language processing component in many E-commerce applications. As businesses continue to evolve, the pre-trained models should be able to adopt new domain knowledge and new tasks. This paper proposes a novel sequential multi-task pre-trained language framework, ICL-BERT (In-loop Continual Learning BERT), which enables evolving the current model with new knowledge and new tasks. The contributions of ICL-BERT are (1) vocabularies and entities are optimized on E-commerce corpus; (2) a new glyph embedding is introduced to learn glyph information for vocabularies and entities; (3) specific and general tasks are designed to encode E-commerce knowledge for pre-training ICL-BERT; and (4) a new task-gating mechanism, called ICL (In-loop continual Learning), is proposed for sequential multi-task learning, which evolves the current model effectively and efficiently. Our evaluation results demonstrate that ICL-BERT outperforms existing models in both CLUE and e-commerce tasks, with an average accuracy improvement of 1.73% and 3.5%, respectively. Furthermore, ICL-BERT serves as a fundamental pre-trained language model that runs online in JingDong’s daily business.
Chiman Wong, Sanpeng Wang, Danyang Zhu, Chi-Man Vong
Intell. Data Anal.6
2026 Fast doubly reconstructed affinity propagation for semi-supervised classification
Guanjin Wang, Chi-Man Vong, Shitong Wang 0001
Neurocomputing3
2026 DDCL-Net: dual-domain collaborative learning network for EEG-based emotion recognition
Zhenrong Ruan, Chi-Man Vong, Huqin Weng, Chuangquan Chen
Knowl. Based Syst.3
2026 Exploiting the potential supervision information of clean samples in partial label learning
Guangtai Wang, Chi-Man Vong
Pattern Recognit.2
2026 Guiding multimodal LLMs for efficient visual place recognition
Zhijian He, Jintao Cheng, Yipu Zhang 0002, Chi-Man Vong, Jin Wu 0002, Xieyuanli Chen
Pattern Recognit. Lett.5
2026 CGSI: Context-Guided and UAV's Status Informed Multimodal Framework for Generalizable Cross-View Geo-Localization
abstract
Cross-View Geo-Localization is essential for drone visual localization and navigation, which aims at establishing correlation between images collected by unmanned aerial vehicle (UAV) and satellite platforms in the same geographic area. Drastic changes in the drone’s viewpoints pose a significant challenge for methods based on image representation mining. Previous studies attempt to learn fine-grained image appearance features from various perspectives; however, they tend to underutilize the various state information of the UAV. This paper proposes a novel multimodal framework, CGSI (Context-Guided and UAV’s Status Informed), which leverages UAV state textual descriptions to mitigate scene bias caused by viewpoint differences. The following two issues are addressed to achieve more accurate and reliable multimodal geo-localization: 1) The domain gap across different datasets caused by the fixed UAV altitudes. We propose a Context-Guided Multimodal Tokenizer, which learns contextual vectors from multi-altitude visual features and utilizes them as adaptive text tokens. 2) Multimodal features are susceptible to state-feature ambiguity. We propose a Drone Group Graph Attention method to enhance the association between UAV visual feature with the same location ID but different states and exploit the intrinsic relationships to extract discriminative multimodal features. Extensive experiments on the University-1652 and SUES benchmark demonstrate that our CGSI significantly outperforms existing algorithms, achieving state-of-the-art performance. The substantial improvements observed in cross-region ablation experiments further showcase the superior domain generalization capability of our method.
Jian Sun 0038, Junlang Huang, Yimin Zhou 0001, Chi-Man Vong
IEEE Trans. Circuits Syst. Video Technol.5
2026 Nested Stacking TSK Fuzzy Classifier by Additive-Rule-Decomposition and Antinoisy-Labeling Based Learning on Large-Scale Noisy Labeling Data
abstract
This study explores how to develop a novel deep stacking Takagi–Sugeno–Kang (TSK) fuzzy classifier to efficiently realize interpretable classification for large-scale multi-class complex data contaminated with noisy labels. To this end, the nested stacking TSK fuzzy classifier NSARD-TSK and its additive-rule-decomposition & anti-noisy-labeling based learning method are proposed to embody their completely distinctive design methodology: (1) each quasi-higher-order TSK fuzzy subclassifier of NSARD-TSK is built in deep stacking way for several interpretable zero-order TSK fuzzy subclassifiers to roughly behave like a higher-order TSK classifier, such that both strong uncertainty-handling capability and generalization are provided. (2) the parameters in all the fuzzy rules of each quasi-higher-order TSK fuzzy subclassifier except the last one are transformatively trained to have their desired and undesired parts such that the resultant fuzzy rules with desired parameters for clean data are acquired in early training and simultaneously difficult-to-classify data containing noisy labeling data are fixed. (3) each successive quasi-higher-order TSK fuzzy subclassifier is stacked on difficult-to-classify data from the previous subclassifier so as to form NSARD-TSK's nested stacking structure with enhanced generalization capability. In particular, the last quasi-higher-order TSK fuzzy subclassifier adopts the proposed anti-noisy-labeling squared loss function as its unique learning objective with a theoretical guarantee for suppressing noisy labels. (4) NSARD-TSK linearly aggregates all the subclassifiers to further improve its final classification performance without affecting NSARD-TSK's interpretability. Experiments on 12 benchmarking datasets with and without noisy labels validate NSARD-TSK's efficiency over the comparative methods in terms of average testing performance and model complexity.
Zhengxun Guo, Chi-Man Vong, Shitong Wang 0001
IEEE Trans. Fuzzy Syst.3
2026 A Fuzzy Large TabNet-Based Model and Its Distillation Learning for Noisy-Labeled Data via Consequent Additive Decomposition
abstract
In this study, in order to exploit an interpretable fuzzy large model on clean dataset from large and/or complex noisy-labeled data and then generate its student model appropriately through knowledge distillation for its lightweight usability, anInterpretablefuzzylargeTabNet-basedmodel, called IFLTNM, based on the well-proven deep neural network TabNet and interpretable fuzzy rules is first proposed as a type of fuzzy large models. By means of both the well-established memorization effect and the proposed consequent additive decomposition, after interpretable antecedents are fixed, the consequent parameters of all fuzzy rules in the IFLTNM model are appropriately determined at early training stage for almost clean training subset from the whole noisy-labeled training set. After that, by regularizing the learning objective of the proposed knowledge distillation on almost clean training subset with an extra loss function on almost noisy-labeled training subset obtained after the IFLTNM's training, the proposed learning objective is optimized to have IFLTNM's student model s-IFLTNM consisting of the same number of fuzzy rules sharing the same interpretable antecedents as in IFLTNM yet having linear consequents. IFLTNM has its structural and training novelty in the sense of leveraging both the memorization effect and consequent additive decomposition to train a fuzzy large model and acquire almost noisy-labeled training samples, and s-IFLTNM is distilled from IFLTNM creatively through leveraging both almost clean and noisy-labeled training samples instead of only clean samples. Experimental results on real-world benchmark datasets demonstrate the effectiveness of IFLTNM over the state-of-the-art methods,e.g., achieving 84.94% average testing accuracy yet preserving interpretability for thecovertypedataset even with 30% symmetric label noise, and the power of s- IFLTNM in the sense of both a significant reduction of computational burden and model complexity of less than 400 on the adopted datasets with only a marginal drop in performance.
Chi-Man Vong, Shitong Wang 0001
IEEE Trans. Fuzzy Syst.2
2026 Correlative Fusion-Based Multi-Instance Partial Multi-Label Learning
abstract
Multi-instance partial multi-label learning (MIPML) addresses a challenging scenario wherein each training sample comprises a bag of multiple instances associated with a candidate label set comprising several true labels alongside noisy labels simultaneously. Current MIPML methods typically neglect the essential correlations between labels and instances at both the instance and bag levels, which limits their effectiveness in disambiguation and predictive accuracy. To address these limitations, we present Correlation-Fusion MIPML (CF-MIPML), an innovative framework that integrates Label Confidence Generation (LCG) and Candidate Label Disambiguation (CLD). The LCG module systematically constructs a robust label confidence matrix by capturing correlation structures within and across bags, thereby providing a foundation for precise label disambiguation. The CLD module utilizes the comprehensive confidence matrix to further improve label predictions, utilizing an optimized iterative fusion loss function that incorporates partial loss and interaction loss. This joint-loss strategy enables ongoing refinement of label confidence during training, thereby improving the robustness and accuracy of predictions. Comprehensive experimental results on various benchmark and real-world datasets demonstrate that CF-MIPML outperforms existing state-of-the-art methods, enhancing handling of complex label ambiguity and improving overall model generalization in practical MIPML scenarios.
Chi-Man Vong, Yiu-Ming Cheung
IEEE Trans. Multim.2
2026 Federated Class Incremental Learning Method With High Accuracy and Extremely Low Communication Cost Based on Broad Learning System
abstract
Federated Class Incremental Learning (FCIL) enables distributed clients to collaboratively train a global model based on their private sequential tasks without compromising data privacy. Currently, some FCIL methods have been proposed, and most are designed based on deep models. However, enabling these FCIL models to converge requires numerous communication rounds, significantly increasing communication costs. Recently, the Broad Learning System (BLS), an effective and efficient shallow model, was proposed and adapted for CIL tasks [i.e., BLS-Class Incremental Learning (CIL)]. BLS-CIL exhibits fast updates and high retainability. However, it requires prior knowledge of when new class data arrives and cannot be directly used in federated scenarios due to the global catastrophic forgetting in FCIL. Thus, an innovative Federated Class incremental learning method based on BLS (FedCBLS) is proposed, which extends BLS-CIL within the federated scenario and provides three advantages: 1) high accuracy from the local perspective, achieved by integrating BLS-CIL with a newly designed automatic decision-making (ADM) method to detect novel classes and learn them incrementally for local clients; 2) high accuracy from the global perspective, attained through the newly proposed local model refinement (LMR) and global model projection (GMP) methods, mitigating global catastrophic forgetting stemming from heterogeneous data across clients; and 3) extremely low communication costs due to the newly derived closed-form solutions without iterative optimization for both local and global models. Comprehensive experimental results show that our FedCBLS outperforms the state-of-the-art (SOTA) FCIL methods by up to 8.15%, while drastically reducing communication costs to 1% of SOTA’s. Our code is available athttps://github.com/dujie-szu/FedCBLS.git
Jie Du 0001, Wenbing Chen, Peng Liu 0070, Chi-Man Vong, Tianfu Wang 0001, C. L. Philip Chen
IEEE Trans. Syst. Man Cybern. Syst.4
2026 SemanticStitch: enhancing image coherence through foreground-aware seam carving
Ji-Ping Jin, Chen-Bin Feng, Chi-Man Vong
Vis. Comput.4
2025 GBRIP: Granular Ball Representation for Imbalanced Partial Label Learning
abstract
Partial label learning (PLL) is a complicated weakly supervised multi-classification task compounded by class imbalance. Currently, existing methods only rely on inter-class pseudo-labeling from inter-class features, often overlooking the significant impact of the intra-class imbalanced features combined with the inter-class. To address these limitations, we introduce Granular Ball Representation for Imbalanced PLL (GBRIP), a novel framework for imbalanced PLL. GBRIP utilizes coarse-grained granular ball representation and multi-center loss to construct a granular ball-based feature space through unsupervised learning, effectively capturing the feature distribution within each class. GBRIP mitigates the impact of confusing features by systematically refining label disambiguation and estimating imbalance distributions. The novel multi-center loss function enhances learning by emphasizing the relationships between samples and their respective centers within the granular balls. Extensive experiments on standard benchmarks demonstrate that GBRIP outperforms existing state-of-the-art methods, offering a robust solution to the challenges of imbalanced PLL.
Yiu-Ming Cheung, Chi-Man Vong, Wenbin Qian
AAAI3
2025 C-Adapter: Adapting Deep Classifiers for Efficient Conformal Prediction Sets
abstract
Conformal prediction, as an emerging uncertainty quantification technique, typically functions as post-hoc processing for the outputs of trained classifiers. To optimize the classifier for maximum predictive efficiency, Conformal Training rectifies the training objective of base classifiers with a regularization that minimizes the average prediction set size at a specific error rate. However, the regularization term inevitably deteriorates the classification accuracy of classifiers, thereby leading to suboptimal efficiency of conformal predictors. To address this issue, we introduce Conformal Adapter (C-Adapter), an adapter-based tuning method to enhance the efficiency of conformal predictors without sacrificing accuracy. In particular, we implement the adapter as a class of intra order-preserving functions and tune it with our proposed loss that maximizes the discriminability of non-conformity scores between correctly and randomly matched data-label pairs. Using C-Adapter, the model tends to produce higher non-conformity scores for incorrect labels than for correct ones, thereby enhancing predictive efficiency across different coverage rates. Extensive experiments demonstrate that C-Adapter can effectively adapt various classifiers for efficient conformal prediction sets, as well as enhance the conformal training method.
Kangdao Liu, Hao Zeng 0005, Jianguo Huang, Huiping Zhuang, Chi-Man Vong, Hongxin Wei
ECAI5
2025 Towards Fully Test-Time Adaptation via Variance Balancing and Semantic Augmentation
abstract
Fully test-time adaptation (FTTA) is to adapt a model trained on a source domain to a target domain during the testing phase. Traditional methods like entropy minimization primarily focus on reducing uncertainty in output predictions, yet often overlook the diversity in target prediction results, which is critical for unbalanced classes in complex datasets. To address this, our study introduces a new method named Variance Balancing and Semantic Augmentation (VBSA). VBSA begins by maximizing the sum of singular values in predictions, coupled with a novel variance penalization strategy. This strategy not only focuses the model on unbalanced classes but also mitigates the overfitting risk associated with singular value maximization, thereby ensuring a balanced emphasis across various classes and enhancing the diversity of prediction results. Furthermore, VBSA incorporates semantic data augmentation using data from previous batches, offering semantic-level augmentation for all classes, with particular benefits for unbalanced ones. Extensive experiments demonstrate that our VBSA method has produced the most advanced performance.
Houcheng Su, Bingli Wang, Daixian Liu, Chen-Bin Feng, Chi-Man Vong
ICASSP6
2025 FACNet: Feature Alignment Fast Point Cloud Completion Network
abstract
Point cloud completion aims to infer complete point clouds based on partial 3D point cloud inputs. Various previous methods apply coarse-to-fine strategy networks for generating complete point clouds. However, such methods are not only relatively time-consuming but also cannot provide representative complete shape features based on partial inputs. In this paper, a novel feature alignment fast point cloud completion network (FACNet) is proposed to directly and efficiently generate the detailed shapes of objects. FACNet aligns high-dimensional feature distributions of both partial and complete point clouds to maintain global information about the complete shape. During its decoding process, the local features from the partial point cloud are incorporated along with the maintained global information to ensure complete and time-saving generation of the complete point cloud. Experimental results show that FACNet outperforms the state-of-the-art on PCN, Completion3D, and MVP datasets, and achieves competitive performance on ShapeNet-55 and KITTI datasets. Moreover, FACNet and a simplified version, FACNet-slight, achieve a significant speedup of 3–10 times over other state-of-the-art methods.
Xinxing Yu, Chi-Chong Wong, Chi-Man Vong, Yanyan Liang 0001
Comput. Vis. Media4
2025 Hybrid multiple instance learning network for weakly supervised medical image classification and localization
Qi Lai, Chi-Man Vong, Tao Yan 0006, Pak-Kin Wong 0001, Xiaokun Liang
Expert Syst. Appl.2
2025 Joint Prediction of SOH and RUL for Lithium Batteries Considering Capacity Self Recovery and Model Drift
abstract
Accurate monitoring of the state of health (SOH) and remaining useful life (RUL) of lithium batteries is a key technology to realize their stable and economic operation. The current data-driven methods have problems such as local prediction errors caused by capacity self-recovery effect, model drift phenomenon, and inaccurate trend prediction caused by complex degradation process. In this article, we propose a joint SOH and RUL forecasting based on the fusion of kolmogorov-Arnold network (KAN), deep bidirectional long short-term memory network (DBLSTM) and adaptive mechanism method (KAN-HDBLSTM-AM). The method innovatively introduces an adaptive mechanism that considers the capacity degradation trend for different prediction steps, suppresses the accumulation of errors in the model prediction process, and further improves the prediction accuracy by combining the nonlinear approximation capability of KAN and the mastery of data temporal characteristics of DBLSTM. Experimental validation shows that the proposed method can effectively overcome the prediction errors caused by capacity self-recovery and model drift, and can retain sufficient accuracy under long prediction steps, with the prediction trend being more in line with the actual change trend and the prediction error being smaller compared with other methods.
Zhifei Li 0010, Zhengdong Wang, Zhenbao Liu, Chi-Man Vong
IEEE Internet Things J.5
2025 Complexity-optimized sparse Bayesian learning for scalable classification tasks
Jiahua Luo, Junyi Xiang, Chiman Wong, Chi-Man Vong
Inf. Sci.5
2025 Analytical selection of hidden parameters through expanded enhancement matrix stability for functional-link neural networks and broad learning systems
Chi-Man Vong, C. L. Philip Chen, Shitong Wang 0001
Knowl. Based Syst.2
2025 Dealing with partial labels by knowledge distillation
Guangtai Wang, Yiqiang Lai, Chi-Man Vong
Pattern Recognit.4
2025 HSA-Former: Hierarchical Spatial Aggregation Transformer for EEG-Based Emotion Recognition
abstract
Affective brain–computer interfaces (aBCIs) have shown promising applications due to the significant advancements in utilizing electroencephalogram (EEG) signals for emotion recognition. By measuring neuronal activity across various brain regions, EEG provides rich spatial information that is essential for discriminative feature extraction and effective emotion recognition. However, existing deep learning models face challenges in effectively capturing and leveraging local and global spatial dependencies. To address this, we propose a hierarchical spatial aggregation Transformer (HSA-Former) for EEG-based emotion recognition, which explores spatial relationships from multiple levels, including electrodes, intrabrain regions, and interbrain regions. The HSA-Former is characterized by: 1) a multihierarchical spatial information learning architecture that sequentially extracts diverse spatial features from electrodes, through intrabrain regions, and across interbrain regions; 2) a parallel learning approach that captures internal spatial features of different brain regions from different channels; and 3) an effective aggregation method that mitigates the problem of information loss typically caused by direct pooling in Transformers. Extensive experiments on three public datasets (i.e., SEED, SEED-IV, and SEED-V) demonstrate that the proposed HSA-Former outperforms existing state-of-the-art methods. Interestingly, the weights learned by HSA-Former emphasize the frontal, temporal, and occipital lobes as critical brain regions, aligning with established mechanisms underlying emotion recognition.
Jiayang Huang, Chi-Man Vong, Chen Li 0058, Leicai Xu, C. L. Philip Chen, Chuangquan Chen
IEEE Trans. Comput. Soc. Syst.2
2025 Spatial-Aware Conformal Prediction for Trustworthy Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification involves assigning unique labels to each pixel to identify various land cover categories. While deep classifiers have achieved high predictive accuracy in this field, they lack the ability to rigorously quantify confidence in their predictions. This limitation restricts their application in critical contexts where the cost of prediction errors is significant, as quantifying the uncertainty of model predictions is crucial for the safe deployment of predictive models. To address this limitation, a rigorous theoretical proof is presented first, which demonstrates the validity of Conformal Prediction, an emerging uncertainty quantification technique, in the context of HSI classification. Building on this foundation, a conformal procedure is designed to equip any pre-trained HSI classifier with trustworthy prediction sets, ensuring that the true labels are included with a user-defined probability (e.g., 95%). Furthermore, a novel framework of Conformal Prediction specifically designed for HSI data, called Spatial-Aware Conformal Prediction (SACP), is proposed. This framework integrates essential spatial information of HSI by aggregating the non-conformity scores of pixels with high spatial correlation, effectively improving the statistical efficiency of prediction sets. Both theoretical and empirical results validate the effectiveness of the proposed approaches. The source code is available at https://github.com/J4ckLiu/SACP.
Kangdao Liu, Tianhao Sun, Hao Zeng 0005, Yongshan Zhang, Chi-Man Pun, Chi-Man Vong
IEEE Trans. Circuits Syst. Video Technol.6
2025 Fuzzy Overlapping Modularity Clustering for Symmetric-Tensor Based Graph
abstract
While existing graph clustering methods can only be oriented to classical graph data, this study, as the first attempt, focuses on the proposed symmetric-tensor based graph and its clustering algorithm. To this end, the concept of fuzzy overlapping modularity is defined and then is extended into its generalized version for a symmetric-tensor based graph. Subsequently, based on the principle of maximizing the symmetric-tensor based fuzzy overlapping modularity, a novel learning objective is derived for fuzzy clustering, and the corresponding clustering algorithm, fuzzy overlapping modularity clustering (FOMC), is also proposed. In addition, with only one additional hyperparameter, the semisupervised clustering algorithm SFOMC is also derived for a symmetric-tensor based graph with some labeled samples. Extensive experimental results on synthetic and real benchmarking datasets verify the clustering power of both FOMC and SFOMC on symmetric-tensor based graphs. In particular, FOMC achieves 13.94% improvement over the average performance of the comparative methods on the adopted real networks, and SFOMC’s clustering performance increment becomes 1.799 times higher than the average value of the comparative methods when samples have been labeled from 5% to 25% in the adopted graphs.
Chi-Man Vong, Shitong Wang 0001
IEEE Trans. Fuzzy Syst.2
2025 Robust Visual Place Recognition Under Variational Views
abstract
Visual place recognition (VPR) has played an essential role in simultaneous localization and mapping-based mobile robotics and autonomous driving in the past decade, which can identify previously visited places by matching the current observed view against a view database for global localization and loop closure. However, existing VPR methods always suffer fromfalse place recognitiondue to the following issues: insensitive spatial–temporal embedding extraction, lack of multiview descriptor for matching, and misconsideration of false views for model optimization. To address these issues, a novel robust framework calledVPR under variational views (VPR-VV)is proposed. VPR-VV is integrated with: a sequence encoder to extract robust spatial–temporal features from a view sequence, then a hierarchical view retrieval module is employed for multiview feature descriptor aggregation, and a novel enhanced ranking feedback average precision loss with the normalized discounted cumulative gain metric is designed for model optimization. As a result, VPR-VV can significantly enhance the accuracy and robustness of VPR for robot localization under variational views. Experiments with ablation studies are conducted on various challenging indoor and outdoor datasets, and our framework’s superiority is demonstrated: VPR-VV outperforms state-of-the-art (SOTA) methods by up to 9.4% in recall@1, and real-time inference is achieved without additional memory or computational overhead.
Junlang Huang, Jie Du 0001, Chuangquan Chen, Xieyuanli Chen, Yimin Zhou 0001, Chi-Man Vong
IEEE Trans. Ind. Informatics8
2025 Context-CAM: Context-Level Weight-Based CAM With Sequential Denoising to Generate High-Quality Class Activation Maps
abstract
Class activation mapping (CAM) methods have garnered considerable research attention because they can be used to interpret the decision-making of deep convolutional neural network (CNN) models and provide initial masks for weakly supervised semantic segmentation (WSSS) tasks. However, the class activation maps generated by most CAM methods usually have two limitations: 1) a lack of the ability to cover the whole object when using low-level features; and 2) introducing background noise. To mitigate these issues, an innovative Context-level weights-based CAM (Context-CAM) method is proposed, which guarantees: 1) the non-discriminative regions that have similar appearances and are located close to the discriminative regions can also be highlighted by the newly designed Region-Enhanced Mapping (REM) module using context-level weights; and 2) the background noises are gradually eliminated via a newly proposed Semantic-guided Reverse Sequence Fusion (SRSF) strategy that can sequentially denoise and fuse the region-enhanced maps from the last layer to the first layer. Extensive experimental results show that our Context-CAM can generate higher-quality class activation maps than classic and state-of-the-art (SOTA) CAM methods in terms of the Energy-Based Pointing Game (EBPG) score, and the improvements are up to 35.49% when compared to the second-best method. Moreover, for WSSS tasks, our Context-CAM can directly replace the CAM method used in existing WSSS methods without any architectural modification to further improve the segmentation performance. Our code is available at https://github.com/cwb0611/Context-CAM.
Jie Du 0001, Wenbing Chen, Chi-Man Vong, Peng Liu 0070, Tianfu Wang 0001
IEEE Trans. Image Process.3
2025 OE-BevSeg: An Object Informed and Environment Aware Multimodal Framework for Bird's-Eye-View Vehicle Semantic Segmentation
abstract
Bird’s-eye-view (BEV) semantic segmentation is becoming crucial in autonomous driving systems. It realizes ego-vehicle surrounding environment perception by projecting 2D multi-view images into 3D world space. Recently, BEV segmentation has made notable progress, attributed to better view transformation modules, larger image encoders, or more temporal information. However, there are still two issues: 1) a lack of effective understanding and enhancement of BEV space features, particularly in accurately capturing long-distance environmental features and 2) recognizing fine details of target objects. To address these issues, we propose OE-BevSeg, an end-to-end multimodal framework that enhances BEV segmentation performance through global environment-aware perception and local target object enhancement. OE-BevSeg employs an environment-aware BEV compressor. Based on prior knowledge about the main composition of the BEV surrounding environment varying with the increase of distance intervals, long-sequence global modeling is utilized to improve the model’s understanding and perception of the environment. From the perspective of enriching target object information in segmentation results, we introduce the center-informed object enhancement module, using centerness information to supervise and guide the segmentation head, thereby enhancing segmentation performance from a local enhancement perspective. Additionally, we designed a multimodal fusion branch that integrates multi-view RGB image features with radar/LiDAR features, achieving significant performance improvements. Extensive experiments show that, whether in camera-only or multimodal fusion BEV segmentation tasks, our approach achieves state-of-the-art results by a large margin on the nuScenes dataset for vehicle segmentation, demonstrating superior applicability in the field of autonomous driving. Our code will be released at https://github.com/SunJ1025/OE-BevSeghttps://github.com/SunJ1025/OE-BevSeg.
Jian Sun 0038, Yuqi Dai, Chi-Man Vong, Qing Xu 0010, Shengbo Eben Li, Jianqiang Wang 0003, Keqiang Li 0002
IEEE Trans. Intell. Transp. Syst.3
2025 Broad Multitask Learning System With Group Sparse Regularization
abstract
The broad learning system (BLS) featuring lightweight, incremental extension, and strong generalization capabilities has been successful in its applications. Despite these advantages, BLS struggles in multitask learning (MTL) scenarios with its limited ability to simultaneously unravel multiple complex tasks where existing BLS models cannot adequately capture and leverage essential information across tasks, decreasing their effectiveness and efficacy in MTL scenarios. To address these limitations, we proposed an innovative MTL framework explicitly designed for BLS, named group sparse regularization for broad multitask learning system using related task-wise (BMtLS-RG). This framework combines a task-related BLS learning mechanism with a group sparse optimization strategy, significantly boosting BLS's ability to generalize in MTL environments. The task-related learning component harnesses task correlations to enable shared learning and optimize parameters efficiently. Meanwhile, the group sparse optimization approach helps minimize the effects of irrelevant or noisy data, thus enhancing the robustness and stability of BLS in navigating complex learning scenarios. To address the varied requirements of MTL challenges, we presented two additional variants of BMtLS-RG: BMtLS-RG with sharing parameters of feature mapped nodes (BMtLS-RGf), which integrates a shared feature mapping layer, and BMtLS-RGf and enhanced nodes (BMtLS-RGfe), which further includes an enhanced node layer atop the shared feature mapping structure. These adaptations provide customized solutions tailored to the diverse landscape of MTL problems. We compared BMtLS-RG with state-of-the-art (SOTA) MTL and BLS algorithms through comprehensive experimental evaluation across multiple practical MTL and UCI datasets. BMtLS-RG outperformed SOTA methods in 97.81% of classification tasks and achieved optimal performance in 96.00% of regression tasks, demonstrating its superior accuracy and robustness. Furthermore, BMtLS-RG exhibited satisfactory training efficiency, outperforming existing MTL algorithms by 8.04-42.85 times.
Chuangquan Chen, Chi-Man Vong, Yiu-Ming Cheung
IEEE Trans. Neural Networks Learn. Syst.3
2025 Stacked Ensemble of Extremely Interpretable Takagi-Sugeno-Kang Fuzzy Classifiers for High-Dimensional Data
abstract
To overcome the inappropriateness of the recently-developed fully interpretable Takagi–Sugeno–Kang fuzzy systems (FIMG-TSK) for high-dimensional classification tasks, which is caused by their unreliable Gaussian mixture models and their very lengthy fuzzy rules on all the original features, this study attempts to develop a stacked ensemble of extremely interpretable first-order TSK fuzzy classifiers (SEXI-TSK-FC) comprising extremely interpretable FIMG-TSK-based classifiers. SEXI-TSK-FC has structural and algorithmic novelties. In the structural sense, to guarantee enhanced generalizability and short fuzzy rules, the proposed XI-TSK is created as each subclassifier on a subset of the original features. Then it stacks each successive subclassifier on both the outputs and the important features selected, which are fromthe incorrectly classifieddataset by the previous subclassifier. After that, SEXI-TSK-FC linearly aggregates all the outputs of its subclassifiers with a one-step calculation to enhance classification accuracy while preserving extreme interpretability. In the algorithmic sense, each short fuzzy rule of the XI-TSK subclassifier is determined using the proposed fuzzy feature selection and clustering algorithm to select the subset of all the original features and simultaneously fix the antecedent and consequent of each rule. After that, the rule weights in each subclassifier are trained quickly with strong generalizability using the proposed Vapnik–Chervonenkis dimension minimization–based learning. Experimental results on 12 benchmark datasets demonstrate the power of the proposed classifier SEXI-TSK-FC on high-dimensional data in testing accuracy, training time, and extreme interpretability.
Erhao Zhou, Chi-Man Vong, Shitong Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2025 Learning few-shot semantic segmentation with error-filtered segment anything model
Chen-Bin Feng, Qi Lai, Kangdao Liu, Houcheng Su, Kaixi Luo, Chi-Man Vong
Vis. Comput.7
2024 Multi-kernel partial label learning using graph contrast disambiguation
Zhonglin Wan, Chi-Man Vong
Appl. Intell.3
2024 Label correlations-based multi-label feature selection with label enhancement
Wenbin Qian, Yinsong Xiong, Weiping Ding 0001, Chi-Man Vong
Eng. Appl. Artif. Intell.5
2024 Billion-scale pre-trained knowledge graph model for conversational chatbot
Chiman Wong, Wen Zhang 0015, Huajun Chen, Chi-Man Vong, Chuangquan Chen
Neurocomputing5
2024 Stream label distribution learning processing via broad learning system
Guangtai Wang, Chi-Man Vong
Inf. Sci.3
2024 Federated learning using model projection for multi-center disease diagnosis with non-IID data
Jie Du 0001, Peng Liu 0070, Chi-Man Vong, Yongke You, Bai Ying Lei, Tianfu Wang 0001
Neural Networks4
2024 The MorPhEMe Machine: An Addressable Neural Memory for Learning Knowledge-Regularized Deep Contextualized Chinese Embedding
abstract
Deep contextualized embeddings, as learned by large pre-training models, have proven highly effective in various downstream natural language processing tasks. However, the embedding space in these large models lacks explicit regularization, leading to underfitting and substantial costs during large-scale training on huge corpora. In this paper, we present a novel approach to learning deep contextualized embeddings, introducing linguistic knowledge regularization. Specifically, our proposed model, MorPhEMe (Morphology and Phonology Embedding Memory), features an external addressable memory with two additional addressable memories for storing morphology and phonology knowledge. MorPhEMe can be seamlessly stacked into a deep architecture. Notably different from existing pre-training models, MorPhEMe boasts two distinctive features: (1) compositional encoding and decompositional decoding facilitated by a dynamic addressing mechanism; and (2) explicit memory embedding regularization through cross-layer memory sharing. Theoretical analysis suggests that the inclusion of morphology and phonology enables MorPhEMe to reduce the modeling complexity of natural language sequences. We evaluate MorPhEMe across a diverse set of Chinese natural language processing tasks, including language modeling, word similarity computation, word analogy reasoning, relation extraction, and machine reading comprehension. Experimental results demonstrate that MorPhEMe, in contrast to state-of-the-art models, achieves remarkable improvements with fewer parameters and rapid convergence.
Zhibin Quan, Chi-Man Vong, Wankou Yang
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Multilayer Stacked Evolving Fuzzy System Combined With Compressed Representation Learning
abstract
In order to process the high-dimensionally complicated problems, the intelligence systems need to go deeper to learn high-level data representation. In this article, based on the stacked generalization principle, a multilayered stacked learning system is proposed. Like the deep networks, the proposed system is organized in a layer-by-layer way with evolving fuzzy systems (EFSs) as its base-building units. Each EFS is designed as an autoencoder (AE) to learn the simpler data representation, then multiple EFS-based AEs are stacked in a feedforward manner for learning more complex data representation. Since there exists the redundant or irrelevant information in the new data representation, which may limit high generalization, a novel feature compressing layer is followed by each EFS-based AE to refine the new data representation and reduce the feature dimension via very sparse random projection (VSRP). The proposed system is featured in the following merits: 1) more expressive and complex data representation can be learned in a stacked multilayer architecture; 2) manually tuning the structure of each AE is avoided in every layer since the EFS can self-adapt both its structure and parameters online; 3) the resulting system can avoid overfitting and obtain higher generalization by removing redundant or irrelevant information via VSRP; 4) the steady-state error analysis of the proposed system is studied based on the separable approximation property, which guarantees the learning convergence of the resulting system. Our experimental results on various benchmark datasets indicate the efficacy of the proposed system.
Hai-Jun Rong, Zhao-Xu Yang, Chi-Man Vong
IEEE Trans. Fuzzy Syst.4
2024 Doubly Interpretable Fuzzy Apriori Classifier by Successive Stacking and One-Step Wide Calculation
abstract
Except for linguistic interpretability and uncertainty-handling ability, fuzzy Apriori method (FAM) is being hurdled by both very expensive computational burdens and low generalization capability caused by serious correlation between short to long fuzzy rules generated. The novel doubly interpretable classifier (DI-FAM) with FAM-based hybrid structure is proposed to circumvent the above shortcomings of FAM. DI-FAM successively stacks the short rule bundles of each FAM subclassifier on both a sampled feature subset and the outputs of the previous stacking layer. DI-FAM then finds out the output weights of short rule bundles at each stacking layer, followed by a linear subclassifier (as a compensator) on all the original input features with one-step wide calculation. DI-FAM has four distinct merits: 1)low computational complexitystemmed from both its fast generation way of short rule bundles by FAM subclassifiers, respectively, on their own features, and its one-step calculation for the output weights; 2)theoretical guaranteeabout no violation of the importance ranking orders of the short rules by each FAM subclassifier on its own features at each stacking layer with regard to all the fuzzy rules by FAM on all the input features; 3)enhanced generalization capabilityby successively stacking short rule bundles at each stacking layer according to the stacked generalization principle; and 4)double interpretabilitythat DI-FAM shares both linguistic interpretability of all FAM subclassifiers and feature-importance-based interpretability of a linear subclassifier. Extensive experimental results indicate the effectiveness of DI-FAM in the sense of classification performance, training speed, incremental learning, and double interpretability.
Runshan Xie, Chi-Man Vong, Shitong Wang 0001
IEEE Trans. Fuzzy Syst.2
2024 Internally and Generatively Decorrelated Ensemble of First-Order Takagi-Sugeno-Kang Fuzzy Regressors With Quintuply Diversity Guarantee
abstract
While the recently developed first-order Takagi–Sugeno–Kang (TSK) fuzzy regressor FIMG-TSK shares its full interpretability, this study leverages the concisely expressed output variance of FIMG-TSK to explore its high feasibility in being a wide-ensemble component. In this way, the regression performance can be enhanced and simultaneously FIMG-TSKs overdependence on the rule weights can be alleviated to a certain extent. To this end, a wide ensemble of all base regressors (i.e., FIMG-TSKs) called EFIMG-TSKs is proposed. In the ensemble-strategic aspect, EFIMG-TSK has its internally and generatively decorrelated ensemble strategy with a quintuply diversity guarantee for its strong generalization capability. In the learning aspect, the learning objective of EFIMG-TSKs reflects the internally and generatively decorrelated ensemble learning of all base FIMG-TSKs and accordingly is optimized globally with an analytical solution to the weights of all fuzzy rules in each base FIMG-TSK. The experimental results on 16 benchmarking datasets demonstrate the effectiveness of EFIMG-TSKs in terms of regression performance, training time, and interpretability.
Erhao Zhou, Chi-Man Vong, Yusuke Nojima, Shitong Wang 0001
IEEE Trans. Fuzzy Syst.2
2024 DDIO-Mapping: A Fast and Robust Visual-Inertial Odometry for Low-Texture Environment Challenge
abstract
Accurate localization and pose estimation remain challenging for autonomous robots in low-texture environment. This article proposes a tightly coupled direct depth-inertial odometry and mapping (DDIO-Mapping) framework to simultaneously tackle three crucial issues in such environments: 1) ineffective feature point extraction; 2) inefficient searching of feature points; and 3) imbalanced feature extraction under uneven illumination conditions. In DDIO-Mapping, a novel robust strategy is designed that combines grayscale and depth features for optimization instead of only the RBG features in the existing methods. To improve searching efficiency, a new RGBD feature extraction is applied to directly extract both the depth and grayscale features from the RGBD images, which only requires searching the feature points in the 2-D space rather than the enormous 3-D space in K-dimensional (KD) tree. To deal with imbalanced feature extraction, a feature filtering and selection strategy is proposed to adaptively adjust the depth and grayscale weightage. Finally, with the effectively extracted features from RGBD images, a new nonlinear tightly coupled inverse depth residual function is customized to accurately estimate the optimal pose in low-texture environments. The framework is highly robust, accurate, and efficient. Experiments demonstrate that DDIO-Mapping reduces the root-mean-square error by approximately 30% compared to other state-of-the-art algorithms while retaining the same efficiency of approximately 20–35 ms.
Chuangquan Chen, Yongquan Chen, Junlang Huang, Zuguang Zhou, Yimin Zhou 0001, Chi-Man Vong
IEEE Trans. Ind. Informatics8
2024 Weakly Supervised Semantic Segmentation via Dual-Stream Contrastive Learning of Cross-Image Contextual Information
abstract
Weakly supervised semantic segmentation (WSSS) aims at learning a semantic segmentation model with only image-level tags. Despite intensive research on deep learning approaches over a decade, there is still a significant performance gap between WSSS and full semantic segmentation. Most current WSSS methods always focus on a limited single image (pixel-wise) information while ignoring the valuable interimage (semantic-wise) information. From this perspective, a novel end-to-end WSSS framework called DSCNet is developed along with two innovations: i) pixel-wise group contrast and semantic-wise graph contrast are proposed and introduced into the WSSS framework; ii) a novel dual-stream contrastive learning mechanism is designed to jointly handle pixel-wise and semantic-wise context information for better WSSS performance. Specifically, the pixel-wise group contrast learning and semantic-wise graph contrast learning tasks form a more comprehensive solution. Extensive experiments on PASCAL VOC and MS COCO benchmarks verify the superiority of DSCNet over SOTA approaches and baseline models.
Qi Lai, Chi-Man Vong, Chuangquan Chen
IEEE Trans. Ind. Informatics2
2024 Class-Incremental Learning Method With Fast Update and High Retainability Based on Broad Learning System
abstract
Machine learning aims to generate a predictive model from a training dataset of a fixed number of known classes. However, many real-world applications (such as health monitoring and elderly care) are data streams in which new data arrive continually in a short time. Such new data may even belong to previously unknown classes. Hence, class-incremental learning (CIL) is necessary, which incrementally and rapidly updates an existing model with the data of new classes while retaining the existing knowledge of old classes. However, most current CIL methods are designed based on deep models that require a computationally expensive training and update process. In addition, deep learning based CIL (DCIL) methods typically employ stochastic gradient descent (SGD) as an optimizer that forgets the old knowledge to a certain extent. In this article, a broad learning system-based CIL (BLS-CIL) method with fast update and high retainability of old class knowledge is proposed. Traditional BLS is a fast and effective shallow neural network, but it does not work well on CIL tasks. However, our proposed BLS-CIL can overcome these issues and provide the following: 1) high accuracy due to our novel class-correlation loss function that considers the correlations between old and new classes; 2) significantly short training/update time due to the newly derived closed-form solution for our class-correlation loss without iterative optimization; and 3) high retainability of old class knowledge due to our newly derived recursive update rule for CIL (RULL) that does not replay the exemplars of all old classes, as contrasted to the exemplars-replaying methods with the SGD optimizer. The proposed BLS-CIL has been evaluated over 12 real-world datasets, including seven tabular/numerical datasets and six image datasets, and the compared methods include one shallow network and seven classical or state-of-the-art DCIL methods. Experimental results show that our BIL-CIL can significantly improve the classification performance over a shallow network by a large margin (8.80%-48.42%). It also achieves comparable or even higher accuracy than DCIL methods, but greatly reduces the training time from hours to minutes and the update time from minutes to seconds.
Jie Du 0001, Peng Liu 0070, Chi-Man Vong, Chuangquan Chen, Tianfu Wang 0001, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.3
2024 An Adaptive Deep Metric Learning Loss Function for Class-Imbalance Learning via Intraclass Diversity and Interclass Distillation
abstract
Deep metric learning (DML) has been widely applied in various tasks (e.g., medical diagnosis and face recognition) due to the effective extraction of discriminant features via reducing data overlapping. However, in practice, these tasks also easily suffer from two class-imbalance learning (CIL) problems: data scarcity and data density, causing misclassification. Existing DML losses rarely consider these two issues, while CIL losses cannot reduce data overlapping and data density. In fact, it is a great challenge for a loss function to mitigate the impact of these three issues simultaneously, which is the objective of our proposed intraclass diversity and interclass distillation (IDID) loss with adaptive weight in this article. IDID-loss generates diverse features within classes regardless of the class sample size (to alleviate the issues of data scarcity and data density) and simultaneously preserves the semantic correlations between classes using learnable similarity when pushing different classes away from each other (to reduce overlapping). In summary, our IDID-loss provides three advantages: 1) it can simultaneously mitigate all the three issues while DML and CIL losses cannot; 2) it generates more diverse and discriminant feature representations with higher generalization ability, compared with DML losses; and 3) it provides a larger improvement on the classes of data scarcity and density with a smaller sacrifice on easy class accuracy, compared with CIL losses. Experimental results on seven public real-world datasets show that our IDID-loss achieves the best performances in terms of G-mean, F1-score, and accuracy when compared with both state-of-the-art (SOTA) DML and CIL losses. In addition, it gets rid of the time-consuming fine-tuning process over the hyperparameters of loss function.
Jie Du 0001, Xiaoci Zhang, Peng Liu 0070, Chi-Man Vong, Tianfu Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Fast Broad Multiview Multi-Instance Multilabel Learning (FBM3L) With Viewwise Intercorrelation
abstract
Multiview multi-instance multilabel learning (M3L) is a popular research topic during the past few years in modeling complex real-world objects such as medical images and subtitled video. However, existing M3L methods suffer from relatively low accuracy and training efficiency for large datasets due to several issues: 1) the viewwise intercorrelation (i.e., the correlations of instances and/or bags between different views) are neglected; 2) the diverse correlations (e.g., viewwise intercorrelation, interinstance correlation, and interlabel correlation) are not jointly considered; and 3) high computation burden for training process over bags, instances, and labels across different views. To resolve these issues, a novel framework called fast broad M3L (FBM3L) is proposed with three innovations: 1) utilization of viewwise intercorrelation for better modeling of M3L tasks while existing M3L methods have not considered; 2) based on graph convolutional network (GCN) and broad learning system (BLS), a viewwise subnetwork is newly designed to achieve joint learning among the diverse correlations; and 3) under BLS platform, FBM3L can learn multiple subnetworks jointly across all views with significantly less training time. Experiments show that FBM3L is highly competitive (or even better than) in all evaluation metrics [up to 64% in average precision (AP)] and much faster than most M3L (or MIML) methods (up to 1030 times), especially on large multiview datasets (≥260 K objects).
Qi Lai, Chi-Man Vong, Jianhang Zhou, Yimin Zhou 0001, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.2
2023 DC-DC Buck circuit fault diagnosis with insufficient state data based on deep model and transfer strategy
Zhenbao Liu, Chi-Man Vong, Yongyi Cai
Expert Syst. Appl.3
2023 Multi-graph embedding for partial label learning
Chi-Man Vong, Zhonglin Wan
Neural Comput. Appl.2
2023 A novel sequential structure for lightweight multi-scale feature learning under limited available images
Peng Liu 0070, Jie Du 0001, Chi-Man Vong
Neural Networks3
2023 Fast AUC Maximization Learning Machine With Simultaneous Outlier Detection
abstract
While AUC maximizing support vector machine (AUCSVM) has been developed to solve imbalanced classification tasks, its huge computational burden will make AUCSVM become impracticable and even computationally forbidden for medium or large-scale imbalanced data. In addition, minority class sometimes means extremely important information for users or is corrupted by noises and/or outliers in practical application scenarios such as medical diagnosis, which actually inspires us to generalize the AUC concept to reflect such importance or upper bound of noises or outliers. In order to address these issues, by means of both the generalized AUC metric and the core vector machine (CVM) technique, a fast AUC maximizing learning machine, called ρ -AUCCVM, with simultaneous outlier detection is proposed in this study. ρ -AUCCVM has its notorious merits: 1) it indeed shares the CVM's advantage, that is, asymptotically linear time complexity with respect to the total number of sample pairs, together with space complexity independent on the total number of sample pairs and 2) it can automatically determine the importance of the minority class (assuming no noise) or the upper bound of noises or outliers. Extensive experimental results about benchmarking imbalanced datasets verify the above advantages of ρ -AUCCVM.
Chi-Man Vong, Shitong Wang 0001
IEEE Trans. Cybern.2
2023 Joint Label Enhancement and Label Distribution Learning via Stacked Graph Regularization-Based Polynomial Fuzzy Broad Learning System
abstract
Label distribution learning (LDL), leveraging the label significance (LS), is more appropriate for solving label ambiguity problems than multilabel learning (MLL). However, directly obtaining the LS of LDL is extremely expensive and challenging. Thus, label enhancement (LE) algorithms are effectively proposed to acquire inherent LS from MLL for training the LDL models. Nevertheless, most existing LE models will suffer from low accuracy and low efficiency with following issues: ignoring mapping relationship between feature and label space, resulting in inaccurate enhanced data; designing independently apart from LDL models, resulting in an inability of unified LE-LDL learning; and requiring to optimize numerous parameters iteratively, resulting in worse training efficiency. Consequently, a novel unified LE-LDL learning framework, namely stacked graph-regularized polynomial-based fuzzy broad learning system (SGP-FBLS), is proposed by following three innovations: polynomial-based fuzzy system is introduced to enhance feature mapping ability while improving the learning performance effectively; graph regularized-based optimization objective function (GP-FBLS) is presented by considering interinstance correlation and label correlation to mine potential LS, thereby improving the accuracy of subsequent LDL tasks; and a weight stacked strategy is innovatively proposed to directly transmit LS and weighted parameters from GP-FBLS to SGP-FBLS without retraining, achieving the most satisfactory performance while significantly improving the training efficiency. Finally, comparative studies on 19 practical datasets demonstrate the effectiveness and superiority of proposed methods.
Chi-Man Vong, Guangtai Wang, Wenbin Qian, Yimin Zhou 0001, C. L. Philip Chen
IEEE Trans. Fuzzy Syst.2
2023 A Fully Interpretable First-Order TSK Fuzzy System and Its Training With Negative Entropic and Rule-Stability-Based Regularization
abstract
While interpretable antecedent parts of first-order Takagi–Sugeno–Kang (TSK) fuzzy rules can be properly acquired by adopting some clustering methods, this study aims at avoiding the commonly used yet fully incomprehensive consequent parts and their intractable training, and simultaneously seeking for enhanced generalization performance by determining the weight of each rule. The central idea is to build a mathematically equivalent bridge between a Gaussian mixture model (GMM) and a fully interpretable first-order TSK fuzzy system called FIMG-TSK, with the help of Gaussian-mixture's mean. The resultant FIMG-TSK has a simple expected output expression without summation-to-one defuzzification, which will be helpful in inducing both smaller output variance and a negative entropic and rule-stability-based regularizer for enhancing the generalization performance. After revealing three factors affecting the output stability of FIMG-TSK, the negative entropic and rule-stability-based regularizer is designed through both these factors and the squared entropy to make the output variance of FIMG-TSK as small as possible. Accordingly, a novel training method, whose objective function takes the proposed regularizer as an additional term and hence compromises both accuracy and output stability of FIMG-TSK, is developed to quickly provide an analytical solution to the weight of each rule. The effectiveness of the proposed training method is manifested by the experimental results on ten regression datasets.
Erhao Zhou, Chi-Man Vong, Yusuke Nojima, Shitong Wang 0001
IEEE Trans. Fuzzy Syst.2
2023 Parameter-Free Loss for Class-Imbalanced Deep Learning in Image Classification
abstract
Current state-of-the-art class-imbalanced loss functions for deep models require exhaustive tuning on hyperparameters for high model performance, resulting in low training efficiency and impracticality for nonexpert users. To tackle this issue, a parameter-free loss (PF-loss) function is proposed, which works for both binary and multiclass-imbalanced deep learning for image classification tasks. PF-loss provides three advantages: 1) training time is significantly reduced due to NO tuning on hyperparameter(s); 2) it dynamically pays more attention on minority classes (rather than outliers compared to the existing loss functions) with NO hyperparameters in the loss function; and 3) higher accuracy can be achieved since it adapts to the changes of data distribution in each mini-batch instead of the fixed hyperparameters in the existing methods during training, especially when the data are highly skewed. Experimental results on some classical image datasets with different imbalance ratios (IR, up to 200) show that PF-loss reduces the training time down to 1/148 of that spent by compared state-of-the-art losses and simultaneously achieves comparable or even higher accuracy in terms of both G-mean and area under receiver operating characteristic (ROC) curve (AUC) metrics, especially when the data are highly skewed.
Jie Du 0001, Yanhong Zhou, Peng Liu 0070, Chi-Man Vong, Tianfu Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Accurate and Efficient Large-Scale Multi-Label Learning With Reduced Feature Broad Learning System Using Label Correlation
abstract
Multi-label learning for large-scale data is a grand challenge because of a large number of labels with a complex data structure. Hence, the existing large-scale multi-label methods either have unsatisfactory classification performance or are extremely time-consuming for training utilizing a massive amount of data. A broad learning system (BLS), a flat network with the advantages of succinct structures, is appropriate for addressing large-scale tasks. However, existing BLS models are not directly applicable for large-scale multi-label learning due to the large and complex label space. In this work, a novel multi-label classifier based on BLS (called BLS-MLL) is proposed with two new mechanisms: kernel-based feature reduction module and correlation-based label thresholding. The kernel-based feature reduction module contains three layers, namely, the feature mapping layer, enhancement nodes layer, and feature reduction layer. The feature mapping layer employs elastic network regularization to solve the randomness of features in order to improve performance. In the enhancement nodes layer, the kernel method is applied for high-dimensional nonlinear conversion to achieve high efficiency. The newly constructed feature reduction layer is used to further significantly improve both the training efficiency and accuracy when facing high-dimensionality with abundant or noisy information embedded in large-scale data. The correlation-based label thresholding enables BLS-MLL to generate a label-thresholding function for effective conversion of the final decision values to logical outputs, thus, improving the classification performance. Finally, experimental comparisons among six state-of-the-art multi-label classifiers on ten datasets demonstrate the effectiveness of the proposed BLS-MLL. The results of the classification performance show that BLS-MLL outperforms the compared algorithms in 86% of cases with better training efficiency in 90% of cases.
Chi-Man Vong, C. L. Philip Chen, Yimin Zhou 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 A Novel Copy-Move Forgery Detection Algorithm via Feature Label Matching and Hierarchical Segmentation Filtering
Chi-Man Vong
Inf. Process. Manag.3
2022 Recursive least mean dual p-power solution to the generalization of evolving fuzzy system under multiple noises
Hai-Jun Rong, Zhao-Xu Yang, Chi-Man Vong
Inf. Sci.4
2022 Easy Domain Adaptation for cross-subject multi-view emotion recognition
Chuangquan Chen, Chi-Man Vong, Shitong Wang 0001, Hongtao Wang 0001, Miaoqi Pang
Knowl. Based Syst.2
2022 Multi-scale Multi-instance Multi-feature Joint Learning Broad Network (M3JLBN) for gastric intestinal metaplasia subtype classification
Qi Lai, Chi-Man Vong, Pak-Kin Wong 0001, Shitong Wang 0001, Tao Yan 0006, I. Cheong Choi, Hon Ho Yu
Knowl. Based Syst.2
2022 Effective and efficient pixel-level detection for diverse video copy-move forgery types
Chi-Man Vong, Ji-Xiang Yang, Jing-Hong Zhao, Jiahua Luo
Pattern Recognit.3
2022 Fuzzy KNN Method With Adaptive Nearest Neighbors
abstract
Due to its strong performance in handling uncertain and ambiguous data, the fuzzy k -nearest-neighbor method (FKNN) has realized substantial success in a wide variety of applications. However, its classification performance would be heavily deteriorated if the number k of nearest neighbors was unsuitably fixed for each testing sample. This study examines the feasibility of using only one fixed k value for FKNN on each testing sample. A novel FKNN-based classification method, namely, fuzzy KNN method with adaptive nearest neighbors (A-FKNN), is devised for learning a distinct optimal k value for each testing sample. In the training stage, after applying a sparse representation method on all training samples for reconstruction, A-FKNN learns the optimal k value for each training sample and builds a decision tree (namely, A-FKNN tree) from all training samples with new labels (the learned optimal k values instead of the original labels), in which each leaf node stores the corresponding optimal k value. In the testing stage, A-FKNN identifies the optimal k value for each testing sample by searching the A-FKNN tree and runs FKNN with the optimal k value for each testing sample. Moreover, a fast version of A-FKNN, namely, FA-FKNN, is designed by building the FA-FKNN decision tree, which stores the optimal k value with only a subset of training samples in each leaf node. Experimental results on 32 UCI datasets demonstrate that both A-FKNN and FA-FKNN outperform the compared methods in terms of classification accuracy, and FA-FKNN has a shorter running time.
Zekang Bian, Chi-Man Vong, Pak-Kin Wong 0001, Shitong Wang 0001
IEEE Trans. Cybern.2
2022 Fast Training of Adversarial Deep Fuzzy Classifier by Downsizing Fuzzy Rules With Gradient Guided Learning
abstract
While our recent deep fuzzy classifier DSA-FC, which stacks adversarial interpretable Takagi–Sugeno–Kang fuzzy subclassifiers, shares its promising classification, its training speed will become very slow and even intolerable for large-scale datasets, due to successive training on all training samples with their random gradient based updates along each layer of its stacked structure. In order to circumvent this bottleneck issue, a fast training algorithm FTA is developed in this study by downsizing fuzzy rules with the proposed gradient guided learning for each subclassifier at each layer of DSA-FC on large-scale datasets. The core of FTA is to assure fast training of each subclassifier at each layer of DSA-FC, which first generates first-order smooth gradient guided information by means of the proposed top-kfuzzy rules selected from all fuzzy rules in each subclassifier, and then quickly updates the current inputs in terms of such information, which will be taken as the inputs of the subclassifier at the next layer. Our theoretical analysis reveals that the proposed gradient guided learning indeed enhances the generalization capability of a deep fuzzy classifier with or without adversarial attacks on outputs. Experimental results on large datasets demonstrate that FTA indeed trains the deep fuzzy classifier DSA-FC quickly with enhanced generalization capability.
Suhang Gu, Chi-Man Vong, Pak-Kin Wong 0001, Shitong Wang 0001
IEEE Trans. Fuzzy Syst.2
2022 Feature-Based Direct Tracking and Mapping for Real-Time Noise-Robust Outdoor 3D Reconstruction Using Quadcopters
abstract
In this work, we focus on real-time 3D reconstruction or localization and mapping for outdoor scene using an aerial vehicle called quadcopter. Quadcopter provides the advantages of high flexibility and wide view field in spatial movement. However, existing feature-based and direct methods (using dense or semi-dense approach) are not suitable for outdoor environment, in which multiple challenging scenarios arise such as lighting variance, jittering views, high-speed and non-smooth flight trajectory. The main reason is that the existing methods rely on the assumption of brightness constancy across multiple images and only raw pixel intensities are employed for direct image alignment. In order to tackle these scenarios, a novel method called Feature-based Direct Tracking and Mapping (FDTAM) is proposed, which i) incorporates an efficient binary feature descriptor into direct image alignment module to tackle the challenging scenarios, such as drifting issue under lighting variance problem; ii) applies semi-dense approach to obtain high reconstruction quality; iii) provides a framework with low computational complexity for real-time reconstruction. Compared to other state-of-the-art feature-based and direct methods, our proposed method is shown to tackle the challenging scenarios and improve the accuracy and robustness even in CPU (rather than GPU) platform.
Chi-Chong Wong, Chi-Man Vong, Yimin Zhou 0001
IEEE Trans. Intell. Transp. Syst.2
2021 Persistent Homology based Graph Convolution Network for Fine-grained 3D Shape Segmentation
abstract
Fine-grained 3D segmentation is an important task in 3D object understanding, especially in applications such as intelligent manufacturing or parts analysis for 3D objects. However, many challenges involved in such problem are yet to be solved, such as i) interpreting the complex structures located in different regions for 3D objects; ii) capturing fine-grained structures with sufficient topology correctness. Current deep learning and graph machine learning methods fail to tackle such challenges and thus provide inferior performance in fine-grained 3D analysis. In this work, methods in topological data analysis are incorporated with geometric deep learning model for the task of fine-grained segmentation for 3D objects. We propose a novel neural network model called Persistent Homology based Graph Convolution Network (PHGCN), which i) integrates persistent homology into graph convolution network to capture multi-scale structural information that can accurately represent complex structures for 3D objects; ii) applies a novel Persistence Diagram Loss (ℒPD) that provides sufficient topology correctness for segmentation over the fine-grained structures. Extensive experiments on fine-grained 3D segmentation validate the effectiveness of the proposed PHGCN model and show significant improvements over current state-of-the-art methods.
Chi-Chong Wong, Chi-Man Vong
ICCV2
2021 Improving Conversational Recommender System by Pretraining Billion-scale Knowledge Graph
abstract
Conversational Recommender Systems (CRSs) in E-commerce platforms aim to recommend items to users via multiple conversational interactions. Click-through rate (CTR) prediction models are commonly used for ranking candidate items. However, most CRSs are suffer from the problem of data scarcity and sparseness. To address this issue, we propose a novel knowledge-enhanced deep cross network (K-DCN), a two-step (pretrain and fine-tune) CTR prediction model to recommend items. We first construct a billion-scale conversation knowledge graph (CKG) from information about users, items and converations, and then pretrain CKG by introducing knowledge graph embedding method and graph convolution network to encode semantic and structural information respectively. To make the CTR prediction model sensible of current state of users and the relationship between dialogues and items, we introduce user-state and dialogue-interaction representations based on pre-trained CKG and propose K-DCN. In K-DCN, we fuse the user-state representation, dialogue-interaction representation and other normal feature representations via deep cross network, which will give the rank of candidate items to be recommended. We experimentally prove that our proposal significantly outperforms baselines and show it's real application in Alime.
Chiman Wong, Wen Zhang 0015, Chi-Man Vong, Hui Chen 0018, Yichi Zhang 0009, Huajun Chen
ICDE4
2021 Light-weight network for real-time adaptive stereo depth estimation
Wanshui Gan, Pak-Kin Wong 0001, Guokuan Yu, Rongchen Zhao, Chi-Man Vong
Neurocomputing5
2021 Scalable and memory-efficient sparse learning for classification with approximate Bayesian regularization priors
Jiahua Luo, Chi-Man Vong, Chiman Wong, Chuangquan Chen
Neurocomputing3
2021 Multinomial Bayesian extreme learning machine for sparse and accurate classification model
Jiahua Luo, Chiman Wong, Chi-Man Vong
Neurocomputing3
2021 Jointly evolving and compressing fuzzy system for feature reduction and classification
Hai-Jun Rong, Zhao-Xu Yang, Chi-Man Vong
Inf. Sci.4
2021 Novel Efficient RNN and LSTM-Like Architectures: Recurrent and Gated Broad Learning Systems and Their Applications for Text Classification
abstract
High accuracy of text classification can be achieved through simultaneous learning of multiple information, such as sequence information and word importance. In this article, a kind of flat neural networks called the broad learning system (BLS) is employed to derive two novel learning methods for text classification, including recurrent BLS (R-BLS) and long short-term memory (LSTM)-like architecture: gated BLS (G-BLS). The proposed two methods possess three advantages: 1) higher accuracy due to the simultaneous learning of multiple information, even compared to deep LSTM that extracts deeper but single information only; 2) significantly faster training time due to the noniterative learning in BLS, compared to LSTM; and 3) easy integration with other discriminant information for further improvement. The proposed methods have been evaluated over 13 real-world datasets from various types of text classification. From the experimental results, the proposed methods achieve higher accuracies than LSTM while taking significantly less training time on most evaluated datasets, especially when the LSTM is in deep architecture. Compared to R-BLS, G-BLS has an extra forget gate to control the flow of information (similar to LSTM) to further improve the accuracy on text classification so that G-BLS is more effective while R-BLS is more efficient.
Jie Du 0001, Chi-Man Vong, C. L. Philip Chen
IEEE Trans. Cybern.2
2021 Ground Plane Context Aggregation Network for Day-and-Night on Vehicular Pedestrian Detection
abstract
Ground plane context is an essential semantic information in on-road pedestrian detection task. Due to viewpoint geometry constraints, pedestrians only appear in certain regions of the image, which is close to the horizon area of the ground plane. As a result, the lacking to ground plane context information may cause pedestrian detection system suffering from severe false alarm (i.e. high false positive (FP) rate). For Advanced Driver Assistance System (ADAS), high FP rate not only distracts the driver, but also causes frequent unexpected braking to damage the vehicle’s hardware. In this paper, a novel pedestrian detection method called ground plane context aggregation network (GPCAnet) is proposed, which integrates ground plane context information into deep learning based detector to drastically reduce the FP rate. The proposed GPCAnet consists of two modules: i) a ground area predication (GAP) branch is appended on top of convolutional feature map of the backbone network, in parallel with existing branches, for region proposal, classification and bounding box regression; ii) based on GAP, a ground-region proposal network (GRPN) is designed to filter FP cases in order to reduce computations. To evaluate the effectiveness of proposed GPCAnet, experiments on day and night on-road pedestrian detection are performed on both visible and far infrared pedestrian detection datasets, e.g. Caltech and SCUT. Experimental results show that GPCAnet achieves better performance than state-of-the-art methods, while drastically reducing FP rate in pedestrian detection.
Zhewei Xu, Chi-Man Vong, Chi-Chong Wong, Qiong Liu 0006
IEEE Trans. Intell. Transp. Syst.2
2020 Efficient Outdoor 3D Point Cloud Semantic Segmentation for Critical Road Objects and Distributed Contexts
Chi-Chong Wong, Chi-Man Vong
ECCV (27)2
2020 Extreme semi-supervised learning for multiclass classification
Chuangquan Chen, Chi-Man Vong
Neurocomputing3
2020 Novel up-scale feature aggregation for object detection in aerial images
Jingkai Zhou, Chi-Man Vong, Qiong Liu 0006
Neurocomputing4
2020 SCNet: Scale-aware coupling-structure network for efficient video object detection
Fengchao Wang, Zhewei Xu, Chi-Man Vong, Qiong Liu 0006
Neurocomputing4
2020 Tracking objects with partial occlusion by background alignment
Feng Wu 0004, Chi-Man Vong, Qiong Liu 0006
Neurocomputing2
2020 Approximate empirical kernel map-based iterative extreme learning machine for clustering
Chuangquan Chen, Chi-Man Vong, Pak-Kin Wong 0001, Keng Iam Tai
Neural Comput. Appl.2
2020 Adaptive neural tracking control for automotive engine idle speed regulation using extreme learning machine
Pak-Kin Wong 0001, Chi-Man Vong, Zhi-Xin Yang 0001
Neural Comput. Appl.3
2020 Accurate and efficient sequential ensemble learning for highly imbalanced multi-class data
Chi-Man Vong, Jie Du 0001
Neural Networks1
2020 Robust Online Multilabel Learning Under Dynamic Changes in Data Distribution With Labels
abstract
In this paper, a robust online multilabel learning method dealing with dynamically changing multilabel data streams is proposed. The proposed method has three advantages: 1) higher accuracy due to a newly defined objective function based on labels ranking; 2) fast training and update based on a newly derived closed-form (rather than gradient descent based) solution for the new objective function; and 3) high robustness to a newly identified concept drift in multilabel data streams, namely, changes in data distribution with labels (CDDL). The high robustness benefits from two novel works: 1) a new sequential update rule that preserves the labels ranking information learned from all old (but discarded) samples while updating the model only based on new incoming samples and 2) a fixed threshold for label bipartition that is insensitive to any kind of changes in data distribution including CDDL. The proposed method has been evaluated over 13 benchmark datasets from various domains. As shown in the experimental results, the proposed work is highly robust to CDDL in both the sequential model update and multilabel thresholding. Furthermore, the proposed method improves the performance in different evaluation measures, including Hamming loss, F1-measure, Precision, and Recall while taking short training time on most evaluated datasets.
Jie Du 0001, Chi-Man Vong
IEEE Trans. Cybern.2
2020 Efficient Outdoor Video Semantic Segmentation Using Feedback-Based Fully Convolution Neural Network
abstract
In this article, we focus on efficient semantic segmentation problem from sequential two-dimensional images, in which all pixels are classified into certain classes for scene understanding. Such problem is challenging because it involves constraints of both spatial and temporal consistencies, which have large difficulties in explicitly determining such structural constraints. Traditionally, such a problem is tackled using structured prediction method, such as conditional random field (CRF). However, pure CRF method suffers from very high complexity in computing high-order potentials and slow performance during inference step, which is unsuitable for efficient video segmentation in real scenario. In this article, a novel feedback-based deep fully convolutional neural network (CNN) is proposed to inherently incorporate spatial context through appending output feedback mechanism. The proposed method has the following contributions: 1) spatial context in images are easily captured through iterative feedback refinement, without the expensive postprocess step such as CRF refinement; 2) easily integrated with generic deep CNN structure; and 3) the inference time is greatly reduced for efficient image segmentation. Compared to current state-of-the-art methods, our proposed method was shown to provide up to 14% better accuracy on semantic segmentation task in challenging Camvid and Cityscapes datasets, while taking up to relatively 980% shorter inference time. The proposed method also shows its effectiveness for real-time road detection task of autonomous driving.
Chi-Chong Wong, Chi-Man Vong
IEEE Trans. Ind. Informatics3
2019 An Enhanced Hierarchical Extreme Learning Machine with Random Sparse Matrix Based Autoencoder
abstract
Recently, by employing the stacked extreme learning machine (ELM) based autoencoders (ELM-AE) and sparse AEs (SAE), multilayer ELM (ML-ELM) and hierarchical ELM (H-ELM) has been developed. Compared to the conventional stacked AEs, the ML-ELM and H-ELM usually achieve better generalization performance with a significantly reduced training time. However, the ℓ1-norm based SAE may suffer the overfitting problem and it is unable to provide analytical solution leading to long training time for big data. To alleviate these deficiencies, we propose an enhanced H-ELM (EH-ELM) with a novel random sparse matrix based AE (SMA) in this paper. The contributions are in two aspects, 1) utilizing the random sparse matrix, the sparse features can be obtained; 2) the proposed SMA can provide an analytical solution so that the high computational complexity issue in SAE can be addressed. Experimental results on benchmark datasets show that the proposed EH-ELM achieves a higher recognition rate and a faster training speed than H-ELM and ML-ELM.
Tianlei Wang, Xiaoping Lai, Jiuwen Cao, Chi-Man Vong, Badong Chen
ICASSP4
2019 3DViewGraph: Learning Global Features for 3D Shapes from A Graph of Unordered Views with Attention
abstract
Learning global features by aggregating information over multiple views has been shown to be effective for 3D shape analysis. For view aggregation in deep learning models, pooling has been applied extensively. However, pooling leads to a loss of the content within views, and the spatial relationship among views, which limits the discriminability of learned features. We propose 3DViewGraph to resolve this issue, which learns 3D global features by more effectively aggregating unordered views with attention. Specifically, unordered views taken around a shape are regarded as view nodes on a view graph. 3DViewGraph first learns a novel latent semantic mapping to project low-level view features into meaningful latent semantic embeddings in a lower dimensional space, which is spanned by latent semantic patterns. Then, the content and spatial information of each pair of view nodes are encoded by a novel spatial pattern correlation, where the correlation is computed among latent semantic patterns. Finally, all spatial pattern correlations are integrated with attention weights learned by a novel attention mechanism. This further increases the discriminability of learned features by highlighting the unordered view nodes with distinctive characteristics and depressing the ones with appearance ambiguity. We show that 3DViewGraph outperforms state-of-the-art methods under three large-scale benchmarks.
Zhizhong Han, Xiyang Wang 0001, Chi-Man Vong, Yu-Shen Liu, Matthias Zwicker, C. L. Philip Chen
IJCAI3
2019 Scale adaptive image cropping for UAV object detection
Jingkai Zhou, Chi-Man Vong, Qiong Liu 0006, Zhenyu Wang 0001
Neurocomputing2
2019 Unsupervised Learning of 3-D Local Features From Raw Voxels Based on a Novel Permutation Voxelization Strategy
abstract
Effective 3-D local features are significant elements for 3-D shape analysis. Existing hand-crafted 3-D local descriptors are effective but usually involve intensive human intervention and prior knowledge, which burdens the subsequent processing procedures. An alternative resorts to the unsupervised learning of features from raw 3-D representations via popular deep learning models. However, this alternative suffers from several significant unresolved issues, such as irregular vertex topology, arbitrary mesh resolution, orientation ambiguity on the 3-D surface, and rigid and slightly nonrigid transformation invariance. To tackle these issues, we propose an unsupervised 3-D local feature learning framework based on a novel permutation voxelization strategy to learn high-level and hierarchical 3-D local features from raw 3-D voxels. Specifically, the proposed strategy first applies a novel voxelization which discretizes each 3-D local region with irregular vertex topology and arbitrary mesh resolution into regular voxels, and then, a novel permutation is applied to permute the voxels to simultaneously eliminate the effect of rotation transformation and orientation ambiguity on the surface. Based on the proposed strategy, the permuted voxels can fully encode the geometry and structure of each local region in regular, sparse, and binary vectors. These voxel vectors are highly suitable for the learning of hierarchical common surface patterns by stacked sparse autoencoder with hierarchical abstraction and sparse constraint. Experiments are conducted on three aspects for evaluating the learned local features: 1) global shape retrieval; 2) partial shape retrieval; and 3) shape correspondence. The experimental results show that the learned local features outperform the other state-of-the-art 3-D shape descriptors.
Zhizhong Han, Zhenbao Liu, Junwei Han 0001, Chi-Man Vong, Shuhui Bu, C. L. Philip Chen
IEEE Trans. Cybern.4
2019 3D2SeqViews: Aggregating Sequential Views for 3D Global Feature Learning by CNN With Hierarchical Attention Aggregation
abstract
Learning 3D global features by aggregating multiple views is important. Pooling is widely used to aggregate views in deep learning models. However, pooling disregards a lot of content information within views and the spatial relationship among the views, which limits the discriminability of learned features. To resolve this issue, 3D to Sequential Views (3D2SeqViews) is proposed to more effectively aggregate the sequential views using convolutional neural networks with a novel hierarchical attention aggregation. Specifically, the content information within each view is first encoded. Then, the encoded view content information and the sequential spatiality among the views are simultaneously aggregated by the hierarchical attention aggregation, where view-level attention and class-level attention are proposed to hierarchically weight sequential views and shape classes. View-level attention is learned to indicate how much attention is paid to each view by each shape class, which subsequently weights sequential views through a novel recursive view integration. Recursive view integration learns the semantic meaning of view sequence, which is robust to the first view position. Furthermore, class-level attention is introduced to describe how much attention is paid to each shape class, which innovatively employs the discriminative ability of the fine-tuned network. 3D2SeqViews learns more discriminative features than the state-of-the-art, which leads to the outperforming results in shape classification and retrieval under three large-scale benchmarks.
Zhizhong Han, Honglei Lu, Zhenbao Liu, Chi-Man Vong, Yu-Shen Liu, Matthias Zwicker, Junwei Han 0001, C. L. Philip Chen
IEEE Trans. Image Process.4
2019 SeqViews2SeqLabels: Learning 3D Global Features via Aggregating Sequential Views by RNN With Attention
abstract
Learning 3D global features by aggregating multiple views has been introduced as a successful strategy for 3D shape analysis. In recent deep learning models with end-to-end training, pooling is a widely adopted procedure for view aggregation. However, pooling merely retains the max or mean value over all views, which disregards the content information of almost all views and also the spatial information among the views. To resolve these issues, we propose Sequential Views To Sequential Labels (SeqViews2SeqLabels) as a novel deep learning model with an encoder-decoder structure based on recurrent neural networks (RNNs) with attention. SeqViews2SeqLabels consists of two connected parts, an encoder-RNN followed by a decoder-RNN, that aim to learn the global features by aggregating sequential views and then performing shape classification from the learned global features, respectively. Specifically, the encoder-RNN learns the global features by simultaneously encoding the spatial and content information of sequential views, which captures the semantics of the view sequence. With the proposed prediction of sequential labels, the decoder-RNN performs more accurate classification using the learned global features by predicting sequential labels step by step. Learning to predict sequential labels provides more and finer discriminative information among shape classes to learn, which alleviates the overfitting problem inherent in training using a limited number of 3D shapes. Moreover, we introduce an attention mechanism to further improve the discriminative ability of SeqViews2SeqLabels. This mechanism increases the weight of views that are distinctive to each shape class, and it dramatically reduces the effect of selecting the first view position. Shape classification and retrieval results under three large-scale benchmarks verify that SeqViews2SeqLabels learns more discriminative global features by more effectively aggregating sequential views than state-of-the-art methods.
Zhizhong Han, Mingyang Shang, Zhenbao Liu, Chi-Man Vong, Yu-Shen Liu, Matthias Zwicker, Junwei Han 0001, C. L. Philip Chen
IEEE Trans. Image Process.4
2018 DAliM: Machine Learning Based Intelligent Lucky Money Determination for Large-Scale E-Commerce Businesses
Min Fu 0001, Chiman Wong, Yanjun Huang, Yuanping Li, James Xi Zheng, Jia Wu 0001, Jian Yang 0001, Chi-Man Vong
ICSOC9
2018 Empirical kernel map-based multilayer extreme learning machines for representation learning
Chi-Man Vong, Chuangquan Chen, Pak-Kin Wong 0001
Neurocomputing1
2018 Online extreme learning machine based modeling and optimization for point-by-point engine calibration
Pak-Kin Wong 0001, Xiang Hui Gao, Ka In Wong, Chi-Man Vong
Neurocomputing4
2018 Efficient extreme learning machine via very sparse random projection
Chuangquan Chen, Chi-Man Vong, Chiman Wong, Weiru Wang 0001, Pak-Kin Wong 0001
Soft Comput.2
2018 Deep Spatiality: Unsupervised Learning of Spatially-Enhanced Global and Local 3D Features by Deep Neural Network With Coupled Softmax
abstract
The discriminability of Bag-of-Words representations can be increased via encoding the spatial relationship among virtual words on 3D shapes. However, this encoding task involves several issues, including arbitrary mesh resolutions, irregular vertex topology, orientation ambiguity on 3D surface, invariance to rigid and non-rigid shape transformations. To address these issues, a novel unsupervised spatial learning framework based on deep neural network, deep spatiality (DS), is proposed. Specifically, DS employs two novel components: spatial context extractor and deep context learner. Spatial context extractor extracts the spatial relationship among virtual words in a local region into a raw spatial representation. Along a consistent circular direction, a directed circular graph is constructed to encode relative positions between pairwise virtual words in each face ring into a relative spatial matrix. By decomposing each relative spatial matrix using SVD, the raw spatial representation is formed, from which deep context learner conducts unsupervised learning of global and local features. Deep context learner is a deep neural network with a novel model structure to adapt the proposed coupled softmax layer, which encodes not only the discriminative information among local regions but also the one among global shapes. Experimental results show that DS outperforms state-of-the-art methods.
Zhizhong Han, Zhenbao Liu, Chi-Man Vong, Yu-Shen Liu, Shuhui Bu, Junwei Han 0001, C. L. Philip Chen
IEEE Trans. Image Process.3
2018 Postboosting Using Extended G-Mean for Online Sequential Multiclass Imbalance Learning
abstract
In this paper, a novel learning method called postboosting using extended G-mean (PBG) is proposed for online sequential multiclass imbalance learning (OS-MIL) in neural networks. PBG is effective due to three reasons. 1) Through postadjusting a classification boundary under extended G-mean, the challenging issue of imbalanced class distribution for sequentially arriving multiclass data can be effectively resolved. 2) A newly derived update rule for online sequential learning is proposed, which produces a high G-mean for current model and simultaneously possesses almost the same information of its previous models. 3) A dynamic adjustment mechanism provided by extended G-mean is valid to deal with the unresolved challenging dense-majority problem and two dynamic changing issues, namely, dynamic changing data scarcity (DCDS) and dynamic changing data diversity (DCDD). Compared to other OS-MIL methods, PBG is highly effective on resolving DCDS, while PBG is the only method to resolve dense-majority and DCDD. Furthermore, PBG can directly and effectively handle unscaled data stream. Experiments have been conducted for PBG and two popular OS-MIL methods for neural networks under massive binary and multiclass data sets. Through the analyses of experimental results, PBG is shown to outperform the other compared methods on all data sets in various aspects including the issues of data scarcity, dense-majority, DCDS, DCDD, and unscaled data.
Chi-Man Vong, Jie Du 0001, Chiman Wong, Jiuwen Cao
IEEE Trans. Neural Networks Learn. Syst.1
2018 Kernel-Based Multilayer Extreme Learning Machines for Representation Learning
abstract
Recently, multilayer extreme learning machine (ML-ELM) was applied to stacked autoencoder (SAE) for representation learning. In contrast to traditional SAE, the training time of ML-ELM is significantly reduced from hours to seconds with high accuracy. However, ML-ELM suffers from several drawbacks: 1) manual tuning on the number of hidden nodes in every layer is an uncertain factor to training time and generalization; 2) random projection of input weights and bias in every layer of ML-ELM leads to suboptimal model generalization; 3) the pseudoinverse solution for output weights in every layer incurs relatively large reconstruction error; and 4) the storage and execution time for transformation matrices in representation learning are proportional to the number of hidden layers. Inspired by kernel learning, a kernel version of ML-ELM is developed, namely, multilayer kernel ELM (ML-KELM), whose contributions are: 1) elimination of manual tuning on the number of hidden nodes in every layer; 2) no random projection mechanism so as to obtain optimal model generalization; 3) exact inverse solution for output weights is guaranteed under invertible kernel matrix, resulting to smaller reconstruction error; and 4) all transformation matrices are unified into two matrices only, so that storage can be reduced and may shorten model execution time. Benchmark data sets of different sizes have been employed for the evaluation of ML-KELM. Experimental results have verified the contributions of the proposed ML-KELM. The improvement in accuracy over benchmark data sets is up to 7%.
Chiman Wong, Chi-Man Vong, Pak-Kin Wong 0001, Jiuwen Cao
IEEE Trans. Neural Networks Learn. Syst.2
2017 Advances in extreme learning machines (ELM2015)
Amaury Lendasse, Chi-Man Vong, Kar-Ann Toh, Yoan Miché, Guang-Bin Huang
Neurocomputing2
2017 A novel meta-cognitive fuzzy-neural model with backstepping strategy for adaptive control of uncertain nonlinear systems
Hai-Jun Rong, Zhao-Xu Yang, Pak-Kin Wong 0001, Chi-Man Vong, Guang-She Zhao
Neurocomputing4
2017 Efficient shape classification using region descriptors
Cong Lin 0001, Chi-Man Pun, Chi-Man Vong, Donald A. Adjeroh
Multim. Tools Appl.3
2017 Post-boosting of classification boundary for imbalanced data using geometric mean
Jie Du 0001, Chi-Man Vong, Chi-Man Pun, Pak-Kin Wong 0001, Weng-Fai Ip
Neural Networks2
2017 Capturing High-Discriminative Fault Features for Electronics-Rich Analog System via Deep Learning
abstract
Fault detection and isolation (FDI) is very difficult for electronics-rich analog systems due to its sophisticated mechanism and variable operational conditions. Traditionally, FDI in such systems is done through the monitoring of deviation of output signals in voltage or current at system level, which commonly arises from the degradation of one or more critical components. Therefore, FDI can be transformed to a multiclass classification task given the extracted features of the output signals in voltage or current of the circuit. Traditional feature extraction on the circuit output is mostly based on time-domain, frequency-domain, or time-frequency signal processing, which collapse high-dimensional raw signals into a lower dimensional feature set. Such low-dimensional feature set usually suffers from information loss so as to affect the accuracy of the later fault diagnosis. In order to retain as much information as possible, deep learning is proposed which employs a hierarchical structure to capture the different levels of semantic representations of the signals. In this paper, a novel fault diagnostic application of Gaussian-Bernoulli deep belief network (GB-DBN) for electronics-rich analog systems is developed which can more effectively capture the high-order semantic features within the raw output signals. The novel fault diagnosis is validated experimentally on two typical analog filter circuits. Experimental results show the fault diagnosis based on GB-DBN is with superior diagnostic performance than the traditional feature extraction methods.
Zhenbao Liu, Chi-Man Vong, Shuhui Bu, Junwei Han 0001
IEEE Trans. Ind. Informatics3
2017 BoSCC: Bag of Spatial Context Correlations for Spatially Enhanced 3D Shape Representation
abstract
Highly discriminative 3D shape representations can be formed by encoding the spatial relationship among virtual words into the Bag of Words (BoW) method. To achieve this challenging task, several unresolved issues in the encoding procedure must be overcome for 3D shapes, including: 1) arbitrary mesh resolution; 2) irregular vertex topology; 3) orientation ambiguity on the 3D surface; and 4) invariance to rigid and non-rigid shape transformations. In this paper, a novel spatially enhanced 3D shape representation called bag of spatial context correlations (BoSCCs) is proposed to address all these issues. Adopting a novel local perspective, BoSCC is able to describe a 3D shape by an occurrence frequency histogram of spatial context correlation patterns, which makes BoSCC become more compact and discriminative than previous global perspective-based methods. Specifically, the spatial context correlation is proposed to simultaneously encode the geometric and spatial information of a 3D local region by the correlation among spatial contexts of vertices in that region, which effectively resolves the aforementioned issues. The spatial context of each vertex is modeled by Markov chains in a multi-scale manner, which thoroughly captures the spatial relationship by the transition probabilities of intra-virtual words and the ones of inter-virtual words. The high discriminability and compactness of BoSCC are effective for classification and retrieval, especially in the scenarios of limited samples and partial shape retrieval. Experimental results show that BoSCC outperforms the state-of-the-art spatially enhanced BoW methods in three common applications: global shape retrieval, shape classification, and partial shape retrieval.
Zhizhong Han, Zhenbao Liu, Chi-Man Vong, Yu-Shen Liu, Shuhui Bu, Junwei Han 0001, C. L. Philip Chen
IEEE Trans. Image Process.3
2017 Mesh Convolutional Restricted Boltzmann Machines for Unsupervised Learning of Features With Structure Preservation on 3-D Meshes
abstract
Discriminative features of 3-D meshes are significant to many 3-D shape analysis tasks. However, handcrafted descriptors and traditional unsupervised 3-D feature learning methods suffer from several significant weaknesses: 1) the extensive human intervention is involved; 2) the local and global structure information of 3-D meshes cannot be preserved, which is in fact an important source of discriminability; 3) the irregular vertex topology and arbitrary resolution of 3-D meshes do not allow the direct application of the popular deep learning models; 4) the orientation is ambiguous on the mesh surface; and 5) the effect of rigid and nonrigid transformations on 3-D meshes cannot be eliminated. As a remedy, we propose a deep learning model with a novel irregular model structure, called mesh convolutional restricted Boltzmann machines (MCRBMs). MCRBM aims to simultaneously learn structure-preserving local and global features from a novel raw representation, local function energy distribution. In addition, multiple MCRBMs can be stacked into a deeper model, called mesh convolutional deep belief networks (MCDBNs). MCDBN employs a novel local structure preserving convolution (LSPC) strategy to convolve the geometry and the local structure learned by the lower MCRBM to the upper MCRBM. LSPC facilitates resolving the challenging issue of the orientation ambiguity on the mesh surface in MCDBN. Experiments using the proposed MCRBM and MCDBN were conducted on three common aspects: global shape retrieval, partial shape retrieval, and shape correspondence. Results show that the features learned by the proposed methods outperform the other state-of-the-art 3-D shape features.
Zhizhong Han, Zhenbao Liu, Junwei Han 0001, Chi-Man Vong, Shuhui Bu, C. L. Philip Chen
IEEE Trans. Neural Networks Learn. Syst.4
2016 Adaptive control of rapidly time-varying discrete-time system using initial-training-free online extreme learning machine
Xiang Hui Gao, Ka In Wong, Pak-Kin Wong 0001, Chi-Man Vong
Neurocomputing4
2016 Advances in extreme learning machines (ELM2014)
Amaury Lendasse, Chi-Man Vong, Yoan Miché, Guang-Bin Huang
Neurocomputing2
2016 Sparse Bayesian extreme learning committee machine for engine simultaneous fault diagnosis
Pak-Kin Wong 0001, Jianhua Zhong, Zhi-Xin Yang 0001, Chi-Man Vong
Neurocomputing4
2016 Fast detection of impact location using kernel extreme learning machine
Heming Fu, Chi-Man Vong, Pak-Kin Wong 0001, Zhi-Xin Yang 0001
Neural Comput. Appl.2
2016 Model predictive engine air-ratio control using online sequential extreme learning machine
Pak-Kin Wong 0001, Hang-Cheong Wong, Chi-Man Vong, Zhengchao Xie, Shaojia Huang
Neural Comput. Appl.3
2016 Unsupervised 3D Local Feature Learning by Circle Convolutional Restricted Boltzmann Machine
abstract
Extracting local features from 3D shapes is an important and challenging task that usually requires carefully designed 3D shape descriptors. However, these descriptors are hand-crafted and require intensive human intervention with prior knowledge. To tackle this issue, we propose a novel deep learning model, namely circle convolutional restricted Boltzmann machine (CCRBM), for unsupervised 3D local feature learning. CCRBM is specially designed to learn from raw 3D representations. It effectively overcomes obstacles such as irregular vertex topology, orientation ambiguity on the 3D surface, and rigid or slightly non-rigid transformation invariance in the hierarchical learning of 3D data that cannot be resolved by the existing deep learning models. Specifically, by introducing the novel circle convolution, CCRBM holds a novel ring-like multi-layer structure to learn 3D local features in a structure preserving manner. Circle convolution convolves across 3D local regions via rotating a novel circular sector convolution window in a consistent circular direction. In the process of circle convolution, extra points are sampled in each 3D local region and projected onto the tangent plane of the center of the region. In this way, the projection distances in each sector window are employed to constitute a novel local raw 3D representation called projection distance distribution (PDD). In addition, to eliminate the initial location ambiguity of a sector window, the Fourier transform modulus is used to transform the PDD into the Fourier domain, which is then conveyed to CCRBM. Experiments using the learned local features are conducted on three aspects: global shape retrieval, partial shape retrieval, and shape correspondence. The experimental results show that the learned local features outperform other state-of-the-art 3D shape descriptors.
Zhizhong Han, Zhenbao Liu, Junwei Han 0001, Chi-Man Vong, Shuhui Bu, Xuelong Li 0001
IEEE Trans. Image Process.4
2015 Sparse Bayesian extreme learning machine and its application to biofuel engine performance prediction
Ka In Wong, Chi-Man Vong, Pak-Kin Wong 0001, Jiahua Luo
Neurocomputing2
2015 Fast and accurate face detection by sparse Bayesian extreme learning machine
Chi-Man Vong, Keng Iam Tai, Chi-Man Pun, Pak-Kin Wong 0001
Neural Comput. Appl.1
2014 Hybrid model predictive controller for engine air-ratio control
abstract
Air-ratio is an important engine parameter which relates closely to engine emissions, power, and brake-specific fuel consumption. Model predictive controller (MPC) is a well-known technique for air-ratio control. This paper utilizes two advanced techniques, discrete wavelet transformation (DWT) and relevance vector machine (RVM), to develop wavelet relevance vector machine model predictive controller (W-MPC) for air-ratio regulation. To compensate for the modelling error of W-MPC and system disturbances, the W-MPC is proposed to connect in parallel with a proportional-integral (PI) controller so as to form a new hybrid model predictive controller (H-MPC). The proposed H-MPC is implemented on a real engine to evaluate its effectiveness. Its control performance is also compared with the W-MPC without PI controller and the latest MPC for engine air-ratio control in the literature. Experimental results show the superiority of the proposed H-MPC over the other two controllers, which can more effectively regulate the air-ratio to target values under external disturbance. Therefore, the proposed H-MPC is a promising scheme for engine air-ratio control.
Pak-Kin Wong 0001, Hang-Cheong Wong, Tong Meng Iong, Chi-Man Vong
ICARCV4
2014 Predicting minority class for suspended particulate matters level by extreme learning machine
Chi-Man Vong, Weng-Fai Ip, Pak-Kin Wong 0001, Chi-Chong Chiu
Neurocomputing1
2014 Real-time fault diagnosis for gas turbine generator systems using extreme learning machine
Pak-Kin Wong 0001, Zhi-Xin Yang 0001, Chi-Man Vong, Jianhua Zhong
Neurocomputing3
2014 Sparse Bayesian Extreme Learning Machine for Multi-classification
abstract
Extreme learning machine (ELM) has become a popular topic in machine learning in recent years. ELM is a new kind of single-hidden layer feedforward neural network with an extremely low computational cost. ELM, however, has two evident drawbacks: 1) the output weights solved by Moore-Penrose generalized inverse is a least squares minimization issue, which easily suffers from overfitting and 2) the accuracy of ELM is drastically sensitive to the number of hidden neurons so that a large model is usually generated. This brief presents a sparse Bayesian approach for learning the output weights of ELM in classification. The new model, called Sparse Bayesian ELM (SBELM), can resolve these two drawbacks by estimating the marginal likelihood of network outputs and automatically pruning most of the redundant hidden neurons during learning phase, which results in an accurate and compact model. The proposed SBELM is evaluated on wide types of benchmark classification problems, which verifies that the accuracy of SBELM model is relatively insensitive to the number of hidden neurons; and hence a much more compact model is always produced as compared with other state-of-the-art neural network classifiers.
Jiahua Luo, Chi-Man Vong, Pak-Kin Wong 0001
IEEE Trans. Neural Networks Learn. Syst.2
2012 Modelling and prediction of automotive engine airratio using relevance vector machine
abstract
Fuel efficiency and pollution reduction relate closely to air-ratio (i.e. lambda) among all of the automotive engine control variables. Accurate lambda prediction is essential for effective lambda control. This paper presents an online sequential algorithm for relevance vector machine (RVM) to build a time-dependent RVM lambda function which can be continually updated whenever a sample is added to, or removed from, the training dataset. In order to evaluate the effectiveness of the online sequential algorithm, three lambda time series obtained from experiments under different engine operating conditions were employed. The prediction results under the online sequential algorithm over unseen cases were compared with those under decremental least-squares support vector machine. From the experiments, the online sequential RVM shows promising results and is superior to the typical online algorithm.
Pak-Kin Wong 0001, Hang-Cheong Wong, Chi-Man Vong
ICARCV3
2011 Case-based expert system using wavelet packet transform and kernel-based feature manipulation for engine ignition system diagnosis
Chi-Man Vong, Pak-Kin Wong 0001, Weng-Fai Ip
Eng. Appl. Artif. Intell.1
2011 Engine ignition signal diagnosis with Wavelet Packet Transform and Multi-class Least Squares Support Vector Machines
Chi-Man Vong, Pak-Kin Wong 0001
Expert Syst. Appl.1
2010 Case-based adaptation for automotive engine electronic control unit calibration
Chi-Man Vong, Pak-Kin Wong 0001
Expert Syst. Appl.1
2006 Prediction of automotive engine power and torque using least squares support vector machines and Bayesian inference
Chi-Man Vong, Pak-Kin Wong 0001
Eng. Appl. Artif. Intell.1