VLDB 2026 Research / reviewers in the wild / expert
Kezhi Mao
dblp:m/KezhiMao · also Ke Zhi Mao
· DBLP profile ↗
80ranked-venue papers
9as first author
30since 2021 · last 2025
0000-0002-9191-8604ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 4 first-author · 28 since 2021Databases, data management, data science and information retrieval · 12Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-authorSystems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HFD-Teacher: High-Frequency Depth Distillation From Depth Foundation Models for Enhanced Depth Completion
Anqi Cheng, Haiyue Zhu, Pey Yuen Tao, Kezhi Mao |
ICCV | 6 |
| 2025 | Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token PredictionabstractJunlang Qian, Zixiao Zhu, Hanzhang Zhou, Zijian Feng, Zepeng Zhai, Kezhi Mao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Junlang Qian, Zixiao Zhu, Hanzhang Zhou, Zijian Feng, Zepeng Zhai, Kezhi Mao |
NAACL (Long Papers) | 6 |
| 2025 | Logit Separability-Driven Samples and Multiple Class-Related Words Selection for Advancing In-Context LearningabstractZixiao Zhu, Zijian Feng, Hanzhang Zhou, Junlang Qian, Kezhi Mao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zixiao Zhu, Zijian Feng, Hanzhang Zhou, Junlang Qian, Kezhi Mao |
NAACL (Long Papers) | 5 |
| 2025 | Restoring Pruned Large Language Models via Lost Component CompensationabstractPruning is a widely used technique to reduce the size and inference cost of large language models (LLMs), but it often causes performance degradation. To mitigate this, existing restoration methods typically employ parameter-efficient fine-tuning (PEFT), such as LoRA, to recover the pruned model's performance. However, most PEFT methods are designed for dense models and overlook the distinct properties of pruned models, often resulting in suboptimal recovery. In this work, we propose a targeted restoration strategy for pruned models that restores performance while preserving their low cost and high efficiency. We observe that pruning-induced information loss is reflected in attention activations, and selectively reintroducing components of this information can significantly recover model performance. Based on this insight, we introduce RestoreLCC (Restoring Pruned LLMs via Lost Component Compensation), a plug-and-play method that contrastively probes critical attention heads via activation editing, extracts lost components from activation differences, and finally injects them back into the corresponding pruned heads for compensation and recovery. RestoreLCC is compatible with structured, semi-structured, and unstructured pruning schemes. Extensive experiments demonstrate that RestoreLCC consistently outperforms state-of-the-art baselines in both general and task-specific performance recovery, without compromising the sparsity or inference efficiency of pruned models. Zijian Feng, Hanzhang Zhou, Zixiao Zhu, Chua Jia Jim Deryl, Lee Onn Mak, Gee Wah Ng, Kezhi Mao |
NeurIPS | 8 |
| 2024 | FreeCtrl: Constructing Control Centers with Feedforward Layers for Learning-Free Controllable Text GenerationabstractControllable text generation (CTG) seeks to craft texts adhering to specific attributes, traditionally employing learning-based techniques such as training, fine-tuning, or prefix-tuning with attribute-specific datasets.These approaches, while effective, demand extensive computational and data resources.In contrast, some proposed learning-free alternatives circumvent learning but often yield inferior results, exemplifying the fundamental machine learning trade-off between computational expense and model efficacy.To overcome these limitations, we propose FreeCtrl, a learningfree approach that dynamically adjusts the weights of selected feedforward neural network (FFN) vectors to steer the outputs of large language models (LLMs).FreeCtrl hinges on the principle that the weights of different FFN vectors influence the likelihood of different tokens appearing in the output.By identifying and adaptively adjusting the weights of attributerelated FFN vectors, FreeCtrl can control the output likelihood of attribute keywords in the generated content.Extensive experiments on single-and multi-attribute control reveal that the learning-free FreeCtrl outperforms other learning-free and learning-based methods, successfully resolving the dilemma between learning costs and model performance 1 . Zijian Feng, Hanzhang Zhou, Kezhi Mao, Zixiao Zhu |
ACL (1) | 3 |
| 2024 | LLMs Learn Task Heuristics from Demonstrations: A Heuristic-Driven Prompting Strategy for Document-Level Event Argument ExtractionabstractIn this study, we explore in-context learning (ICL) in document-level event argument extraction (EAE) to alleviate the dependency on large-scale labeled data for this task.We introduce the Heuristic-Driven Link-of-Analogy (HD-LoA) prompting tailored for the EAE task.Specifically, we hypothesize and validate that LLMs learn task-specific heuristics from demonstrations in ICL.Building upon this hypothesis, we introduce an explicit heuristicdriven demonstration construction approach, which transforms the haphazard example selection process into a systematic method that emphasizes task heuristics.Additionally, inspired by the analogical reasoning of human, we propose the link-of-analogy prompting, which enables LLMs to process new situations by drawing analogies to known situations, enhancing their performance on unseen classes beyond limited ICL examples.Experiments show that our method outperforms existing prompting methods and few-shot supervised learning methods on document-level EAE datasets.Additionally, the HD-LoA prompting shows effectiveness in other tasks like sentiment analysis and natural language inference, demonstrating its broad adaptability 1 . Hanzhang Zhou, Junlang Qian, Zijian Feng, Zixiao Zhu, Kezhi Mao |
ACL (1) | 6 |
| 2024 | Unveiling and Manipulating Prompt Influence in Large Language ModelsabstractPrompts play a crucial role in guiding the responses of Large Language Models (LLMs). However, the intricate role of individual tokens in prompts, known as input saliency, in shaping the responses remains largely underexplored. Existing saliency methods either misalign with LLM generation objectives or rely heavily on linearity assumptions, leading to potential inaccuracies. To address this, we propose Token Distribution Dynamics (TDD), an elegantly simple yet remarkably effective approach to unveil and manipulate the role of prompts in generating LLM outputs. TDD leverages the robust interpreting capabilities of the language model head (LM head) to assess input saliency. It projects input tokens into the embedding space and then estimates their significance based on distribution dynamics over the vocabulary. We introduce three TDD variants: forward, backward, and bidirectional, each offering unique insights into token relevance. Extensive experiments reveal that the TDD surpasses state-of-the-art baselines with a big margin in elucidating the causal relationships between prompts and LLM outputs. Beyond mere interpretation, we apply TDD to two prompt manipulation tasks for controlled text generation: zero-shot toxic language suppression and sentiment steering. Empirical results underscore TDD's proficiency in identifying both toxic and sentimental cues in prompts, subsequently mitigating toxicity or modulating sentiment in the generated content. Zijian Feng, Hanzhang Zhou, Zixiao Zhu, Junlang Qian, Kezhi Mao |
ICLR | 5 |
| 2024 | GAM-Depth: Self-Supervised Indoor Depth Estimation Leveraging a Gradient-Aware Mask and Semantic ConstraintsabstractSelf-supervised depth estimation has evolved into an image reconstruction task that minimizes a photometric loss. While recent methods have made strides in indoor depth estimation, they often produce inconsistent depth estimation in textureless areas and unsatisfactory depth discrepancies at object boundaries. To address these issues, in this work, we propose GAM-Depth, developed upon two novel components: gradient-aware mask and semantic constraints. The gradient-aware mask enables adaptive and robust supervision for both key areas and textureless regions by allocating weights based on gradient magnitudes. The incorporation of semantic constraints for indoor self-supervised depth estimation improves depth discrepancies at object boundaries, leveraging a co-optimization network and proxy semantic labels derived from a pretrained segmentation model. Experimental studies on three indoor datasets, including NYUv2, ScanNet, and InteriorNet, show that GAM-Depth outperforms existing methods and achieves state-of-the-art performance, signifying a meaningful step forward in indoor depth estimation. Our code will be available at https://github.com/AnqiCheng1234/GAM-Depth. Anqi Cheng, Haiyue Zhu, Kezhi Mao |
ICRA | 4 |
| 2024 | Reinforced Cross-Domain Knowledge Distillation on Time Series DataabstractUnsupervised domain adaptation methods have demonstrated superior capabilities in handling the domain shift issue which widely exists in various time series tasks. However, their prominent adaptation performances heavily rely on complex model architectures, posing an unprecedented challenge in deploying them on resource-limited devices for real-time monitoring. Existing approaches, which integrates knowledge distillation into domain adaptation frameworks to simultaneously address domain shift and model complexity, often neglect network capacity gap between teacher and student and just coarsely align their outputs over all source and target samples, resulting in poor distillation efficiency. Thus, in this paper, we propose an innovative framework named Reinforced Cross-Domain Knowledge Distillation (RCD-KD) which can effectively adapt to student's network capability via dynamically selecting suitable target domain samples for knowledge transferring. Particularly, a reinforcement learning-based module with a novel reward function is proposed to learn optimal target sample selection policy based on student's capacity. Meanwhile, a domain discriminator is designed to transfer the domain invariant knowledge. Empirical experimental results and analyses on four public time series datasets demonstrate the effectiveness of our proposed method over other state-of-the-art benchmarks. Qing Xu 0015, Min Wu 0008, Xiaoli Li 0001, Kezhi Mao, Zhenghua Chen |
NeurIPS | 4 |
| 2024 | UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN ManipulationabstractLarge language models (LLMs) have demonstrated impressive capabilities in various tasks using the in-context learning (ICL) paradigm. However, their effectiveness is often compromised by inherent bias, leading to prompt brittleness—sensitivity to design settings such as example selection, order, and prompt formatting. Previous studies have addressed LLM bias through external adjustment of model outputs, but the internal mechanisms that lead to such bias remain unexplored. Our work delves into these mechanisms, particularly investigating how feedforward neural networks (FFNs) and attention heads result in the bias of LLMs. By Interpreting the contribution of individual FFN vectors and attention heads, we identify the biased LLM components that skew LLMs' prediction toward specific labels. To mitigate these biases, we introduce UniBias, an inference-only method that effectively identifies and eliminates biased FFN vectors and attention heads. Extensive experiments across 12 NLP datasets demonstrate that UniBias significantly enhances ICL performance and alleviates prompt brittleness of LLMs. Hanzhang Zhou, Zijian Feng, Zixiao Zhu, Junlang Qian, Kezhi Mao |
NeurIPS | 5 |
| 2024 | Explicit and implicit knowledge-enhanced model for event causality identification
Kezhi Mao |
Expert Syst. Appl. | 2 |
| 2024 | Adaptive micro- and macro-knowledge incorporation for hierarchical text classification
Zijian Feng, Kezhi Mao, Hanzhang Zhou |
Expert Syst. Appl. | 2 |
| 2024 | Self-Supervised Video Representation Learning by Video Incoherence DetectionabstractThis article introduces a novel self-supervised method that leverages incoherence detection for video representation learning. It stems from the observation that the visual system of human beings can easily identify video incoherence based on their comprehensive understanding of videos. Specifically, we construct the incoherent clip by multiple subclips hierarchically sampled from the same raw video with various lengths of incoherence. The network is trained to learn the high-level representation by predicting the location and length of incoherence given the incoherent clip as input. Additionally, we introduce intravideo contrastive learning to maximize the mutual information between incoherent clips from the same raw video. We evaluate our proposed method through extensive experiments on action recognition and video retrieval using various backbone networks. Experiments show that our proposed method achieves remarkable performance across different backbone networks and different datasets compared to previous coherence-based methods. Haozhi Cao, Yuecong Xu, Kezhi Mao, Lihua Xie 0001, Jianxiong Yin, Simon See, Qianwen Xu 0001, Jianfei Yang 0001 |
IEEE Trans. Cybern. | 3 |
| 2024 | Aligning Correlation Information for Domain Adaptation in Action RecognitionabstractDomain adaptation (DA) approaches address domain shift and enable networks to be applied to different scenarios. Although various image DA approaches have been proposed in recent years, there is limited research toward video DA. This is partly due to the complexity in adapting the different modalities of features in videos, which includes the correlation features extracted as long-range dependencies of pixels across spatiotemporal dimensions. The correlation features are highly associated with action classes and proven their effectiveness in accurate video feature extraction through the supervised action recognition task. Yet correlation features of the same action would differ across domains due to domain shift. Therefore, we propose a novel adversarial correlation adaptation network (ACAN) to align action videos by aligning pixel correlations. ACAN aims to minimize the distribution of correlation information, termed as pixel correlation discrepancy (PCD). Additionally, video DA research is also limited by the lack of cross-domain video datasets with larger domain shifts. We, therefore, introduce a novel HMDB-ARID dataset with a larger domain shift caused by a larger statistical difference between domains. This dataset is built in an effort to leverage current datasets for dark video classification. Empirical results demonstrate the state-of-the-art performance of our proposed ACAN for both existing and the new video DA datasets. Yuecong Xu, Haozhi Cao, Kezhi Mao, Zhenghua Chen, Lihua Xie 0001, Jianfei Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Distilling Universal and Joint Knowledge for Cross-Domain Model Compression on Time Series DataabstractFor many real-world time series tasks, the computational complexity of prevalent deep leaning models often hinders the deployment on resource limited environments (e.g., smartphones). Moreover, due to the inevitable domain shift between model training (source) and deploying (target) stages, compressing those deep models under cross-domain scenarios becomes more challenging. Although some of existing works have already explored cross-domain knowledge distillation for model compression, they are either biased to source data or heavily tangled between source and target data. To this end, we design a novel end-to-end framework called UNiversal and joInt Knowledge Distillation (UNI-KD) for cross-domain model compression. In particular, we propose to transfer both the universal feature-level knowledge across source and target domains and the joint logit-level knowledge shared by both domains from the teacher to the student model via an adversarial learning scheme. More specifically, a feature-domain discriminator is employed to align teacher’s and student’s representations for universal knowledge transfer. A data-domain discriminator is utilized to prioritize the domain-shared samples for joint knowledge transfer. Extensive experimental results on four time series datasets demonstrate the superiority of our proposed method over state-of-the-art (SOTA) benchmarks. The source code is available at https://github.com/ijcai2023/UNI KD. Qing Xu 0015, Min Wu 0008, Xiaoli Li 0001, Kezhi Mao, Zhenghua Chen |
IJCAI | 4 |
| 2023 | A graph attention network utilizing multi-granular information for emotion-cause pair extraction
Kezhi Mao |
Neurocomputing | 2 |
| 2023 | Feature-aware conditional GAN for category text generation
Kezhi Mao, Fanfan Lin, Zijian Feng |
Neurocomputing | 2 |
| 2023 | Knowledge-based BERT word embedding fine-tuning for emotion recognition
Zixiao Zhu, Kezhi Mao |
Neurocomputing | 2 |
| 2022 | A Comparative Study on Machine Learning algorithms for Knowledge DiscoveryabstractFor centuries, the process of formulating new knowledge from observations has driven scientific discoveries. With rapid advancements in machine learning, it is natural to question the possibility of automating knowledge discovery in the scientific field. A benchmark task for automated knowledge discovery is called symbolic regression. The task aims to predict a mathematical equation that best describes the observational data. The advancements in symbolic regression have significant potential to aid research in understanding unexplored systems' dynamics and governing properties. However, the combinatorial nature of the problem makes it an expensive and challenging problem to solve efficiently. Several types of symbolic regression algorithms exist, from genetic programming and sparse regression to deep generative models. However, no survey collates these prominent algorithms. Therefore, this paper aims to summarize key research works in symbolic regression and perform a comparative study to understand the strength and limitations of each method. Finally, we highlight the challenges in the current methods and future research directions in the application of machine learning in knowledge discovery. Siddesh Sambasivam Suseela, Feng Yang 0011, Kezhi Mao |
ICARCV | 3 |
| 2022 | Calibrating Class Weights with Multi-Modal Information for Partial Video Domain AdaptationabstractAssuming the source label space subsumes the target one, Partial Video Domain Adaptation (PVDA) is a more general and practical scenario for cross-domain video classification problems. The key challenge of PVDA is to mitigate the negative transfer caused by the source-only outlier classes. To tackle this challenge, a crucial step is to aggregate target predictions to assign class weights by up-weighing target classes and down-weighing outlier classes. However, the incorrect predictions of class weights can mislead the network and lead to negative transfer. Previous works improve the class weight accuracy by utilizing temporal features and attention mechanisms, but these methods may fall short when trying to generate accurate class weight when domain shifts are significant, as in most real-world scenarios. To deal with these challenges, we first propose the Multi-modality partial Adversarial Network (MAN), which utilizes multi-scale and multi-modal information to enhance PVDA performance. Based on MAN, we then propose Multi-modality Cluster-calibrated partial Adversarial Network (MCAN). It utilizes a novel class weight calibration method to alleviate the negative transfer caused by incorrect class weights. Specifically, the calibration method tries to identify and weigh correct and incorrect predictions using distributional information implied by unsupervised clustering. Extensive experiments are conducted on prevailing PVDA benchmarks, and the proposed MCAN achieves significant improvements when compared to state-of-the-art PVDA methods. Yuecong Xu, Jianfei Yang 0001, Kezhi Mao |
ACM Multimedia | 4 |
| 2022 | Document-Level Event Argument Extraction by Leveraging Redundant Information and Closed Boundary LossabstractIn document-level event argument extraction, an argument is likely to appear multiple times in different expressions in the document.The redundancy of arguments underlying multiple sentences is beneficial but is often overlooked.In addition, in event argument extraction, most entities are regarded as class "others", i.e.Universum class, which is defined as a collection of samples that do not belong to any class of interest.Universum class is composed of heterogeneous entities without typical common features.Classifiers trained by cross entropy loss could easily misclassify the Universum class because of their open decision boundary.In this paper, to make use of redundant event information underlying a document, we build an entity coreference graph with the graph2token module to produce a comprehensive and coreference-aware representation for every entity and then build an entity summary graph to merge the multiple extraction results.To better classify Universum class, we propose a new loss function to build classifiers with closed boundaries.Experimental results show that our model outperforms the previous state-of-the-art models by 3.35% in F1-score. Hanzhang Zhou, Kezhi Mao |
NAACL-HLT | 2 |
| 2022 | Tailored text augmentation for sentiment analysis
Zijian Feng, Hanzhang Zhou, Zixiao Zhu, Kezhi Mao |
Expert Syst. Appl. | 4 |
| 2022 | A novel end-to-end neural network for simultaneous filtering of task-unrelated named entities and fine-grained typing of task-related named entities
Kezhi Mao, Yuecong Xu, Edmond Yat-Man Lo |
Expert Syst. Appl. | 2 |
| 2022 | RF-HoDRF: High-Order Hybrid Discriminative Random Field Improved by Two-Layer Random Forest for SAR Image Change DetectionabstractFor better exploiting discriminative texture features and encoding high-level structures, this letter presents a high- order hybrid discriminative random field improved by two-layer random forest, abbreviated as RF-HoDRF, for synthetic aperture radar (SAR) image change detection. First, RF-HoDRF constructs a two-layer random forest (TL-RF) model to realize the selection of high-dimensional texture features, and then provides the class probabilities for constructing the unary potential in RF-HoDRF. Second, it defines a high-order potential on high-order cliques generated by superpixels to encode the high-level structures and maintain the region consistency. Finally, considering the pairwise potential by improved generalized Ising model and the statistics by generalized Gamma distribution (GΓD), the RF-HoDRF model is derived under the discriminative model framework. Then, by iteratively maximizing the local posterior probabilities, the class labels and the parameters are optimally estimated until they converge. Extensive comparisons and ablation experiments on measured SAR images verify the effectiveness of our method, and demonstrate that discriminative features selection and high-order structures maintenance have great contributions to improving change detection performances. Wanying Song, Yan Wu 0003, Peng Zhang 0003, Kezhi Mao |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2021 | ACT: an Attentive Convolutional Transformer for Efficient Text ClassificationabstractRecently, Transformer has been demonstrating promising performance in many NLP tasks and showing a trend of replacing Recurrent Neural Network (RNN). Meanwhile, less attention is drawn to Convolutional Neural Network (CNN) due to its weak ability in capturing sequential and long-distance dependencies, although it has excellent local feature extraction capability. In this paper, we introduce an Attentive Convolutional Transformer (ACT) that takes the advantages of both Transformer and CNN for efficient text classification. Specifically, we propose a novel attentive convolution mechanism that utilizes the semantic meaning of convolutional filters attentively to transform text from complex word space to a more informative convolutional filter space where important n-grams are captured. ACT is able to capture both local and global dependencies effectively while preserving sequential information. Experiments on various text classification tasks and detailed analyses show that ACT is a lightweight, fast, and effective universal text classifier, outperforming CNNs, RNNs, and attentive models including Transformer. Peixiang Zhong, Kezhi Mao, Dongzhe Wang, Xuefeng Yang, Jianxiong Yin, Simon See |
AAAI | 3 |
| 2021 | Partial Video Domain Adaptation with Partial Adversarial Temporal Attentive NetworkabstractPartial Domain Adaptation (PDA) is a practical and general domain adaptation scenario, which relaxes the fully shared label space assumption such that the source label space subsumes the target one. The key challenge of PDA is the issue of negative transfer caused by source-only classes. For videos, such negative transfer could be triggered by both spatial and temporal features, which leads to a more challenging Partial Video Domain Adaptation (PVDA) problem. In this paper, we propose a novel Partial Adversarial Temporal Attentive Network (PATAN) to address the PVDA problem by utilizing both spatial and temporal features for filtering source-only classes. Besides, PATAN constructs effective overall temporal features by attending to local temporal features that contribute more toward the class filtration process. We further introduce new benchmarks to facilitate research on PVDA problems, covering a wide range of PVDA scenarios. Empirical results demonstrate the state-of-the-art performance of our proposed PATAN across the multiple PVDA benchmarks. Code will be provided at: https://github.com/xuyu0010/PATAN. Yuecong Xu, Jianfei Yang 0001, Haozhi Cao, Zhenghua Chen, Kezhi Mao |
ICCV | 6 |
| 2021 | Exploiting inter-frame regional correlation for efficient action recognition
Yuecong Xu, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See |
Expert Syst. Appl. | 3 |
| 2021 | Particle swarm optimization with state-based adaptive velocity limit strategy
Kezhi Mao, Fanfan Lin, Xin Zhang 0034 |
Neurocomputing | 2 |
| 2021 | PNL: Efficient long-range dependencies extraction with pyramid non-local module for action recognition
Yuecong Xu, Haozhi Cao, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See |
Neurocomputing | 4 |
| 2021 | Effective action recognition with embedded key point shifts
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See |
Pattern Recognit. | 4 |
| 2020 | Convolutional Transformer with Sentiment-aware Attention for Sentiment AnalysisabstractGiven certain data available for training, the keys to improving a sentiment analysis system lie in developing a good model that is capable of capturing both local and global features of texts, as well as incorporating external knowledge into the model effectively. In this paper, we propose a multi-window Convolutional Transformer (ConvTransformer) that takes the advantages of both Transformer and CNN for sentiment analysis. The proposed ConvTransformer is able to capture important local n-gram features effectively while preserving sequential information of texts. Furthermore, we propose a sentiment-aware attention mechanism to incorporate the sentiment intensity information of each word by utilizing an external knowledge base, SentiWordNet. The sentiment-aware attention mechanism takes both sentiment and position information of each token into consideration when computing attention weights, resulting in a global feature for final classification. Comparing with CNN, RNN and attention-based baseline models, our model achieves the best performance on multiple sentiment analysis datasets. Peixiang Zhong, Jiaheng Zhang, Kezhi Mao |
IJCNN | 4 |
| 2020 | Improving convolutional neural network for text classification by recursive data pruning
Kezhi Mao, Edmond Yat-Man Lo |
Neurocomputing | 3 |
| 2020 | Bag-of-Concepts representation for document classification based on automatic knowledge acquisition from probabilistic knowledge base
Kezhi Mao, Yuecong Xu, Jiaheng Zhang |
Knowl. Based Syst. | 2 |
| 2019 | Improving Relation Extraction with Knowledge-attentionabstractPengfei Li, Kezhi Mao, Xuefeng Yang, Qi Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Kezhi Mao, Xuefeng Yang |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Knowledge-oriented convolutional neural network for causal relation extraction from natural language textsabstractCausal relation extraction is a challenging yet very important task for Natural Language Processing (NLP). There are many existing approaches developed to tackle this task, either rule-based (non-statistical) or machine-learning-based (statistical) method. For rule-based method, extensive manual work is required to construct handcrafted patterns, however, the precision and recall are low due to the complexity of causal relation expressions in natural language. For machine-learning-based method, current approaches either rely on sophisticated feature engineering which is error-prone, or rely on large amount of labeled data which is impractical for causal relation extraction problem. To address the above issues, we propose a Knowledge-oriented Convolutional Neural Network (K-CNN) for causal relation extraction in this paper. K-CNN consists of a knowledge-oriented channel that incorporates human prior knowledge to capture the linguistic clues of causal relationship , and a data-oriented channel that learns other important features of causal relation from the data. The convolutional filters in knowledge-oriented channel are automatically generated from lexical knowledge bases such as WordNet and FrameNet. We propose filter selection and clustering techniques to reduce dimensionality and improve the performance of K-CNN. Furthermore, additional semantic features that are useful for identifying causal relations are created. Three datasets have been used to evaluate the ability of K-CNN to effectively extract causal relation from texts, and the model outperforms current state-of-art models for relation extraction. Kezhi Mao |
Expert Syst. Appl. | 2 |
| 2019 | Task-generic semantic convolutional neural network for web text-aided image classification
Dongzhe Wang, Kezhi Mao |
Neurocomputing | 2 |
| 2019 | Enhanced feature fusion through irrelevant redundancy elimination in intra-class and extra-class discriminative correlation analysis
Zuobin Wu, Kezhi Mao, Gee Wah Ng |
Neurocomputing | 2 |
| 2019 | Semantic-filtered Soft-Split-Aware video captioning with audio-augmented feature
Yuecong Xu, Jianfei Yang 0001, Kezhi Mao |
Neurocomputing | 3 |
| 2019 | Learning Semantic Text Features for Web Text-Aided Image ClassificationabstractThe good generalization performance of conventional pattern classifiers often relies on the size of training data labeled by costly human labor. These days, publicly available web resources grow explosively, and this allows us to easily obtain abundant and cheap web data. Yet, web data are usually not as cooperative as human labeled data. In this paper, we explore the use of web text data to aid image classification. Without requiring the previous collection of auxiliary data from the web, we directly retrieve the web text information with the aid of the powerful reverse image search engine. We develop a novel textual modeling method namedsemantic matching neural network(SMNN) that is capable of learning semantic features from the associated text of web images. The SMNN text features have improved reliability and applicability, compared to the text features obtained from other methods. The SMNN text features and convolutional neural network (CNN) visual features are merged into a shared representation, which learns to capture the correlations between the two modalities. Experimental results on benchmark UIUC-Sports, Scene-15, Caltech-256, and Pascal VOC-2012 data sets show that the visual and text modalities of data from different sources are remarkably complementary and the fusion of them achieves substantial performance improvement. Dongzhe Wang, Kezhi Mao |
IEEE Trans. Multim. | 2 |
| 2018 | Feature Regrouping for CCA - Based Feature Fusion and Extraction Through Normalized CutabstractFeature fusion is important for providing enhancements of data authenticity in both traditional and deep learning pattern analysis. Classical serial fusion concatenates multiple feature sets followed by dimensionality reduction using principal component analysis (PCA), linear discriminant analysis (LDA), canonical correlation analysis (CCA) etc. CCA-based feature fusion is a main technique for exploring the mutual relationships of multiple feature sets. It considers the correlation of multiple feature sets during dimensionality reduction. In traditional CCA-based feature fusion and extraction, the natural groupings of features are directly used. It is still unclear whether the natural groupings of features are optimal for CCA-based fusion. In this paper, we propose a feature regrouping algorithm for CCA-based feature fusion and extraction through normalized cut (FR-NC). Feature correlation analysis is incorporated into normalized cut, in which the intra-group correlation is maximized, and the extra-group correlation is minimized simultaneously. CCA-based feature fusion is performed on the regrouped features. The proposed feature regrouping algorithm aims to provide enhanced fused features for pattern classification. Extensive experiments have proved its effectiveness. Zuobin Wu, Kezhi Mao, Gee Wah Ng |
FUSION | 2 |
| 2018 | Fuzzy Bag-of-Words Model for Document RepresentationabstractOne key issue in text mining and natural language processing is how to effectively represent documents using numerical vectors. One classical model is the Bag-of-Words (BoW). In a BoW-based vector representation of a document, each element denotes the normalized number of occurrence of a basis term in the document. To count the number of occurrence of a basis term, BoW conducts exact word matching, which can be regarded as a hard mapping from words to the basis term. BoW representation suffers from its intrinsic extreme sparsity, high dimensionality, and inability to capture high-level semantic meanings behind text data. To address the aforementioned issues, we propose a new document representation method named fuzzy Bag-of-Words (FBoW) in this paper. FBoW adopts a fuzzy mapping based on semantic correlation among words quantified by cosine similarity measures between word embeddings. Since word semantic matching instead of exact word string matching is used, the FBoW could encode more semantics into the numerical representation. In addition, we propose to use word clusters instead of individual words as basis terms and develop fuzzy Bag-of-WordClusters (FBoWC) models. Three variants under the framework of FBoWC are proposed based on three different similarity measures between word clusters and words, which are named as FBoWCmean, FBoWCmax, and FBoWCmin , respectively. Document representations learned by the proposed FBoW and FBoWC are dense and able to encode high-level semantics. The task of document categorization is used to evaluate the performance of learned representation by the proposed FBoW and FBoWC methods. The results on seven real-word document classification datasets in comparison with six document representation learning methods have shown that our methods FBoW and FBoWC achieve the highest classification accuracies. Rui Zhao 0004, Kezhi Mao |
IEEE Trans. Fuzzy Syst. | 2 |
| 2017 | Convolutional neural networks and multimodal fusion for text aided image classificationabstractWith the exponential growth of web meta-data, exploiting multimodal online sources via standard search engine has become a trend in visual recognition as it effectively alleviates the shortage of training data. However, the web meta-data such as text data is usually not as cooperative as expected due to its unstructured nature. To address this problem, this paper investigates the numerical representation of web text data. We firstly adopt convolutional neural network (CNN) for web text modeling on top of word vectors. Combined with CNN for image, we present a multimodal fusion to maximize the discriminative power of visual and textual modality data for decision level and feature level simultaneously. Experimental results show that the proposed framework achieves significant improvement in large-scale image classification on Pascal VOC-2007 and VOC-2012 datasets. Dongzhe Wang, Kezhi Mao, Gee Wah Ng |
FUSION | 2 |
| 2017 | Effective feature fusion for pattern classification based on intra-class and extra-class discriminative correlation analysisabstractInformation fusion aims to exploit truthful knowledge from various sources in a reliable and accurate way. Fusion of information can be conducted at three abstraction levels including feature level, score level and decision level. The feature fusion approaches have the advantages of preserving effective discriminative structure underlying various features. In this paper, we propose an effective feature fusion algorithm based on intra-class and extra-class discriminative correlation analysis (IEDCA), aiming to eliminate between-class correlation and retain enough feature dimension for correlation analysis. IEDCA explores the intra-class correlation including both the pairs-wise correlation like CCA-based feature fusion approaches and the correlation across different features within the same class. Our proposed method can be used in unimodal feature fusion as well as multimodal feature fusion, and extensive experiments have proved its effectiveness. Zuobin Wu, Kezhi Mao, Gee Wah Ng |
FUSION | 2 |
| 2017 | Cyberbullying Detection Based on Semantic-Enhanced Marginalized Denoising Auto-EncoderabstractAs a side effect of increasingly popular social media, cyberbullying has emerged as a serious problem afflicting children, adolescents and young adults. Machine learning techniques make automatic detection of bullying messages in social media possible, and this could help to construct a healthy and safe social media environment. In this meaningful research area, one critical issue is robust and discriminative numerical representation learning of text messages. In this paper, we propose a new representation learning method to tackle this problem. Our method named semantic-enhanced marginalized denoising auto-encoder (smSDA) is developed via semantic extension of the popular deep learning model stacked denoising autoencoder (SDA). The semantic extension consists of semantic dropout noise and sparsity constraints, where the semantic dropout noise is designed based on domain knowledge and the word embedding technique. Our proposed method is able to exploit the hidden feature structure of bullying information and learn a robust and discriminative representation of text. Comprehensive experiments on two public cyberbullying corpora (Twitter and MySpace) are conducted, and the results show that our proposed approaches outperform other baseline text representation learning methods. Rui Zhao 0004, Kezhi Mao |
IEEE Trans. Affect. Comput. | 2 |
| 2017 | Task Independent Fine Tuning for Word EmbeddingsabstractRepresentation learning of words, also known as word embedding technique, is based on the distributional hypothesis that words with similar semantic meanings have similar context. The selection of context window naturally has an influence on word vectors learned. However, it is found that the word vectors are often very sensitive to the defined context window, and unfortunately there is no unified optimal context window for all words. One impact of this issues is that, under a predefined context window, the semantic meanings of some words may not be well represented by the learned vectors. To alleviate the problem and improve word embeddings, we propose a task-independent fine-tuning framework in this paper. The main idea of the task-independent fine tuning is to integrate multiple word embeddings and lexical semantic resources to fine tune a target word embedding. The effectiveness of the proposed framework is tested by tasks of semantic similarity prediction, analogical reasoning, and sentence completion. Experiments results on six word embeddings and eight datasets show that the proposed fine-tuning framework could significantly improve word embeddings. Xuefeng Yang, Kezhi Mao |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Topic-Aware Deep Compositional Models for Sentence ClassificationabstractIn recent years, deep compositional models have emerged as a popular technique for representation learning of sentence in computational linguistic and natural language processing. These models normally train various forms of neural networks on top of pretrained word embeddings using a task-specific corpus. However, most of these works neglect the multisense nature of words in the pretrained word embeddings. In this paper we introduce topic models to enrich the word embeddings for multisenses of words. The integration of the topic model with various semantic compositional processes leads to topic-aware convolutional neural network and topic-aware long short term memory networks. Different from previous multisense word embeddings models that assign multiple independent and sense-specific embeddings to each word, our proposed models are lightweight and have flexible frameworks that regard word sense as the composition of two parts: a general sense derived from a large corpus and a topic-specific sense derived from a task-specific corpus. In addition, our proposed models focus on semantic composition instead of word understanding. With the help of topic models, we can integrate the topic-specific sense at word-level before the composition and sentence-level after the composition. Comprehensive experiments on five public sentence classification datasets are conducted and the results show that our proposed topic-aware deep compositional models produce competitive or better performance than other text representation learning methods. Rui Zhao 0004, Kezhi Mao |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Event-Based Hough Transform in a Spiking Neural Network for Multiple Line Detection and Tracking Using a Dynamic Vision Sensor
Sajjad Seifozzakerini, Weiyun Yau, Kezhi Mao |
BMVC | 4 |
| 2016 | Adaptive multimodal fusion with web resources for scene classification
Dongzhe Wang, Kezhi Mao, Gee Wah Ng, Tien Pham |
FUSION | 2 |
| 2016 | Constructing Bayesian networks by harvesting knowledge from online resources
Zhibo Xiao, Tharini Nayanika de Silva, Kezhi Mao, Gee Wah Ng |
FUSION | 4 |
| 2016 | Learning multi-prototype word embedding from single-prototype word embedding with integrated knowledge
Xuefeng Yang, Kezhi Mao |
Expert Syst. Appl. | 2 |
| 2015 | Improving scene classification by fusion of training data and web resources
Dongzhe Wang, Kezhi Mao, Gee Wah Ng |
FUSION | 2 |
| 2015 | Emphasizing Minority Class in LDA for Feature Subset Selection on High-Dimensional Small-Sized ProblemsabstractAlthough mostly used for pattern classification, linear discriminant analysis (LDA) can also be used in feature selection as an effective measure to evaluate the separative ability of a feature subset. When applied to feature selection on high-dimensional small-sized (HDSS) data (generally) with class-imbalance, LDA encounters four problems, including singularity of scatter matrix, overfitting, overwhelming and prohibitively computational complexity. In this study, we propose the LDA-based feature selection method minority class emphasized linear discriminant analysis (MCE-LDA) with a new regularization technique to address the first three problems. Different to giving equal or more emphasis to majority class in conventional forms of regularization, the proposed regularization emphasizes more on minority class, with the expectation of improving overall performance by alleviating overwhelming of majority class to minority class as well as overfitting in minority class. In order to reduce computational overhead, an incremental implementation of LDA-based feature selection has been introduced. Comparative studies with other forms of regularization to LDA as well as with other popular feature selection methods on five HDSS problems show that MCE-LDA can produce feature subsets with excellent performance in both classification and robustness. Further experimental results of true positive rate (TPR) and true negative rate (TNR) have also verified the effectiveness of the proposed technique in alleviating overwhelming and overfitting problems. Feng Yang 0011, Kezhi Mao, Gary Kee Khoon Lee, Wenyin Tang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | Semantic-level fusion of heterogenous sensor network and other sources based on Bayesian network
Kui Wu 0002, Wenyin Tang, Kezhi Mao, Gee Wah Ng, Lee Onn Mak |
FUSION | 3 |
| 2014 | Multi level causal relation identification using extended features
Xuefeng Yang, Kezhi Mao |
Expert Syst. Appl. | 2 |
| 2013 | Model-Based Online Learning With KernelsabstractNew optimization models and algorithms for online learning with Kernels (OLK) in classification, regression, and novelty detection are proposed in a reproducing Kernel Hilbert space. Unlike the stochastic gradient descent algorithm, called the naive online Reg minimization algorithm (NORMA), OLK algorithms are obtained by solving a constrained optimization problem based on the proposed models. By exploiting the techniques of the Lagrange dual problem like Vapnik's support vector machine (SVM), the solution of the optimization problem can be obtained iteratively and the iteration process is similar to that of the NORMA. This further strengthens the foundation of OLK and enriches the research area of SVM. We also apply the obtained OLK algorithms to problems in classification, regression, and novelty detection, including real time background substraction, to show their effectiveness. It is illustrated that, based on the experimental results of both classification and regression, the accuracy of OLK algorithms is comparable with traditional SVM-based algorithms, such as SVM and least square SVM (LS-SVM), and with the state-of-the-art algorithms, such as Kernel recursive least square (KRLS) method and projectron method, while it is slightly higher than that of NORMA. On the other hand, the computational cost of the OLK algorithm is comparable with or slightly lower than existing online methods, such as above mentioned NORMA, KRLS, and projectron methods, but much lower than that of SVM-based algorithms. In addition, different from SVM and LS-SVM, it is possible for OLK algorithms to be applied to non-stationary problems. Also, the applicability of OLK in novelty detection is illustrated by simulation results. Guoqi Li 0002, Changyun Wen, Zhengguo Li, Feng Yang 0011, Kezhi Mao |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2012 | A cognitively inspired rule-plus-exemplar framework for interpretable pattern classification
Wing Yee Sit, Kezhi Mao |
FUSION | 2 |
| 2012 | Adaptive Fuzzy Rule-Based Classification System Integrating Both Expert Knowledge and DataabstractThis paper presents an adaptive fuzzy rule-based classification system using a new hybrid modeling method that integrates both expert knowledge and new knowledge learnt from data. Inspired by human learning, the membership functions of fuzzy rules are optimized based on a hybrid error function that combines errors caused by the class predefined by expert knowledge and nearby historical data. The weights of the two errors can be adjusted by a conservative parameter. Experimental results show that our method significantly reduces classification ambiguity in 9 datasets. Wenyin Tang, Kezhi Mao, Lee Onn Mak, Gee Wah Ng |
ICTAI | 2 |
| 2011 | Regularized linear discriminant analysis and its recursive implementation for gene subset selectionabstractAlthough mostly used for pattern classification, linear discriminant analysis (LDA) may also be used for feature selection. When employed to select genes for microarray data, which has high dimensionality and small sample size, LDA encounters three problems, including singularity of scatter matrix, overfitting and prohibitive computational complexity. In this study, we propose a new regularization technique to address the singularity and overfitting problem. In addition, we develop a recursive implementation for LDA to reduce computational overhead. Experimental studies on 5 gene microarray problems show that the regularized linear discriminant analysis (RLDA) and its recursive implementation produce gene subsets with excellent classification performance. Kezhi Mao, Feng Yang 0011, Wenyin Tang |
CIBCB | 1 |
| 2011 | Target classification using knowledge-based probabilistic model
Wenyin Tang, Kezhi Mao, Lee Onn Mak, Gee Wah Ng, Zhaoyang Sun, Ji Hua Ang, Godfrey Lim |
FUSION | 2 |
| 2011 | Recursive Mahalanobis Separability Measure for Gene Subset SelectionabstractMahalanobis class separability measure provides an effective evaluation of the discriminative power of a feature subset, and is widely used in feature selection. However, this measure is computationally intensive or even prohibitive when it is applied to gene expression data. In this study, a recursive approach to Mahalanobis measure evaluation is proposed, with the goal of reducing computational overhead. Instead of evaluating Mahalanobis measure directly in high-dimensional space, the recursive approach evaluates the measure through successive evaluations in 2D space. Because of its recursive nature, this approach is extremely efficient when it is combined with a forward search procedure. In addition, it is noted that gene subsets selected by Mahalanobis measure tend to overfit training data and generalize unsatisfactorily on unseen test data, due to small sample size in gene expression problems. To alleviate the overfitting problem, a regularized recursive Mahalanobis measure is proposed in this study, and guidelines on determination of regularization parameters are provided. Experimental studies on five gene expression problems show that the regularized recursive Mahalanobis measure substantially outperforms the nonregularized Mahalanobis measures and the benchmark recursive feature elimination (RFE) algorithm in all five problems. Kezhi Mao, Wenyin Tang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | Robust Feature Selection for Microarray Data Based on Multicriterion FusionabstractFeature selection often aims to select a compact feature subset to build a pattern classifier with reduced complexity, so as to achieve improved classification performance. From the perspective of pattern analysis, producing stable or robust solution is also a desired property of a feature selection algorithm. However, the issue of robustness is often overlooked in feature selection. In this study, we analyze the robustness issue existing in feature selection for high-dimensional and small-sized gene-expression data, and propose to improve robustness of feature selection algorithm by using multiple feature selection evaluation criteria. Based on this idea, a multicriterion fusion-based recursive feature elimination (MCF-RFE) algorithm is developed with the goal of improving both classification performance and stability of feature selection results. Experimental studies on five gene-expression data sets show that the MCF-RFE algorithm outperforms the commonly used benchmark feature selection algorithm SVM-RFE. Feng Yang 0011, Kezhi Mao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2010 | Improving robustness of gene ranking by resampling and permutation based score correction and normalizationabstractFeature ranking, which ranks features via their individual importance, is one of the frequently used feature selection techniques. Traditional feature ranking criteria are apt to produce inconsistent ranking results even with light perturbations in training samples when applied to high dimensional and small-sized gene expression data. A widely used strategy for solving the inconsistencies is the multi-criterion combination. But one problem encountered in combining multiple criteria is the score normalization. In this paper, problems in existing methods are first analyzed, and a new gene importance transformation algorithm is then proposed. Experimental studies on three popular gene expression datasets show that the multi-criterion combination based on the proposed score correction and normalization produces gene rankings with improved robustness. Feng Yang 0011, Kezhi Mao |
BIBM | 2 |
| 2010 | Classification for overlapping classes using optimized overlapping region detection and soft decision
Wenyin Tang, Kezhi Mao, Lee Onn Mak, Gee Wah Ng |
FUSION | 2 |
| 2010 | Designing compact Gabor filter banks for efficient texture feature extractionabstractTexture feature has been widely used in image segmentation, classification, retrieval and many others. Among various approaches to texture feature extraction, Gabor filtering has emerged as one of the most popular in recent years. Gabor filter-based texture feature extractor is in fact a Gabor filter bank defined by its parameters including frequencies, orientations and smoothing parameters of the Gaussian envelope. In the literature, these parameters are often set by trial and error, based on the experience of the user, and the Gabor filter banks thus designed are often over-sized. To address the problem mentioned above, we propose to design compact Gabor filter banks by incorporating filter selection in this study. We develop a new Mahalanobis separability measure-based supervised approach to address the need of texture feature extraction. The strengths of our methods are twofold. Firstly, the proposed method provides a systematic way for Gabor filter bank design to avoid man-made bias. Secondly, the compact filter banks thus designed overcomes the problem of redundant or insignificant/irrelevant filter banks, and this in turn leads to improved performance of texture classification. Experimental results on benchmark datasets demonstrate the effectiveness of our proposed approach. Kezhi Mao, Hong Zhang 0013, Tianyou Chai |
ICARCV | 2 |
| 2010 | Selection of Gabor filters for improved texture feature extractionabstractTexture feature has been widely used in object recognition, image content analysis and many others. Among various approaches to texture feature extraction, Gabor filter has emerged as one of the most popular ones. Gabor filter-based feature extractor is in fact a Gabor filter bank defined by its parameters including frequencies, orientations and smooth parameters of Gaussian envelope. In the literature, different parameter settings have been suggested, and filter banks created by these parameter settings work well in general. From the perspective of pattern classification, however, filter banks thus designed may not be ideal. In the present study, we propose a new approach to Gabor filter bank design, by incorporating feature selection, i.e. filter selection, into the design process. The merits of incorporating filter selection in filter bank design are twofold. Firstly, filter selection produces a compact Gabor filter bank and hence reduces computational complexity of texture feature extraction. Secondly, Gabor filter bank thus designed produces low-dimensional feature representation with improved sample-to-feature ratio, and this in turn leads to improved performance of texture classification. Experiment results on benchmark datasets and a real application have demonstrated the effectiveness of the proposed method. Kezhi Mao, Hong Zhang 0013, Tianyou Chai |
ICIP | 2 |
| 2008 | Fast Gene Selection for Microarray Data Using SVM-Based Evaluation CriterionabstractAn important application of microarrays is to identify the relevant genes, among thousands of genes, for phenotypic classification. The performance of a gene selection algorithm is often assessed in terms of both predictive capacity and computational efficiency, but predictive capacity of selected features receives more attention than does computational efficiency. However, in gene selection problems, the computational efficiency is equally important because of very high dimensionality of gene expression data. We propose an SVM-IRFS algorithm which combines Support Vector Machine (SVM) based criterion, generalized parwpar2measure, with a new search procedure, named as Iterative Reduced Forward Selection (IRFS), to address the gene selection problem. In the IRFS, an adaptive threshold is used to screen the irrelevant feature subsets, thus unnecessary computations can be avoided. The advantage of our proposed SVM-IRFS algorithm is twofold. First, the selection procedure of SVM-IRFS algorithm is computationally very efficient. It can identify tens from thousands of genes in several seconds. Second, benefiting from the good classification performance of support vector machines, SVM-IRFS produces the feature subset with high predictive capacity. Kezhi Mao, David P. Tuck |
BIBM | 3 |
| 2007 | Feature selection algorithm for mixed data with both nominal and continuous features
Wenyin Tang, Kezhi Mao |
Pattern Recognit. Lett. | 2 |
| 2006 | The ties problem resulting from counting-based error estimators and its impact on gene selection algorithmsabstractMOTIVATION: Feature selection approaches, such as filter and wrapper, have been applied to address the gene selection problem in the literature of microarray data analysis. In wrapper methods, the classification error is usually used as the evaluation criterion of feature subsets. Due to the nature of high dimensionality and small sample size of microarray data, however, counting-based error estimation may not necessarily be an ideal criterion for gene selection problem. RESULTS: Our study reveals that evaluating genes in terms of counting-based error estimators such as resubstitution error, leave-one-out error, cross-validation error and bootstrap error may encounter severe ties problem, i.e. two or more gene subsets score equally, and this in turn results in uncertainty in gene selection. Our analysis finds that the ties problem is caused by the discrete nature of counting-based error estimators and could be avoided by using continuous evaluation criteria instead. Experiment results show that continuous evaluation criteria such as generalised the absolute value of w2 measure for support vector machines and modified Relief's measure for k-nearest neighbors produce improved gene selection compared with counting-based error estimators. AVAILABILITY: The companion website is at http://www.ntu.edu.sg/home5/pg02776030/wrappers/ The website contains (1) the source code of all the gene selection algorithms and (2) the complete set of tables and figures of experiments. Kezhi Mao |
Bioinform. | 2 |
| 2006 | Regularization Network-based Gene Selection for Microarray Data AnalysisabstractMicroarray data contains a large number of genes (usually more than 1000) and a relatively small number of samples (usually fewer than 100). This presents problems to discriminant analysis of microarray data. One way to alleviate the problem is to reduce dimensionality of data by selecting important genes to the discriminant problem. Gene selection can be cast as a feature selection problem in the context of pattern classification. Feature selection approaches are broadly grouped into filter methods and wrapper methods. The wrapper method outperforms the filter method but at the cost of more intensive computation. In the present study, we proposed a wrapper-like gene selection algorithm based on the Regularization Network. Compared with classical wrapper method, the computational costs in our gene selection algorithm is significantly reduced, because the evaluation criterion we proposed does not demand repeated training in the leave-one-out procedure. Kezhi Mao |
Int. J. Neural Syst. | 2 |
| 2005 | Gene Selection of DNA Microarray Data Based on Regularization Networks
Kezhi Mao |
IDEAL | 2 |
| 2005 | Feature Selection Algorithm for Data with Both Nominal and Continuous Features
Wenyin Tang, Kezhi Mao |
PAKDD | 2 |
| 2005 | LS Bound based gene selection for DNA microarray dataabstractMOTIVATION: One problem with discriminant analysis of DNA microarray data is that each sample is represented by quite a large number of genes, and many of them are irrelevant, insignificant or redundant to the discriminant problem at hand. Methods for selecting important genes are, therefore, of much significance in microarray data analysis. In the present study, a new criterion, called LS Bound measure, is proposed to address the gene selection problem. The LS Bound measure is derived from leave-one-out procedure of LS-SVMs (least squares support vector machines), and as the upper bound for leave-one-out classification results it reflects to some extent the generalization performance of gene subsets. RESULTS: We applied this LS Bound measure for gene selection on two benchmark microarray datasets: colon cancer and leukemia. We also compared the LS Bound measure with other evaluation criteria, including the well-known Fisher's ratio and Mahalanobis class separability measure, and other published gene selection algorithms, including Weighting factor and SVM Recursive Feature Elimination. The strength of the LS Bound measure is that it provides gene subsets leading to more accurate classification results than the filter method while its computational complexity is at the level of the filter method. AVAILABILITY: A companion website can be accessed at http://www.ntu.edu.sg/home5/pg02776030/lsbound/. The website contains: (1) the source code of the gene selection algorithm; (2) the complete set of tables and figures regarding the experimental study; (3) proof of the inequality (9). CONTACT: [email protected]. Kezhi Mao |
Bioinform. | 2 |
| 2005 | Fast Modular network implementation for support vector machinesabstractSupport vector machines (SVMs) have been extensively used. However, it is known that SVMs face difficulty in solving large complex problems due to the intensive computation involved in their training algorithms, which are at least quadratic with respect to the number of training examples. This paper proposes a new, simple, and efficient network architecture which consists of several SVMs each trained on a small subregion of the whole data sampling space and the same number of simple neural quantizer modules which inhibit the outputs of all the remote SVMs and only allow a single local SVM to fire (produce actual output) at any time. In principle, this region-computing based modular network method can significantly reduce the learning time of SVM algorithms without sacrificing much generalization performance. The experiments on a few real large complex benchmark problems demonstrate that our method can be significantly faster than single SVMs without losing much generalization performance. Guang-Bin Huang, Kezhi Mao, Chee Kheong Siew, De-Shuang Huang |
IEEE Trans. Neural Networks | 2 |
| 2005 | Neuron selection for RBF neural network classifier based on data structure preserving criterionabstractThe central problem in training a radial basis function neural network is the selection of hidden layer neurons. In this paper, we propose to select hidden layer neurons based on data structure preserving criterion. Data structure denotes relative location of samples in the high-dimensional space. By preserving the data structure of samples including those that are close to separation boundaries between different classes, the neuron subset selected retains the separation margin underlying the full set of hidden layer neurons. As a direct result, the network obtained tends to generalize well. Kezhi Mao, Guang-Bin Huang |
IEEE Trans. Neural Networks | 1 |
| 2005 | Identifying critical variables of principal components for unsupervised feature selectionabstractPrincipal components analysis (PCA) is probably the best-known approach to unsupervised dimensionality reduction. However, axes of the lower-dimensional space, ie., principal components (PCs), are a set of new variables carrying no clear physical meanings. Thus, interpretation of results obtained in the lower-dimensional PCA space and data acquisition for test samples still involve all of the original measurements. To deal with this problem, we develop two algorithms to link the physically meaningless PCs back to a subset of original measurements. The main idea of the algorithms is to evaluate and select feature subsets based on their capacities to reproduce sample projections on principal axes. The strength of the new algorithms is that the computaion complexity involved is significantly reduced, compared with the data structural similarity-based feature evaluation. Kezhi Mao |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2004 | Feature subset selection for support vector machines through discriminative function pruning analysisabstractIn many pattern classification applications, data are represented by high dimensional feature vectors, which induce high computational cost and reduce classification speed in the context of support vector machines (SVMs). To reduce the dimensionality of pattern representation, we develop a discriminative function pruning analysis (DFPA) feature subset selection method in the present study. The basic idea of the DFPA method is to learn the SVM discriminative function from training data using all input variables available first, and then to select feature subset through pruning analysis. In the present study, the pruning is implement using a forward selection procedure combined with a linear least square estimation algorithm, taking advantage of linear-in-the-parameter structure of the SVM discriminative function. The strength of the DFPA method is that it combines good characters of both filter and wrapper methods. Firstly, it retains the simplicity of the filter method avoiding training of a large number of SVM classifier. Secondly, it inherits the good performance of the wrapper method by taking the SVM classification algorithm into account. Kezhi Mao |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2004 | Orthogonal forward selection and backward elimination algorithms for feature subset selectionabstractSequential forward selection (SFS) and sequential backward elimination (SBE) are two commonly used search methods in feature subset selection. In the present study, we derive an orthogonal forward selection (OFS) and an orthogonal backward elimination (OBE) algorithms for feature subset selection by incorporating Gram-Schmidt and Givens orthogonal transforms into forward selection and backward elimination procedures, respectively. The basic idea of the orthogonal feature subset selection algorithms is to find an orthogonal space in which to express features and to perform feature subset selection. After selection, the physically meaningless features in the orthogonal space are linked back to the same number of input variables in the original measurement space. The strength of employing orthogonal transforms is that features are decorrelated in the orthogonal space, hence individual features can be evaluated and selected independently. The effectiveness of our algorithms to deal with real world problems is finally demonstrated. Kezhi Mao |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2002 | RBF neural network center selection based on Fisher ratio class separability measureabstractFor classification applications, the role of hidden layer neurons of a radial basis function (RBF) neural network can be interpreted as a function which maps input patterns from a nonlinear separable space to a linear separable space. In the new space, the responses of the hidden layer neurons form new feature vectors. The discriminative power is then determined by RBF centers. In the present study, we propose to choose RBF centers based on Fisher ratio class separability measure with the objective of achieving maximum discriminative power. We implement this idea using a multistep procedure that combines Fisher ratio, an orthogonal transform, and a forward selection search method. Our motivation of employing the orthogonal transform is to decouple the correlations among the responses of the hidden layer neurons so that the class separability provided by individual RBF neurons can be evaluated independently. The strengths of our method are double fold. First, our method selects a parsimonious network architecture. Second, this method selects centers that provide large class separation. Kezhi Mao |
IEEE Trans. Neural Networks | 1 |
| 2002 | Fast orthogonal forward selection algorithm for feature subset selectionabstractFeature selection is an important issue in pattern classification. In the presented study, we develop a fast orthogonal forward selection (FOFS) algorithm for feature subset selection. The FOFS algorithm employs an orthogonal transform to decompose correlations among candidate features, but it performs the orthogonal decomposition in an implicit way. Consequently, the fast algorithm demands less computational effort as compared with conventional orthogonal forward selection (OFS). Kezhi Mao |
IEEE Trans. Neural Networks | 1 |
| 2000 | Probabilistic neural-network structure determination for pattern classificationabstractNetwork structure determination is an important issue in pattern classification based on a probabilistic neural network. In this study, a supervised network structure determination algorithm is proposed. The proposed algorithm consists of two parts and runs in an iterative way. The first part identifies an appropriate smoothing parameter using a genetic algorithm, while the second part determines suitable pattern layer neurons using a forward regression orthogonal algorithm. The proposed algorithm is capable of offering a fairly small network structure with satisfactory classification accuracy. Kezhi Mao, Kah-Chye Tan, Wee Ser |
IEEE Trans. Neural Networks Learn. Syst. | 1 |