Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xue-wen Chen 0001

dblp:18/3298-1 · also Xue-wen (William) Chen · DBLP profile ↗
← Back
65ranked-venue papers
14as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 10 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 3 first-authorDatabases, data management, data science and information retrieval · 13 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorTheory of computation · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 36% Representation and self-supervised learning · 24% Trustworthy machine learning · 14%
Databases, data mining, and information retrieval
4 papers
Information retrieval · 79% Data mining · 13% Recommender systems · 8%
Interdisciplinary, comprehensive, and emerging computing
7 papers
Bioinformatics and computational biology · 100%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
generative adversarial network
0.712023
DivGAN: A diversity enforcing generative adversarial network for mode collapse reduction · Artif. Intell. 2023
Machine learning › Generative modeling › generative adversarial network › GAN training
mode collapse
0.712023
DivGAN: A diversity enforcing generative adversarial network for mode collapse reduction · Artif. Intell. 2023
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.312017
Manifold Learning by Curved Cosine Mapping · IEEE Trans. Knowl. Data Eng. 2017
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.312017
Manifold Learning by Curved Cosine Mapping · IEEE Trans. Knowl. Data Eng. 2017
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
nonlinear dimensionality reduction
0.312017
Manifold Learning by Curved Cosine Mapping · IEEE Trans. Knowl. Data Eng. 2017
Information retrieval › image retrieval
content-based image retrieval
0.322013
iLike: Bridging the Semantic Gap in Vertical Image Search by Integrating Text and Visual Features · IEEE Trans. Knowl. Data Eng. 2013
iLike: integrating visual and textual features for vertical search · ACM Multimedia 2010
Information retrieval
image retrieval
0.322013
iLike: Bridging the Semantic Gap in Vertical Image Search by Integrating Text and Visual Features · IEEE Trans. Knowl. Data Eng. 2013
iLike: integrating visual and textual features for vertical search · ACM Multimedia 2010
Machine learning › Deep learning architectures and training › regularization
dropout
0.212016
Learning Deep Networks from Noisy Labels with Dropout Regularization · ICDM 2016
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.212016
Learning Deep Networks from Noisy Labels with Dropout Regularization · ICDM 2016
Machine learning › Deep learning architectures and training
regularization
0.212016
Learning Deep Networks from Noisy Labels with Dropout Regularization · ICDM 2016
Machine learning › Trustworthy machine learning › robustness
robust learning
0.212016
Learning Deep Networks from Noisy Labels with Dropout Regularization · ICDM 2016
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene ontology analysis
0.212016
GOAL: the comprehensive gene ontology analysis layer · Sci. China Inf. Sci. 2016
Machine learning › Learning paradigms
class imbalance
0.222010
Combating the Small Sample Class Imbalance Problem Using Feature Selection · IEEE Trans. Knowl. Data Eng. 2010
FAST: a roc-based feature selection metric for small samples and imbalanced data classification problems · KDD 2008
Information retrieval
multimodal retrieval
0.212013
iLike: Bridging the Semantic Gap in Vertical Image Search by Integrating Text and Visual Features · IEEE Trans. Knowl. Data Eng. 2013
Information retrieval
retrieval models
0.212013
iLike: Bridging the Semantic Gap in Vertical Image Search by Integrating Text and Visual Features · IEEE Trans. Knowl. Data Eng. 2013
Data mining › dimensionality reduction
feature selection
0.222008
FAST: a roc-based feature selection metric for small samples and imbalanced data classification problems · KDD 2008
Minimum reference set based feature selection for small sample classifications · ICML 2007
Bioinformatics and computational biology › protein-protein interaction prediction
domain-domain interaction inference
0.122009
Knowledge-guided inference of domain-domain interactions from incomplete protein-protein interaction networks · Bioinform. 2009
Prediction of protein-protein interactions using random decision forest framework · Bioinform. 2005
Information retrieval
e-commerce search
0.112010
iLike: integrating visual and textual features for vertical search · ACM Multimedia 2010
Recommender systems › multimodal recommendation
multimodal fusion
0.112010
iLike: integrating visual and textual features for vertical search · ACM Multimedia 2010
Information retrieval › web search
vertical search
0.112010
iLike: integrating visual and textual features for vertical search · ACM Multimedia 2010
Bioinformatics and computational biology › functional genomics
gene function prediction
0.112009
Identification of genes involved in the same pathways using a Hidden Markov Model-based approach · Bioinform. 2009
Bioinformatics and computational biology › protein analysis › protein-protein interaction
protein-protein interaction analysis
0.112009
Knowledge-guided inference of domain-domain interactions from incomplete protein-protein interaction networks · Bioinform. 2009
Bioinformatics and computational biology › protein-protein interaction prediction
protein-protein interaction site prediction
0.112009
Sequence-based prediction of protein interaction sites with an integrative method · Bioinform. 2009
Bioinformatics and computational biology › protein-protein interaction prediction › protein-protein interaction site prediction
sequence-based protein interaction site prediction
0.112009
Sequence-based prediction of protein interaction sites with an integrative method · Bioinform. 2009
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian network
0.112008
Improving Bayesian Network Structure Learning with Mutual Information-Based Node Ordering in the K2 Algorithm · IEEE Trans. Knowl. Data Eng. 2008
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
structure learning
0.112008
Improving Bayesian Network Structure Learning with Mutual Information-Based Node Ordering in the K2 Algorithm · IEEE Trans. Knowl. Data Eng. 2008
Machine learning › Probabilistic and Bayesian machine learning › noise modeling
label noise modeling
0.112016
Learning Deep Networks from Noisy Labels with Dropout Regularization · ICDM 2016
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference
0.112006
An effective structure learning method for constructing gene networks · Bioinform. 2006
Bioinformatics and computational biology › protein-protein interaction prediction
domain-domain interaction prediction
0.112005
Prediction of protein-protein interactions using random decision forest framework · Bioinform. 2005
Bioinformatics and computational biology
protein-protein interaction prediction
0.112005
Prediction of protein-protein interactions using random decision forest framework · Bioinform. 2005

Methods — techniques the papers use, named apart from their topics

generative adversarial network · 0.7law of cosines · 0.3visual feature re-weighting · 0.3stochastic gradient descent · 0.2software tool · 0.2softmax noise model · 0.2dropout · 0.2visual thesaurus construction · 0.2ROC curve analysis · 0.2minimum reference set · 0.1textual-visual feature combination · 0.1resampling · 0.1precision-recall curve · 0.1feature selection metrics · 0.1AUC · 0.1random forest · 0.1microarray expression analysis · 0.1imbalanced data classification · 0.1
YearPublicationVenuePosition
2023 DivGAN: A diversity enforcing generative adversarial network for mode collapse reduction
Manal Nawar Allahyani, Rahaf Alsulami, Taif Alwafi, Tarik K. Alafif, Heyfa Ammar, Sari Sabban, Xue-wen Chen 0001
Artif. Intell.7
2020 Locality preserving difference component analysis based on the Lq norm
Zhizheng Liang, Xue-wen Chen 0001, Lei Zhang 0029, Jin Liu 0006, Yong Zhou 0003
Pattern Anal. Appl.2
2020 Correlation classifiers based on data perturbation: New formulations and algorithms
Zhizheng Liang, Xue-wen Chen 0001, Lei Zhang 0029, Jin Liu 0006, Yong Zhou 0003
Pattern Recognit.2
2019 "Stream loss": ConvNet learning for face verification using unlabeled videos in the wild
Elaheh Rashedi, Elaheh Barati, Matthew S. Nokleby, Xue-wen Chen 0001
Neurocomputing4
2018 Optimizing Taxi Carpool Policies via Reinforcement Learning and Spatio-Temporal Mining
abstract
In this paper, we develop a reinforcement learning (RL) based system to learn an effective policy for carpooling that maximizes transportation efficiency so that fewer cars are required to fulfill the given amount of trip demand. For this purpose, first, we develop a deep neural network model, called ST-NN (Spatio-Temporal Neural Network), to predict taxi trip time from the raw GPS trip data. Secondly, we develop a carpooling simulation environment for RL training, with the output of ST-NN and using the NYC taxi trip dataset. In order to maximize transportation efficiency and minimize traffic congestion, we choose the effective distance covered by the driver on a carpool trip as the reward. Therefore, the more effective distance a driver achieves over a trip (i.e. to satisfy more trip demand) the higher the efficiency and the less will be the traffic congestion. We compared the performance of RL learned policy to a fixed policy (which always accepts carpool) as a baseline and obtained promising results that are interpretable and demonstrate the advantage of our RL approach. We also compare the performance of ST-NN to that of state-of-the-art travel time estimation methods and observe that ST-NN significantly improves the prediction performance and is more robust to outliers.
Ishan Jindal, Zhiwei (Tony) Qin, Xue-wen Chen 0001, Matthew S. Nokleby, Jieping Ye
IEEE BigData3
2018 Teacher/Student Deep Semi-Supervised Learning for Training with Noisy Labels
abstract
Deep learning methods are at the forefront of leading state-of-the-art methods in a wide range of machine learning applications. In particular, convolutional neural networks (CNNs) attain topmost performance assuming a sufficiently large number of labeled training examples. Unfortunately, labeled data is artificially curated, and it requires human labor, which consequently makes it expensive and time-consuming. Moreover, there are no guarantees that the obtained labels are noise-free. In fact, the performance of CNNs is influenced by the level of noisy labels in the training dataset. Although the literature lacks attention to train learning methods with noisy labels, few semi-supervised learning methods mitigate this obstacle. In this paper, we propose a new teacher/student deep semi-supervised learning (TS-DSSL) method that employs self-training on noisy labels training dataset. We measure the efficiency of TS-DSSL on semi-supervised visual object classification tasks on the benchmark datasets CIFAR10 and MNIST. TS-DSSL achieves impressive results even in the presence of high-level noisy labels. It also sets a record on datasets with various levels of noisy labels created from the previous datasets with uniform and non-uniform noise distributions.
Zeyad Hailat, Xue-wen Chen 0001
ICMLA2
2018 Deep Semi-Supervised Learning
abstract
Convolutional neural networks (CNNs) attain state-of-the-art performance on various classification tasks assuming a sufficiently large number of labeled training examples. Unfortunately, curating sufficiently large labeled training dataset requires human involvement, which is expensive and time consuming. Semi-supervised methods can alleviate this problem by utilizing a limited number of labeled data in conjunction with sufficiently large unlabeled data to construct a classification model. Self-training techniques are among the earliest semi-supervised methods proposed to enhance learning by utilizing unlabeled data. In this paper, we propose a deep semi-supervised learning (DSSL) self-training method that utilizes the strengths of both supervised and unsupervised learning within a single model. We measure the efficacy of the proposed method on semi-supervised visual object classification tasks using the datasets CIFAR-10, CIFAR-100, STL-10, MNIST, and SVHN. The experiments show that DSSL surpasses semi-supervised state-of-the-art methods for most of the aforementioned datasets.
Zeyad Hailat, Artem Komarichev, Xue-wen Chen 0001
ICPR3
2017 On Classifying Facial Races with Partial Occlusions and Pose Variations
abstract
Many biometrics and security systems use facial information to obtain an individual identification and recognition. Classifying a race from a face image can provide a strong hint to search for facial identity and criminal identification. Current facial race classification methods are confined only to constrained non-partially occluded frontal faces. Challenges remain under unconstrained environments such as partial occlusions and pose variations. In this paper, we propose a Convolutional Neural Network (CNN) model to classify facial races with partial occlusions and pose variations. The proposed model is trained using a broad and balanced racial distributed face image dataset. The model is trained on four major human races, Caucasian, Indian, Mongolian, and Negroid. Our model is evaluated against the state-of-the-art methods on a constrained face test dataset. Also, an evaluation of the proposed model and human performance is conducted and compared on our new unconstrained facial race benchmark (CIMN) dataset. Our results show that our model achieves 95.1% of race classification accuracy on constrained frontal faces. Also, the proposed model achieves a comparable classification accuracy result compared to human performance with a margin of 6.2% under the current challenges in the unconstrained environment.
Tarik K. Alafif, Zeyad Hailat, Melih S. Aslan, Xue-wen Chen 0001
ICMLA4
2017 Multi-channel multi-model feature learning for face recognition
Melih S. Aslan, Zeyad Hailat, Tarik K. Alafif, Xue-wen Chen 0001
Pattern Recognit. Lett.4
2017 Manifold Learning by Curved Cosine Mapping
abstract
In the field of pattern recognition, data analysis, and machine learning, data points are usually modeled as high-dimensional vectors. Due to the curse-of-dimensionality, it is non-trivial to efficiently process the orginal data directly. Given the unique properties of nonlinear dimensionality reduction techniques, nonlinear learning methods are widely adopted to reduce the dimension of data. However, existing nonlinear learning methods fail in many real applications because of the too-strict requirements (for real data) or the difficulty in parameters tuning. Therefore, in this paper, we investigate the manifold learning methods which belong to the family of nonlinear dimensionality reduction methods. Specifically, we proposed a new manifold learning principle for dimensionality reduction named Curved Cosine Mapping (CCM). Based on the law of cosines in Euclidean space, CCM applies a brand new mapping pattern to manifold learning. In CCM, the nonlinear geometric relationships are obtained by utlizing the law of cosines, and then quantified as the dimensionality-reduced features. Compared with the existing approaches, the model has weaker theoretical assumptions over the input data. Moreover, to further reduce the computation cost, an optimized version of CCM is developed. Finally, we conduct extensive experiments over both artificial and real-world datasets to demonstrate the performance of proposed techniques.
Huamao Gu, Xun Wang 0007, Xue-wen Chen 0001, Shaoping Deng, Jin-Qin Shi
IEEE Trans. Knowl. Data Eng.3
2016 Learning Deep Networks from Noisy Labels with Dropout Regularization
abstract
Large datasets often have unreliable labels-such as those obtained from Amazon's Mechanical Turk or social media platforms-and classifiers trained on mislabeled datasets often exhibit poor performance. We present a simple, effective technique for accounting for label noise when training deep neural networks. We augment a standard deep network with a softmax layer that models the label noise statistics. Then, we train the deep network and noise model jointly via end-to-end stochastic gradient descent on the (perhaps mislabeled) dataset. The augmented model is underdetermined, so in order to encourage the learning of a non-trivial noise model, we apply dropout regularization to the weights of the noise model during training. Numerical experiments on noisy versions of the CIFAR-10 and MNIST datasets show that the proposed dropout technique outperforms state-of-the-art methods.
Ishan Jindal, Matthew S. Nokleby, Xue-wen Chen 0001
ICDM3
2016 Latent Topic-Semantic Indexing Based Automatic Text Summarization
abstract
Automatic summarization, a difficult but pressing problem in natural language processing, aims at shortening source documents while retaining main information. In recent years, more statistical machine learning methods have been applied to automatic summarization. In this paper, we propose a novel approach for summarization, based on hierarchical Bayesian model of topic-semantic indexing (TSI) and extraction strategy of average log-likelihood. The new method is tested on Brown corpus, and its performance is analyzed by a well-designed blind experiment of one-way ANOVA on human reviews. The experimental results show that TSI model is promising on topic-driven summarization.
Jiangsheng Yu, Xue-wen Chen 0001
ICMLA2
2016 GOAL: the comprehensive gene ontology analysis layer
Jong Cheol Jeong, George Li 0003, Xue-wen Chen 0001
Sci. China Inf. Sci.3
2015 Unsupervised Learning and Image Classification in High Performance Computing Cluster
abstract
Feature learning and object classification in machine learning have become very active research areas in recent decades. Identifying good features has various benefits for object classification in respect to reducing the computational cost and increasing the classification accuracy. We propose using a multimodal learning and object identification framework with an alternative platform, called High Performance Computing Cluster (HPCC Systems®), to speed up the optimization stages and to handle data of any dimension. Our framework first learns representative bases (or centroids) over unlabeled data for each model through the K-means unsupervised learning method. Then, to extract the desired features from the labeled data, the correlation between the labeled data and representative bases is calculated. These labeled features are fused to represent the identity and then fed to the classifiers to make the final recognition. In addition, many research studies have focused on improving optimization methods and the use of Graphics Processing Units (GPUs) to improve the training time for machine learning algorithms. This study is aimed at exploring feature learning and object classification ideas in HPCC Systems platform. HPCC Systems is a Big Data processing and massively parallel processing (MPP) computing platform used for solving Big Data problems. Algorithms are implemented in HPCC Systems® with a language called Enterprise Control Language (ECL) which is a declarative, data-centric programming language. It is a powerful, high-level, parallel programming language ideal for Big Data intensive applications. We evaluate our proposed framework in this new platform on various databases such as the CALTECH-101, AR databases, and a subset of wild PubFig83 data that we add multimedia content. For instance, we are able to improve on the classification accuracy result of [3] from 74.3% to 78.9% on AR database using Decision Tree C4.5 classifier.
Itauma Itauma, Melih S. Aslan, Flavio Villanustre, Xue-wen Chen 0001
ICMLA4
2015 A New Semantic Functional Similarity over Gene Ontology
abstract
Identifying functionally similar or closely related genes and gene products has significant impacts on biological and clinical studies as well as drug discovery. In this paper, we propose an effective and practically useful method measuring both gene and gene product similarity by integrating the topology of gene ontology, known functional domains and their functional annotations. The proposed method is comprehensively evaluated through statistical analysis of the similarities derived from sequence, structure and phylogenetic profiles, and clustering analysis of disease genes clusters. Our results show that the proposed method clearly outperforms other conventional methods. Furthermore, literature analysis also reveals that the proposed method is both statistically and biologically promising for identifying functionally similar genes or gene products. In particular, we demonstrate that the proposed functional similarity metric is capable of discoverying new disease related genes or gene products.
Jong Cheol Jeong, Xue-wen Chen 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2014 Learning sparse and scale-free networks
abstract
Gaussian networks study undirected interactions between random variables, through the estimation of the precision matrices. Recently, it has been demonstrated that some of the important networks display features similar to scale-free graphs. There have been few works on the learning of the sparse Gaussian graphical models aiming to preserve properties of networks which are believed to be scale-free or have dominating hubs. We prefer to name both networks as ‘scale-free’ networks for simplicity in this paper. We propose a new log-likelihood formulation, which promotes the sparseness of the precision matrix and features of scale-free graphical topology. We used the alternating direction method of multipliers (ADMM) form, which is used for the convex optimization, to solve the general L1regularized loss optimization. Our proposed method exhibits better estimation performance on various data sets and various number of samples, N. Also, the proposed method and some of the state of the arts methods are tested under various penalty constants to validate the robustness.
Melih S. Aslan, Xue-wen Chen 0001, Hong Cheng 0002
DSAA2
2014 Pseudo labels for imbalanced multi-label learning
abstract
The classification with instances which can be tagged with any of the 2Lpossible subsets from the predefined L labels is called multi-label classification. Multi-label classification is commonly applied in domains, such as multimedia, text, web and biological data analysis. The main challenge lying in multi-label classification is the dilemma of optimising label correlations over exponentially large label powerset and the ignorance of label correlations using binary relevance strategy (1-vs-all heuristic). The classification with label powerset usually encounters with highly skewed data distribution, called imbalanced problem. While binary relevance strategy reduces the problem from exponential to linear, it totally neglects the label correlations. In this artical, we propose a novel strategy of introducing Balanced Pseudo-Labels (BPL) which build more robust classifiers for imbalanced multi-label classification, which embeds imbalanced data in the problems innately. By incorporating the new balanced labels we aim to increase the average distances among the distinct label vectors. In this way, we also code the label correlation implicitly in the algorithm. Another advantage of the proposed method is that it can combined with any classifier and it is proportional to linear label transformation. In the experiment, we choose five multi-label benchmark data sets and compare our algorithm with the most state-of-art algorithms. Our algorithm outperforms them in standard multi-label evaluation in most scenarios.
Wenrong Zeng, Xue-wen Chen 0001, Hong Cheng 0002
DSAA2
2014 A windowed dynamic time warping approach for 3D continuous hand gesture recognition
abstract
Detecting the beginning and end of a specific gesture from an infinite trajectory gesture sequence has gained considerable interests in the past several years. Traditional begin-end dynamic time warping approach for gesture recognition could provide multiple different gesture labels for one trajectory segment. This paper presents a Windowed Dynamic Time Warping (WDTW) approach for 3D continuous hand trajectory gesture recognition. The main contribution is that we introduce a parameterized searching window in the cost matrix of traditional DTW approach to detect the beginning and end of the specific gesture from an infinite trajectory gesture sequence. By doing so, we formulate continuous gesture recognition into online parameter estimation of the searching window. Moreover, the proposed gesture recognition can handle the multilabel issue. We evaluate the proposed windowed dynamic time warping approach in our gesture dataset. The experimental results show that the proposed WDTW can significantly improve the begin-end gesture recognition performance.
Hong Cheng 0002, Xue-wen Chen 0001
ICME3
2014 Pixel-to-Model background modeling in crowded scenes
abstract
Background modeling is an important step for many video surveillance applications such as object detection and scene understanding. In this paper, we present a novel Pixel-to-Model (P2M) paradigm for background modeling in crowded scenes. In particular, the proposed method models the background with a set of context features for each pixel, which are compressively sensed from local patches. We determine whether a pixel belongs to the background according to the minimum P2M distance, which measures the similarity between the pixel and its background model in the space of compressive local descriptors. Moreover, the background updating utilizes minimum and maximum P2M distances to update the pixel feature descriptors in local and neighboring background models, respectively. We evaluate the proposed approach with foreground detection tasks on real crowded surveillance videos. Experiments results show that the proposed P2M approach outperforms the state-of-the-art methods both in indoor and outdoor crowded scenes.
Lu Yang 0002, Hong Cheng 0002, Jianan Su, Xue-wen Chen 0001
ICME4
2014 Kernelized pyramid nearest-neighbor search for object categorization
Hong Cheng 0002, Rongchao Yu, Zicheng Liu 0001, Lu Yang 0002, Xue-wen Chen 0001
Mach. Vis. Appl.5
2013 Evaluating topology-based metrics for GO term similarity measures
abstract
Defining semantic functional similarity measures provides effective means to validate protein function prediction methods and to retrieve biologically relevant information from big biological data. It also improves understanding of interrelationship between genes and gene products (GPs). Currently, one of the most commonly used tools for functionally annotating genes and GPs is the Gene Ontology (GO), which describes genes/GPs using a machine-readable language. To measure the semantic similarity between two GO terms, many studies that are based on GO topology have recently been reported. However, a comprehensive assessment and general guidelines for validating these methods are lacking. In this paper, we collect a large dataset to evaluate five often-used semantic similarity measure methods by estimating sequence similarity, phylogenetic profile similarity, and structural similarity. We further compare the measures in terms of their clustering performance using domains extracted from SCOP database. We describe some key aspects of these measure methods and discuss how the limitations may be addressed as well as some open problems.
Jong Cheol Jeong, Xue-wen Chen 0001
BIBM2
2013 Content-based assessment of the credibility of online healthcare information
abstract
Currently, a large amount of data is produced in healthcare informatics due to the growth of web technologies like social networks, wikis, blogs and RSS feeds. However, not all health information provided online is trustworthy. Even though many experts are involved in publishing trusted information, it is difficult for the general population to determine the credibility of the information. Therefore, a reliable mechanism to automatically determine the trustworthiness of online healthcare information is highly desired. In this paper, we propose two novel approaches based on Topic Modeling and Hidden Markov Models (HMMs), that can be applied over a large volume of online healthcare data to assess its trustworthiness. Traditional Topic Modeling is solely based on the “bag-of-words” model, however, we also consider the semantics of the content to identify the underlying topics in a sentence. For the HMM approach, we built our trustworthy and suspicious models after analyzing the characteristics of sentences from such websites. Both methods perform well to assess the trustworthiness, however HMM is less sophisticated to capture the semantics of sentences. We evaluated our method on randomly chosen real dataset and are able to achieve about 90% accuracy in identifying the trustworthiness of the content.
Meeyoung Park, Hariprasad Sampathkumar, Bo Luo, Xue-wen Chen 0001
IEEE BigData4
2013 Sparse representation and learning in visual recognition: Theory and applications
Hong Cheng 0002, Zicheng Liu 0001, Lu Yang 0002, Xue-wen Chen 0001
Signal Process.4
2013 iLike: Bridging the Semantic Gap in Vertical Image Search by Integrating Text and Visual Features
abstract
With the development of Internet and Web 2.0, large-volume multimedia contents have been made available online. It is highly desired to provide easy accessibility to such contents, i.e., efficient and precise retrieval of images that satisfies users' needs. Toward this goal, content-based image retrieval (CBIR) has been intensively studied in the research community, while text-based search is better adopted in the industry. Both approaches have inherent disadvantages and limitations. Therefore, unlike the great success of text search, web image search engines are still premature. In this paper, we present iLike, a vertical image search engine that integrates both textual and visual features to improve retrieval performance. We bridge the semantic gap by capturing the meaning of each text term in the visual feature space, and reweight visual features according to their significance to the query terms. We also bridge the user intention gap because we are able to infer the "visual meanings" behind the textual queries. Last but not least, we provide a visual thesaurus, which is generated from the statistical similarity between the visual space representation of textual terms. Experimental results show that our approach improves both precision and recall, compared with content-based or text-based image retrieval techniques. More importantly, search results from iLike is more consistent with users' perception of the query terms.
Yuxin Chen 0001, Hariprasad Sampathkumar, Bo Luo, Xue-wen Chen 0001
IEEE Trans. Knowl. Data Eng.4
2012 Large-scale prediction of adverse drug reactions using chemical, biological, and phenotypic properties of drugs
abstract
OBJECTIVE: Adverse drug reaction (ADR) is one of the major causes of failure in drug development. Severe ADRs that go undetected until the post-marketing phase of a drug often lead to patient morbidity. Accurate prediction of potential ADRs is required in the entire life cycle of a drug, including early stages of drug design, different phases of clinical trials, and post-marketing surveillance. METHODS: Many studies have utilized either chemical structures or molecular pathways of the drugs to predict ADRs. Here, the authors propose a machine-learning-based approach for ADR prediction by integrating the phenotypic characteristics of a drug, including indications and other known ADRs, with the drug's chemical structures and biological properties, including protein targets and pathway information. A large-scale study was conducted to predict 1385 known ADRs of 832 approved drugs, and five machine-learning algorithms for this task were compared. RESULTS: This evaluation, based on a fivefold cross-validation, showed that the support vector machine algorithm outperformed the others. Of the three types of information, phenotypic data were the most informative for ADR prediction. When biological and phenotypic features were added to the baseline chemical information, the ADR prediction model achieved significant improvements in area under the curve (from 0.9054 to 0.9524), precision (from 43.37% to 66.17%), and recall (from 49.25% to 63.06%). Most importantly, the proposed model successfully predicted the ADRs associated with withdrawal of rofecoxib and cerivastatin. CONCLUSION: The results suggest that phenotypic information on drugs is valuable for ADR prediction. Moreover, they demonstrate that different models that combine chemical, biological, or phenotypic information can be built from approved drugs, and they have the potential to detect clinically important ADRs in both preclinical and post-marketing phases.
Yonghui Wu 0001, Yukun Chen 0001, Jingchun Sun, Zhongming Zhao, Xue-wen Chen 0001, Michael E. Matheny, Hua Xu 0001
J. Am. Medical Informatics Assoc.6
2011 Functional Gene Detection and Clustering from Seed Gene Sets
abstract
The availability of rapidly increasing repositories of microarray data requires the help of computer-aided analysis techniques. This data combined with a growing knowledge base about molecular processes enables the use of intelligent machine learning algorithms to expand the existing knowledge base. In this paper, we propose a novel algorithm, namely iterated Hidden Markov Model, to query microarray expression data with genes known to be involved in the same function to produce novel genes involved with the same cellular function. We run this algorithm on publicly available benchmark data sets and show that it outperforms comparable machine learning approaches.
Alexander Senf, Xue-wen Chen 0001
BIBM2
2011 FEPI-MB: identifying SNPs-disease association using a Markov Blanket-based approach
abstract
BACKGROUND: The interactions among genetic factors related to diseases are called epistasis. With the availability of genotyped data from genome-wide association studies, it is now possible to computationally unravel epistasis related to the susceptibility to common complex human diseases such as asthma, diabetes, and hypertension. However, the difficulties of detecting epistatic interaction arose from the large number of genetic factors and the enormous size of possible combinations of genetic factors. Most computational methods to detect epistatic interactions are predictor-based methods and can not find true causal factor elements. Moreover, they are both time-consuming and sample-consuming. RESULTS: We propose a new and fast Markov Blanket-based method, FEPI-MB (Fast EPistatic Interactions detection using Markov Blanket), for epistatic interactions detection. The Markov Blanket is a minimal set of variables that can completely shield the target variable from all other variables. Learning of Markov blankets can be used to detect epistatic interactions by a heuristic search for a minimal set of SNPs, which may cause the disease. Experimental results on both simulated data sets and a real data set demonstrate that FEPI-MB significantly outperforms other existing methods and is capable of finding SNPs that have a strong association with common diseases. CONCLUSIONS: FEPI-MB algorithm outperforms other computational methods for detection of epistatic interactions in terms of both the power and sample-efficiency. Moreover, compared to other Markov Blanket learning methods, FEPI-MB is more time-efficient and achieves a better performance.
Bing Han 0001, Xue-wen Chen 0001, Zohreh Talebizadeh
BMC Bioinform.2
2011 On Position-Specific Scoring Matrix for Protein Function Prediction
abstract
While genome sequencing projects have generated tremendous amounts of protein sequence data for a vast number of genomes, substantial portions of most genomes are still unannotated. Despite the success of experimental methods for identifying protein functions, they are often lab intensive and time consuming. Thus, it is only practical to use in silico methods for the genome-wide functional annotations. In this paper, we propose new features extracted from protein sequence only and machine learning-based methods for computational function prediction. These features are derived from a position-specific scoring matrix, which has shown great potential in other bininformatics problems. We evaluate these features using four different classifiers and yeast protein data. Our experimental results show that features derived from the position-specific scoring matrix are appropriate for automatic function annotation.
Jong Cheol Jeong, Xiaotong Lin 0001, Xue-wen Chen 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2010 Detecting SNPs-disease associations using Bayesian networks
abstract
Epistatic interactions play a significant role in improving pathogenesis, prevention, diagnosis and treatment of complex human diseases. A recent study in automatic detection of epistatic interactions showed that Markov Blanket-based methods are capable of finding SNPs (single-nucleotide polymorphism) that have a strong association with common diseases and of reducing false positives when the number of instances is large. Unfortunately, a typical SNP dataset consists of very limited number of examples, where current methods including Markov Blanket-based methods perform poorly. To address small sample problems, we propose a Bayesian network-based approach to detect epistatic interactions. The proposed method also employs a Branch-and-Bound technique for learning. We apply the proposed method to simulated datasets based on four disease models and a real dataset. Experimental results show that our method significantly outperforms Markov Blanket-based methods and other commonly-used methods, especially when the number of samples is small.
Bing Han 0001, Xue-wen Chen 0001
BIBM2
2010 Mr.KNN: soft relevance for multi-label classification
abstract
Multi-label classification refers to learning tasks with each instance belonging to one or more classes simultaneously. It arose from real-world applications such as information retrieval, text categorization and functional genomics. Currently, most of the multi-label learning methods use the strategy called binary relevance, which constructs a classifier for each unique label by grouping data into positives (examples with this label) and negatives (examples without this label). With binary relevance, an example with multiple labels is considered as a positive data for each label it belongs to. For some classes, this data point may behave like an outlier confusing classifiers, especially in the cases of well-separated classes. In this paper, we first introduce a new strategy called soft relevance, where each multi-label example is assigned a relevance score to the labels it belongs to. This soft relevance is then employed in a voting function used in a k nearest neighbor classifier. Furthermore, a voting-margin ratio is introduced to the k nearest neighbor classifier for better performance. We compare the proposed method to other multi-label learning methods over three multi-label datasets and demonstrate that the proposed method provides an effective way to multi-label learning.
Xiaotong Lin 0001, Xue-wen Chen 0001
CIKM2
2010 iLike: integrating visual and textual features for vertical search
abstract
Content-based image search on the Internet is a challenging problem, mostly due to the semantic gap between low-level visual features and high-level content, as well as the excessive computation brought by huge amount of images and high dimensional features. In this paper, we present iLike, a new approach to truly combine textual features from web pages, and visual features from image content for better image search in a vertical search engine. We tackle the first problem by trying to capture the meaning of each text term in the visual feature space, and re-weight visual features according to their significance to the query content. Our experimental results in product search for apparels and accessories demonstrate the effectiveness of iLike and its capability of bridging semantic gaps between visual features and abstract concepts.
Yuxin Chen 0001, Nenghai Yu, Bo Luo, Xue-wen Chen 0001
ACM Multimedia4
2010 A Markov blanket-based method for detecting causal SNPs in GWAS
abstract
BACKGROUND: Detecting epistatic interactions associated with complex and common diseases can help to improve prevention, diagnosis and treatment of these diseases. With the development of genome-wide association studies (GWAS), designing powerful and robust computational method for identifying epistatic interactions associated with common diseases becomes a great challenge to bioinformatics society, because the study of epistatic interactions often deals with the large size of the genotyped data and the huge amount of combinations of all the possible genetic factors. Most existing computational detection methods are based on the classification capacity of SNP sets, which may fail to identify SNP sets that are strongly associated with the diseases and introduce a lot of false positives. In addition, most methods are not suitable for genome-wide scale studies due to their computational complexity. RESULTS: We propose a new Markov Blanket-based method, DASSO-MB (Detection of ASSOciations using Markov Blanket) to detect epistatic interactions in case-control GWAS. Markov blanket of a target variable T can completely shield T from all other variables. Thus, we can guarantee that the SNP set detected by DASSO-MB has a strong association with diseases and contains fewest false positives. Furthermore, DASSO-MB uses a heuristic search strategy by calculating the association between variables to avoid the time-consuming training process as in other machine-learning methods. We apply our algorithm to simulated datasets and a real case-control dataset. We compare DASSO-MB to other commonly-used methods and show that our method significantly outperforms other methods and is capable of finding SNPs strongly associated with diseases. CONCLUSIONS: Our study shows that DASSO-MB can identify a minimal set of causal SNPs associated with diseases, which contains less false positives compared to other existing methods. Given the huge size of genomic dataset produced by GWAS, this is critical in saving the potential costs of biological experiments and being an efficient guideline for pathogenesis research.
Bing Han 0001, Meeyoung Park, Xue-wen Chen 0001
BMC Bioinform.3
2010 Combating the Small Sample Class Imbalance Problem Using Feature Selection
abstract
The class imbalance problem is encountered in real-world applications of machine learning and results in a classifier's suboptimal performance. Researchers have rigorously studied the resampling, algorithms, and feature selection approaches to this problem. No systematic studies have been conducted to understand how well these methods combat the class imbalance problem and which of these methods best manage the different challenges posed by imbalanced data sets. In particular, feature selection has rarely been studied outside of text classification problems. Additionally, no studies have looked at the additional problem of learning from small samples. This paper presents a first systematic comparison of the three types of methods developed for imbalanced data classification problems and of seven feature selection metrics evaluated on small sample data sets from different applications. We evaluated the performance of these metrics using area under the receiver operating characteristic (AUC) and area under the precision-recall curve (PRC). We compared each metric on the average performance across all problems and on the likelihood of a metric yielding the best performance on a specific problem. We examined the performance of these metrics inside each problem domain. Finally, we evaluated the efficacy of these metrics to see which perform best across algorithms. Our results showed that signal-to-noise correlation coefficient (S2N) and Feature Assessment by Sliding Thresholds (FAST) are great candidates for feature selection in most applications, especially when selecting very small numbers of features.
Michael Wasikowski, Xue-wen Chen 0001
IEEE Trans. Knowl. Data Eng.2
2009 DASSO-MB: Detection of Epistatic Interactions in Genome-Wide Association Studies Using Markov Blankets
abstract
With the development of genome-wide association studies (GWAS), computationally identifying the epistatic interactions associated with common diseases presents a significant challenge to bioinformatics society. Most existing computational detection methods are based on the classification capacity of SNP sets, which may fail to identify SNP sets that are strongly associated with the diseases. In addition, most methods are not suitable for genome-wide scale studies due to their computational complexity. To address these issues, we propose the use of a Markov Blanket-based method, DASSO-MB, for epistatic interaction detection. We apply our method to both simulated data sets and a real data set, and demonstrate that DASSO-MB significantly outperforms other existing methods and is capable of finding SNPs that have a strong association with common diseases.
Bing Han 0001, Meeyoung Park, Xue-wen Chen 0001
BIBM3
2009 Learning to rank with a novel kernel perceptron method
abstract
While conventional ranking algorithms, such as the PageRank, rely on the web structure to decide the relevancy of a web page, learning to rank seeks a function capable of ordering a set of instances using a supervised learning approach. Learning to rank has gained increasing popularity in information retrieval and machine learning communities. In this paper, we propose a novel nonlinear perceptron method for rank learning. The proposed method is an online algorithm and simple to implement. It introduces a kernel function to map the original feature space into a nonlinear space and employs a perceptron method to minimize the ranking error by avoiding converging to a solution near the decision boundary and alleviating the effect of outliers in the training dataset. Furthermore, unlike existing approaches such as RankSVM and RankBoost, the proposed method is scalable to large datasets for online learning. Experimental results on benchmark corpora show that our approach is more efficient and achieves higher or comparable accuracies in instance ranking than state of the art methods such as FRank, RankSVM and RankBoost.
Xue-wen Chen 0001, Haixun Wang, Xiaotong Lin 0001
CIKM1
2009 Sequence-based prediction of protein interaction sites with an integrative method
abstract
MOTIVATION: Identification of protein interaction sites has significant impact on understanding protein function, elucidating signal transduction networks and drug design studies. With the exponentially growing protein sequence data, predictive methods using sequence information only for protein interaction site prediction have drawn increasing interest. In this article, we propose a predictive model for identifying protein interaction sites. Without using any structure data, the proposed method extracts a wide range of features from protein sequences. A random forest-based integrative model is developed to effectively utilize these features and to deal with the imbalanced data classification problem commonly encountered in binding site predictions. RESULTS: We evaluate the predictive method using 2829 interface residues and 24,616 non-interface residues extracted from 99 polypeptide chains in the Protein Data Bank. The experimental results show that the proposed method performs significantly better than two other sequence-based predictive methods and can reliably predict residues involved in protein interaction sites. Furthermore, we apply the method to predict interaction sites and to construct three protein complexes: the DnaK molecular chaperone system, 1YUW and 1DKG, which provide new insight into the sequence-function relationship. We show that the predicted interaction sites can be valuable as a first approach for guiding experimental methods investigating protein-protein interactions and localizing the specific interface residues. AVAILABILITY: Datasets and software are available at http://ittc.ku.edu/~xwchen/bindingsite/prediction.
Xue-wen Chen 0001, Jong Cheol Jeong
Bioinform.1
2009 Knowledge-guided inference of domain-domain interactions from incomplete protein-protein interaction networks
abstract
MOTIVATION: Protein-protein interactions (PPIs), though extremely valuable towards a better understanding of protein functions and cellular processes, do not provide any direct information about the regions/domains within the proteins that mediate the interaction. Most often, it is only a fraction of a protein that directly interacts with its biological partners. Thus, understanding interaction at the domain level is a critical step towards (i) thorough understanding of PPI networks; (ii) precise identification of binding sites; (iii) acquisition of insights into the causes of deleterious mutations at interaction sites; and (iv) most importantly, development of drugs to inhibit pathological protein interactions. In addition, knowledge derived from known domain-domain interactions (DDIs) can be used to understand binding interfaces, which in turn can help discover unknown PPIs. RESULTS: Here, we describe a novel method called K-GIDDI (knowledge-guided inference of DDIs) to narrow down the PPI sites to smaller regions/domains. K-GIDDI constructs an initial DDI network from cross-species PPI networks, and then expands the DDI network by inferring additional DDIs using a divide-and-conquer biclustering algorithm guided by Gene Ontology (GO) information, which identifies partial-complete bipartite sub-networks in the DDI network and makes them complete bipartite sub-networks by adding edges. Our results indicate that K-GIDDI can reliably predict DDIs. Most importantly, K-GIDDI's novel network expansion procedure allows prediction of DDIs that are otherwise not identifiable by methods that rely only on PPI data. AVAILABILITY: http://www.ittc.ku.edu/~xwchen/domainNetwork/ddinet.html
Xue-wen Chen 0001, Raja Jothi
Bioinform.2
2009 Identification of genes involved in the same pathways using a Hidden Markov Model-based approach
abstract
MOTIVATION: The sequencing of whole genomes from various species has provided us with a wealth of genetic information. To make use of the vast amounts of data available today it is necessary to devise computer-based analysis techniques. RESULTS: We propose a Hidden Markov Model (HMM) based algorithm to detect groups of genes functionally similar to a set of input genes from microarray expression data. A subset of experiments from a microarray is selected based on a set of related input genes. HMMs are trained from the input genes and a group of random gene input sets to provide significance estimates. Every gene in the microarray is scored using all HMMs and significant matches with the input genes are retained. We ran this algorithm on the life cycle of Drosophila microarray data set with KEGG pathways for cell cycle and translation factors as input data sets. Results show high functional similarity in resulting gene sets, increasing our biological insight into gene pathways and KEGG annotations. The algorithm performed very well compared to the Signature Algorithm and a purely correlation-based approach. AVAILABILITY: Java source codes and data sets are available at http://www.ittc.ku.edu/~xwchen/software.htm
Alexander Senf, Xue-wen Chen 0001
Bioinform.2
2009 Assessing reliability of protein-protein interactions by integrative analysis of data in model organisms
abstract
BACKGROUND: Protein-protein interactions play vital roles in nearly all cellular processes and are involved in the construction of biological pathways such as metabolic and signal transduction pathways. Although large-scale experiments have enabled the discovery of thousands of previously unknown linkages among proteins in many organisms, the high-throughput interaction data is often associated with high error rates. Since protein interaction networks have been utilized in numerous biological inferences, the inclusive experimental errors inevitably affect the quality of such prediction. Thus, it is essential to assess the quality of the protein interaction data. RESULTS: In this paper, a novel Bayesian network-based integrative framework is proposed to assess the reliability of protein-protein interactions. We develop a cross-species in silico model that assigns likelihood scores to individual protein pairs based on the information entirely extracted from model organisms. Our proposed approach integrates multiple microarray datasets and novel features derived from gene ontology. Furthermore, the confidence scores for cross-species protein mappings are explicitly incorporated into our model. Applying our model to predict protein interactions in the human genome, we are able to achieve 80% in sensitivity and 70% in specificity. Finally, we assess the overall quality of the experimentally determined yeast protein-protein interaction dataset. We observe that the more high-throughput experiments confirming an interaction, the higher the likelihood score, which confirms the effectiveness of our approach. CONCLUSION: This study demonstrates that model organisms certainly provide important information for protein-protein interaction inference and assessment. The proposed method is able to assess not only the overall quality of an interaction dataset, but also the quality of individual protein-protein interactions. We expect the method to continually improve as more high quality interaction data from more model organisms becomes available and is readily scalable to a genome-wide application.
Xiaotong Lin 0001, Xue-wen Chen 0001
BMC Bioinform.3
2008 Integrative Neural Network Approach for Protein Interaction Prediction from Heterogeneous Data
Xue-wen Chen 0001, Yong Hu 0002
ADMA1
2008 Protein-Protein Interaction Prediction and Assessment from Model Organisms
abstract
While high throughput technologies provide experimental tools to identify protein-protein interactions (PPIs), their false positive rate can be as high as about 50%. Thus, it is critical to provide computational methods that can assess the experimentally identified PPIs, detect the potential spurious interactions, and reliably predict new PPIs. In this paper, we propose a novel in-silico method for PPI prediction and assessment. Our model is based on new features extracted from different organisms and a Bayesian network that integrates heterogeneous data sources. We successfully apply the novel model to predict human PPIs from three model organisms Saccharomyces cerevisiae, C. elegans, and Drosophila melanogaster. For those PPIs with mapping information from the three model organisms, our model can reach 80% in sensitivity with a specificity of 70%.
Xiaotong Lin 0001, Xue-wen Chen 0001
BIBM3
2008 Gene Network Learning Using Regulated Dynamic Bayesian Network Methods
abstract
Dynamic Bayesian network (DBN) methods have shown great promise in regulatory network reconstruction because of their capability of modeling causality and cyclic networks, and handling data with noises found in biological experiments. However, they tend to produce relative high false positives and are not computationally efficient even for networks of moderate size. This paper presents a novel DBN-based approach to address these issues. For each node, a differential mutual information is used to select potential parents and a Bayesian scoring metric with a Dirichlet prior for regulation is applied to evaluate its parents. The proposed method is applied to recover a network structure from simulated data with higher accuracy and computational efficiency compared to DBNs with other scoring metrics. When applied to infer a cell cycle pathway of Saccharomyces cerevisiae using real time-series expression data, the proposed method is capable of identifying most gene interactions in the pathways.
Xiaotong Lin 0001, Xue-wen Chen 0001
ICMLA2
2008 FAST: a roc-based feature selection metric for small samples and imbalanced data classification problems
abstract
The class imbalance problem is encountered in a large number of practical applications of machine learning and data mining, for example, information retrieval and filtering, and the detection of credit card fraud. It has been widely realized that this imbalance raises issues that are either nonexistent or less severe compared to balanced class cases and often results in a classifier's suboptimal performance. This is even more true when the imbalanced data are also high dimensional. In such cases, feature selection methods are critical to achieve optimal performance. In this paper, we propose a new feature selection method, Feature Assessment by Sliding Thresholds (FAST), which is based on the area under a ROC curve generated by moving the decision boundary of a single feature classifier with thresholds placed using an even-bin distribution. FAST is compared to two commonly-used feature selection methods, correlation coefficient and RELevance In Estimating Features (RELIEF), for imbalanced data classification. The experimental results obtained on text mining, mass spectrometry, and microarray data sets showed that the proposed method outperformed both RELIEF and correlation methods on skewed data sets and was comparable on balanced data sets; when small number of features is preferred, the classification performance of the proposed method was significantly improved compared to correlation and RELIEF-based methods.
Xue-wen Chen 0001, Michael Wasikowski
KDD1
2008 A Bayesian approach to support vector machines for the binary classification
Jiangsheng Yu, Huilin Xiong, Wanling Qu, Xue-wen Chen 0001
Neurocomputing5
2008 Improving Bayesian Network Structure Learning with Mutual Information-Based Node Ordering in the K2 Algorithm
abstract
Structure learning of Bayesian networks is a well-researched but computationally hard task. We present an algorithm that integrates an information-theory-based approach and a scoring-function-based approach for learning structures of Bayesian networks. Our algorithm also makes use of basic Bayesian network concepts like d-separation and condition independence. We show that the proposed algorithm is capable of handling networks with a large number of variables. We present the applicability of the proposed algorithm on four standard network data sets and also compare its performance and computational efficiency with other standard structure-learning methods. The experimental results show that our method can efficiently and accurately identify complex network structures from data.
Xue-wen Chen 0001, Gopalakrishna Anantha, Xiaotong Lin 0001
IEEE Trans. Knowl. Data Eng.1
2007 Human Disease-Gene Classification with Integrative Sequence-Based and Topological Features of Protein-Protein Interaction Networks
abstract
The discovery of human genes that contribute to the appearance and growth of hereditary diseases is an important problem in bioinformatics research. Many techniques have been devised for classifying genes based on information from a variety of sources such as sequence and functional annotation. Recently, the use of topological information in protein-protein interaction networks has shown promise in disease- genes. In this paper, we develop a disease-gene classification system that integrates topological features of protein interaction networks with sequence- derived and other features, utilizing support vector machines for disease-gene classification. We identified several novel topological, sequence, and function-based features that can help to characterize hereditary disease-genes. We also found that using a more complex classifier can contribute to disease-gene classification. We validated our methods by selecting previously unclassified genes that were predicted with high probabilities as disease-genes, and searching for evidence in recent literature of their involvements in disease.
Aaron M. Smalter, Seak Fei Lei, Xue-wen Chen 0001
BIBM3
2007 Minimum reference set based feature selection for small sample classifications
abstract
We address feature selection problems for classification of small samples and high dimensionality. A practical example is microarray-based cancer classification problems, where sample size is typically less than 100 and number of features is several thousands or higher. One of the commonly used methods in addressing this problem is recursive feature elimination (RFE) method, which utilizes the generalization capability embedded in support vector machines and is thus suitable for small samples problems. We propose a novel method using minimum reference set (MRS) generated by the nearest neighbor rule. MRS is the set of minimum number of samples that correctly classify all the training samples. It is related to structural risk minimization principle and thus leads to good generalization. The proposed MRS based method is compared to RFE method with several real datasets, and experimental results show that the MRS method produces better classification performance.
Xue-wen Chen 0001, Jong Cheol Jeong
ICML1
2007 Enhanced recursive feature elimination
abstract
For classification with small training samples and high dimensionality, feature selection plays an important role in avoiding overfitting problems and improving classification performance. One of the commonly used feature selection methods for small samples problems is recursive feature elimination (RFE) method. RFE method utilizes the generalization capability embedded in support vector machines and is thus suitable for small samples problems. Despite its good performance, RFE tends to discard "weak" features, which may provide a significant improvement of performance when combined with other features. In this paper, we propose an enhanced recursive feature elimination (EnRFE) method for feature selection in small training sample classification. Our experimental results show that the proposed method outperforms the original RFE in terms of classification accuracy on various datasets.
Xue-wen Chen 0001, Jong Cheol Jeong
ICMLA1
2007 Normalized Linear Transform for Cross-Platform Microarray Data Integration
abstract
With microarray data being dramatically accumulated, integrating data from related studies represents a natural way to increase sample size so that more reliable statistical analysis may be performed. However, inherent variation among different microarray platforms makes the data integration not a trivial task. In this paper, we present a simple and effective integration scheme, called normalized linear transform (NLT), to combine data from different microarray platforms. The NLT scheme is compared with three other integration schemes for two tasks: classification analysis and gene marker selection. Our experiments demonstrate that the NLT scheme performs best in terms of classification accuracy under various classification settings, and leads to more biologically significant marker genes.
Huilin Xiong, Ya Zhang 0002, Xue-wen Chen 0001
ICMLA3
2007 Machine Learning and Its Applications to Biology
abstract
The term machine learning refers to a set of topics dealing with the creation and evaluation of algorithms that facilitate pattern recognition, classification, and prediction, based on models derived from existing data. Two facets of mechanization should be acknowledged when considering machine learning in broad terms. Firstly, it is intended that the classification and prediction tasks can be accomplished by a suitably programmed computing machine. That is, the product of machine learning is a classifier that can be feasibly used on available hardware. Secondly, it is intended that the creation of the classifier should itself be highly mechanized, and should not involve too much human input. This second facet is inevitably vague, but the basic objective is that the use of automatic algorithm construction methods can minimize the possibility that human biases could affect the selection and performance of the algorithm. Both the creation of the algorithm and its operation to classify objects or predict events are to be based on concrete, observable data. The history of relations between biology and the field of machine learning is long and complex. An early technique [1] for machine learning called the perceptron constituted an attempt to model actual neuronal behavior, and the field of artificial neural network (ANN) design emerged from this attempt. Early work on the analysis of translation initiation sequences [2] employed the perceptron to define criteria for start sites in Escherichia coli. Further artificial neural network architectures such as the adaptive resonance theory (ART) [3] and neocognitron [4] were inspired from the organization of the visual nervous system. In the intervening years, the flexibility of machine learning techniques has grown along with mathematical frameworks for measuring their reliability, and it is natural to hope that machine learning methods will improve the efficiency of discovery and understanding in the mounting volume and complexity of biological data. This tutorial is structured in four main components. Firstly, a brief section reviews definitions and mathematical prerequisites. Secondly, the field of supervised learning is described. Thirdly, methods of unsupervised learning are reviewed. Finally, a section reviews methods and examples as implemented in the open source data analysis and visualization language R (http://www.r-project.org).
Adi L. Tarca, Vincent Carey, Xue-wen Chen 0001, Roberto Romero, Sorin Draghici
PLoS Comput. Biol.3
2007 Data-Dependent Kernel Machines for Microarray Data Classification
abstract
One important application of gene expression analysis is to classify tissue samples according to their gene expression levels. Gene expression data are typically characterized by high dimensionality and small sample size, which makes the classification task quite challenging. In this paper, we present a data-dependent kernel for microarray data classification. This kernel function is engineered so that the class separability of the training data is maximized. A bootstrapping-based resampling scheme is introduced to reduce the possible training bias. The effectiveness of this adaptive kernel for microarray data classification is illustrated with a k-Nearest Neighbor (KNN) classifier. Our experimental study shows that the data-dependent kernel leads to a significant improvement in the accuracy of KNN classifiers. Furthermore, this kernel-based KNN scheme has been demonstrated to be competitive to, if not better than, more sophisticated classifiers such as Support Vector Machines (SVMs) and the Uncorrelated Linear Discriminant Analysis (ULDA) for classifying gene expression data.
Huilin Xiong, Ya Zhang 0002, Xue-wen Chen 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2006 Comparison of One-Class SVM and Two-Class SVM for Fold Recognition
Alexander Senf, Xue-wen Chen 0001, Anne Zhang
ICONIP (2)2
2006 An effective structure learning method for constructing gene networks
abstract
MOTIVATION: Bayesian network methods have shown promise in gene regulatory network reconstruction because of their capability of capturing causal relationships between genes and handling data with noises found in biological experiments. The problem of learning network structures, however, is NP hard. Consequently, heuristic methods such as hill climbing are used for structure learning. For networks of a moderate size, hill climbing methods are not computationally efficient. Furthermore, relatively low accuracy of the learned structures may be observed. The purpose of this article is to present a novel structure learning method for gene network discovery. RESULTS: In this paper, we present a novel structure learning method to reconstruct the underlying gene networks from the observational gene expression data. Unlike hill climbing approaches, the proposed method first constructs an undirected network based on mutual information between two nodes and then splits the structure into substructures. The directional orientations for the edges that connect two nodes are then obtained by optimizing a scoring function for each substructure. Our method is evaluated using two benchmark network datasets with known structures. The results show that the proposed method can identify networks that are close to the optimal structures. It outperforms hill climbing methods in terms of both computation time and predicted structure accuracy. We also apply the method to gene expression data measured during the yeast cycle and show the effectiveness of the proposed method for network reconstruction.
Xue-wen Chen 0001, Gopalakrishna Anantha, Xinkun Wang
Bioinform.1
2006 Kernel-based distance metric learning for microarray data classification
abstract
BACKGROUND: The most fundamental task using gene expression data in clinical oncology is to classify tissue samples according to their gene expression levels. Compared with traditional pattern classifications, gene expression-based data classification is typically characterized by high dimensionality and small sample size, which make the task quite challenging. RESULTS: In this paper, we present a modified K-nearest-neighbor (KNN) scheme, which is based on learning an adaptive distance metric in the data space, for cancer classification using microarray data. The distance metric, derived from the procedure of a data-dependent kernel optimization, can substantially increase the class separability of the data and, consequently, lead to a significant improvement in the performance of the KNN classifier. Intensive experiments show that the performance of the proposed kernel-based KNN scheme is competitive to those of some sophisticated classifiers such as support vector machines (SVMs) and the uncorrelated linear discriminant analysis (ULDA) in classifying the gene expression data. CONCLUSION: A novel distance metric is developed and incorporated into the KNN scheme for cancer classification. This metric can substantially increase the class separability of the data in the feature space and, hence, lead to a significant improvement in the performance of the KNN classifier.
Huilin Xiong, Xue-wen Chen 0001
BMC Bioinform.2
2006 Margin-based wrapper methods for gene identification using microarray
Xue-wen Chen 0001
Neurocomputing1
2006 Multi-class feature selection for texture classification
Xue-wen Chen 0001, Xiang-Yan Zeng, Deborah van Alphen
Pattern Recognit. Lett.1
2005 Optimized Kernel Machine for Cancer Classification Using Gene Expression Data
Huilin Xiong, Xue-wen Chen 0001
CIBCB2
2005 Pruning support vectors for imbalanced data classification
abstract
In many practical applications, learning from imbalanced data poses a significant challenge that is increasingly faced by the machine learning community. The class imbalance problem raises issues that are either nonexistent or less severe compared to balanced class cases. This paper presents a new method for imbalanced data classification. The proposed method is based on support vector machine classifiers and backward pruning technique. The experimental results obtained on two data sets demonstrate the effectiveness of the new algorithm.
Xue-wen Chen 0001, Byron Gerlach, David P. Casasent
IJCNN1
2005 Prediction of protein-protein interactions using random decision forest framework
abstract
MOTIVATION: Protein interactions are of biological interest because they orchestrate a number of cellular processes such as metabolic pathways and immunological recognition. Domains are the building blocks of proteins; therefore, proteins are assumed to interact as a result of their interacting domains. Many domain-based models for protein interaction prediction have been developed, and preliminary results have demonstrated their feasibility. Most of the existing domain-based methods, however, consider only single-domain pairs (one domain from one protein) and assume independence between domain-domain interactions. RESULTS: In this paper, we introduce a domain-based random forest of decision trees to infer protein interactions. Our proposed method is capable of exploring all possible domain interactions and making predictions based on all the protein domains. Experimental results on Saccharomyces cerevisiae dataset demonstrate that our approach can predict protein-protein interactions with higher sensitivity (79.78%) and specificity (64.38%) compared with that of the maximum likelihood approach. Furthermore, our model can be used to infer interactions not only for single-domain pairs but also for multiple domain pairs.
Xue-wen Chen 0001
Bioinform.1
2005 SMO-based pruning methods for sparse least squares support vector machines
abstract
Solutions of least squares support vector machines (LS-SVMs) are typically nonsparse. The sparseness is imposed by subsequently omitting data that introduce the smallest training errors and retraining the remaining data. Iterative retraining requires more intensive computations than training a single nonsparse LS-SVM. In this paper, we propose a new pruning algorithm for sparse LS-SVMs: the sequential minimal optimization (SMO) method is introduced into pruning process; in addition, instead of determining the pruning points by errors, we omit the data points that will introduce minimum changes to a dual objective function. This new criterion is computationally efficient. The effectiveness of the proposed method in terms of computational cost and classification accuracy is demonstrated by numerical experiments.
Xiang-Yan Zeng, Xue-wen Chen 0001
IEEE Trans. Neural Networks2
2004 Locally optimal search method for identifying genes from microarray data
abstract
The gene expression data obtained from microarrays have shown to be useful in cancer classification. DNA microarray data have extremely high dimensionality compared to the small number of available samples. An important step in microarray studies is to remove genes irrelevant to the learning problem and to select a small number of genes expressed in biological samples under specific conditions. In this paper, we propose a novel feature subset selection algorithm, partitional branch and bound (PBB) algorithm. This new algorithm is very efficient for selecting sets of genes in very high dimensional feature space. Two databases are considered: the colon cancer database and the leukemia database. Our experimental results show that the proposed algorithm yields a better subset of features than both forward selection algorithms and individual ranking methods in terms of a criterion function measuring class separability.
Xue-wen Chen 0001, Huaixin Chen
ICARCV1
2003 Confidence-clustering supervised radial basis function neural networks
abstract
We propose a novel technique for the design of radial basis function (RBF) neural networks (NNs). To select various RBF parameters, the class membership information of training samples is utilized to produce a new cluster classes. This allows us to control performance as desired and approximate Neyman-Pearson classification. We show that by properly choosing the desired output neuron levels, then the RBF hidden to output layer performs Fisher discrimination analysis, and the full system performs a nonlinear Fisher analysis. Data on an agricultural product inspection problem and on synthetic data confirm the effectiveness of these methods.
David P. Casasent, Xue-wen Chen 0001
IJCNN2
2003 Radial basis function neural networks for nonlinear Fisher discrimination and Neyman-Pearson classification
David P. Casasent, Xue-wen Chen 0001
Neural Networks2
2003 New training strategies for RBF neural networks for X-ray agricultural product inspection
David P. Casasent, Xue-wen Chen 0001
Pattern Recognit.2
2003 Facial expression recognition: A clustering-based approach
Xue-wen Chen 0001, Thomas S. Huang
Pattern Recognit. Lett.1