Yifeng Li 0001

dblp:65/2432-1 · DBLP profile ↗
← Back
36ranked-venue papers
22as first author
11since 2021 · last 2026
0000-0002-4873-6928ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 11 first-author · 7 since 2021Artificial intelligence and machine learning · 17 · 11 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reply to "Letter to the editor: methodological considerations in the benchmarking of AI-based protein-aptamer complex prediction"
abstract
We thank MD Kazim Okan Dolu for their interest in our published work and for the opportunity to clarify several points regarding software versioning, benchmark independence, secondary-structure generation, binding-free-energy terminology, and negative-control design. We acknowledge that several descriptions in the original manuscript could be stated more precisely. These clarifications improve the transparency and reproducibility of the study, but do not affect the reported analyses or main conclusions.
Jiani Zhao, Kha Tram, Hongbin Yan, Yifeng Li 0001
Briefings Bioinform.4
2026 Comprehensive evaluation of artificial intelligence-empowered approaches for protein-aptamer complex prediction
abstract
Drug discovery is a time-consuming, expensive, and high-risk process. Recent advances in artificial intelligence (AI) have enabled major breakthroughs in small-molecule and protein therapeutics. However, AI-driven design of aptamer drugs remains largely unexplored. Aptamers are short (15-100 nt) single-stranded DNAs or RNAs that exhibit high binding affinity, high specificity, and low immunogenicity, making them promising candidates for disease (such as cancer) therapeutics. Compared with protein-ligand or protein-protein systems, protein-aptamer complexes are under-represented in public structural databases, and aptamers themselves are highly flexible and relatively large molecules. These characteristics present distinct challenges for AI-based structural modeling. Here, we systematically evaluate recent AI frameworks, including AlphaFold3, Chai-1, Boltz-2, and RoseTTAFold2NA, along with a template-based approach, in predicting protein-aptamer complex structures and estimating binding free energies. We establish an independent benchmark to assess their performance in structural accuracy, stability, and energetic consistency. This study provides a foundation for the application of AI in aptamer drug design and offers a reference framework for future research in nucleic-acid therapeutics and biomolecular modeling.
Jiani Zhao, Kha Tram, Hongbin Yan, Yifeng Li 0001
Briefings Bioinform.4
2025 SMORE-DRL: Scalable Multi-Objective Robust and Efficient Deep Reinforcement Learning for Molecular Optimization
abstract
The adoption of AI techniques within the domain of drug design provides an opportunity for systematic and efficient exploration of the vast chemical search space. In recent years, advancements in this domain have been driven by AI frameworks, including deep reinforcement learning (DRL). However, the scalability and performance of existing DRL methodologies are constrained by prolonged training periods and inefficient sample data utilization. Furthermore, generalization capabilities of these models have not been fully investigated. To overcome these limitations, we take a multi-objective optimization perspective and introduce SMORE-DRL, a fragment and transformerbased multi-objective DRL architecture for the optimization of molecules across multiple pharmacological properties, including binding affinity to both single and dual cancer protein targets. Our approach involves pretraining a transformer-encoder model on molecules encoded by a novel hybrid fragment-SMILES representation method. Fine-tuning is performed through a novel gradient-alignment-based DRL, where lead molecules are optimized by selecting and replacing their fragments with alternatives from a fragment dictionary, ultimately resulting in more desirable drug candidates. Our findings indicate that SMOREDRL is superior to current models for lead optimization in terms of quality, efficiency, scalability, and robustness. Furthermore, SMORE-DRL demonstrates the capability of generalizing its optimization process to lead molecules that are not present during the pretraining or fine-tuning phases.
Aws Al Jumaily, Nicholas Aksamit, Yage Zhang, Mohammad Sajjad Ghaemi, Jinqiang Hou, Hsu Kiang Ooi, Yifeng Li 0001
BIBM7
2024 Integrating transformers and many-objective optimization for drug design
abstract
BACKGROUND: Drug design is a challenging and important task that requires the generation of novel and effective molecules that can bind to specific protein targets. Artificial intelligence algorithms have recently showed promising potential to expedite the drug design process. However, existing methods adopt multi-objective approaches which limits the number of objectives. RESULTS: In this paper, we expand this thread of research from the many-objective perspective, by proposing a novel framework that integrates a latent Transformer-based model for molecular generation, with a drug design system that incorporates absorption, distribution, metabolism, excretion, and toxicity prediction, molecular docking, and many-objective metaheuristics. We compared the performance of two latent Transformer models (ReLSO and FragNet) on a molecular generation task and show that ReLSO outperforms FragNet in terms of reconstruction and latent space organization. We then explored six different many-objective metaheuristics based on evolutionary algorithms and particle swarm optimization on a drug design task involving potential drug candidates to human lysophosphatidic acid receptor 1, a cancer-related protein target. CONCLUSION: We show that multi-objective evolutionary algorithm based on dominance and decomposition performs the best in terms of finding molecules that satisfy many objectives, such as high binding affinity and low toxicity, and high drug-likeness. Our framework demonstrates the potential of combining Transformers and many-objective computational intelligence for drug design.
Nicholas Aksamit, Jinqiang Hou, Yifeng Li 0001, Beatrice M. Ombuki-Berman
BMC Bioinform.3
2024 Hybrid fragment-SMILES tokenization for ADMET prediction in drug discovery
abstract
BACKGROUND: Drug discovery and development is the extremely costly and time-consuming process of identifying new molecules that can interact with a biomarker target to interrupt the disease pathway of interest. In addition to binding the target, a drug candidate needs to satisfy multiple properties affecting absorption, distribution, metabolism, excretion, and toxicity (ADMET). Artificial intelligence approaches provide an opportunity to improve each step of the drug discovery and development process, in which the first question faced by us is how a molecule can be informatively represented such that the in-silico solutions are optimized. RESULTS: This study introduces a novel hybrid SMILES-fragment tokenization method, coupled with two pre-training strategies, utilizing a Transformer-based model. We investigate the efficacy of hybrid tokenization in improving the performance of ADMET prediction tasks. Our approach leverages MTL-BERT, an encoder-only Transformer model that achieves state-of-the-art ADMET predictions, and contrasts the standard SMILES tokenization with our hybrid method across a spectrum of fragment library cutoffs. CONCLUSION: The findings reveal that while an excess of fragments can impede performance, using hybrid tokenization with high frequency fragments enhances results beyond the base SMILES tokenization. This advancement underscores the potential of integrating fragment- and character-level molecular features within the training of Transformer models for ADMET property prediction.
Nicholas Aksamit, Alain B. Tchagang, Yifeng Li 0001, Beatrice M. Ombuki-Berman
BMC Bioinform.3
2022 Exploring Multi-Objective Deep Reinforcement Learning Methods for Drug Design
abstract
Drug design and optimization are complex tasks that require strategically efficient exploration of the extremely vast search space. Various fragmentation strategies have been presented in the literature to reduce the complexity of the molecular search space. From the optimization perspective, drug design can be viewed as a multi-objective optimization process. Deep reinforcement learning (DRL) frameworks have displayed promising performances in this field. However, lengthy training periods and inefficient use of sample data limit the scalability of the current frameworks. In this paper, we (1) review the fundamental concepts of deep or multi-objective RL methods and their applications in molecular design, (2) investigate the performance of a recent multi-objective DRL-based and fragment-based drug design framework, named DeepFMPO, in a real application by integrating protein-ligand docking affinity score, and (3) compare this method with a single-objective variant. Through experiments, we find that the DeepFMPO framework (with docking score) can achieve limited success, however, it is incredibly unstable. Our findings encourage further exploration and improvement. Possible sources of the framework's instability and suggestions of further modifications to stabilize the framework are examined.
Aws Al Jumaily, Muhetaer Mukaidaisi, Andrew Vu, Alain B. Tchagang, Yifeng Li 0001
CIBCB5
2022 Transfer Learning-enabled Modelling Framework for Digital Twin
abstract
Recently the machine learning-enabled modeling technology has become a powerful tool to develop data-driven models for explaining, predicting, and describing system behaviors. In particular, it has become a key tool for developing data-driven models for emerging digital twin development which demands the living models for simulating system behaviors. However, such data-driven models carry a fatal deficiency: once the operational environments changed, the model may hardly work well or even becomes useless. This paper attempts to address this issue by proposing to apply transfer learning techniques to develop lifetime robust models for real-world applications. After laying out problems and the reasons of model performance degradation, this paper presents a framework for developing lifetime predictive models for digital twin. A show case from our on-going research project along with the preliminary results demonstrates the feasibility and usefulness of the proposed predictive modeling methods.
Chunsheng Yang, Yifeng Li 0001, Zheng Liu 0002, Min Liao
CSCWD2
2022 Correlation Encoder-Decoder Model for Text Generation
abstract
Text generation is crucial for many applications in natural language processing. With the prevalence of deep learning, the encoder-decoder architecture is dominantly adopted for this task. Accurately encoding the source information is of key importance to text generation, because the target text can be generated only when accurate and complete source information is captured by the encoder and fed into the decoder. However, most existing approaches fail to effectively encode and learn the entire source information, as some features are easy to be missed along with the encoding procedures of the encoder. Similar problems also confuse the implementation of the decoder. How to reduce the problem of information loss in the encoder-decoder model is critical for text generation. To address this issue, we propose a novel correlation encoder-decoder model, which optimizes both the encoder and the decoder to reduce the problem of information loss by enforcing them to minimize the differences between hierarchical layers by maximizing the mutual information. Experimental results on two benchmark datasets demonstrate that the proposed model substantially outperforms the existing state-of-the-art methods. Our source code is publicly available on GitHub1.
Xu Zhang 0053, Yifeng Li 0001, Xueping Peng, Xinxiao Qiao, Wenpeng Lu
IJCNN2
2021 Adversarial Deep Evolutionary Learning for Drug Design
abstract
The design of a new therapeutic agent is a time-consuming and expensive process. The rise of machine intelligence provides a grand opportunity of expeditiously discovering novel drug candidates through smart search in the vast molecular structural space. In this paper, we propose a new approach called adversarial deep evolutionary learning (ADEL) to search for novel molecules in the latent space of an adversarial generative model and keep improving the latent representation space. In ADEL, a custom-made adversarial autoencoder (AAE) model is developed and trained under a deep evolutionary learning (DEL) process. This involves an initial training of the AAE model, followed by an integration of multi-objective evolutionary optimization in the continuous latent representation space of the AAE rather than the discrete structural space of molecules. By using the AAE, an arbitrary distribution can be provided to the training of AAE such that the latent representation space is set to that distribution. This allows for a starting latent space from which new samples can be produced. Throughout the process of learning, new samples of high-quality are generated after each iteration of training and then added back into the full dataset. Therefore, allowing for a more comprehensive procedure of understanding the data structure. This combination of evolving data and continuous learning not only enables improvement in the generative model, but the data as well. By comparing ADEL to the previous work in DEL, we see that ADEL can obtain better property distributions.
Sheriff Abouchekeir, Alain B. Tchagang, Yifeng Li 0001
CIBCB3
2021 Sentence Semantic Matching with Hierarchical CNN Based on Dimension-augmented Representation
abstract
As a fundamental task in natural language processing, sentence semantic matching (SSM) is critical yet challenging due to difficulties in learning expressive sentence representation while capturing complex interactions between sentences. Recent work has shown the great potential of deep neural models in improving the performance of SSM task. However, existing work usually employs recurrent neural networks (RNNs) or 1D (one-dimensional) convolutional neural networks (CNNs) to learn sentence representation, leading to limited performance improvement. Benefiting from the multi-dimensional structure, 2D convolutional neural networks are expected to be more powerful to learn expressive sentence representation by capturing the implicit inter-sentence interactions and thus can further improve the performance of SSM. To this end, in this paper, we propose a novel sentence semantic matching model named Hierarchical CNN based on Dimension-augmented Representation (HiDR). In HiDR, first, bidirectional long short-term memory networks (LSTMs) are utilized to generate dimension-augmented representation for each of the input sentences; then, a hierarchical 2D CNN is devised to learn sentence representation while capturing the inter-sentence interactions, followed by a prediction layer based on sigmoid function to output the matching degree between sentences. To evaluate the performance of our proposed model, we conducted extensive experiments on two public real-world data sets. The empirical results show that HiDR has achieved remarkable results, which demonstrates either better or comparable performance w.r.t. BERT-based models.
Rui Yu 0005, Wenpeng Lu, Yifeng Li 0001, Jiguo Yu, Guoqiang Zhang 0003, Xu Zhang 0053
IJCNN3
2021 Chinese Semantic Matching with Multi-granularity Alignment and Feature Fusion
abstract
Chinese semantic matching is a fundamental task in natural language processing, which is critical and yet challenging for a series of downstream tasks. Although recent work on text representation learning has shown its potential in improving the performance on semantic matching, relatively limited work has been done on exploring the relevant interactive information between two granularity of Chinese text, i.e., character and word. Existing methods usually focus on capturing the interactive features from single granularity, which lead to inefficient text representation. Also, they typically fail to consider the fusion of features from different granularity. As a result, they only achieve limited performance improvement. This paper proposes a novel Chinese semantic matching model based on multi-granularity alignment and feature fusion (MAFFo). To be specific, we first encode the texts from different granularity, which are further handled with soft-alignment attention mechanism to extract relevant interactive information between texts on different granularity. In addition, we devise a feature fusion structure to merge the features from different granularity to generate an ideal representation for the pair of input text sequences, followed by a sigmoid function to judge the semantic matching degree. Extensive experiments on the publicly available dataset BQ demonstrate that our model can effectively improve the performance of semantic matching task and achieve comparable performance with BERT-based methods.
Wenpeng Lu, Yifeng Li 0001, Jiguo Yu, Ping Jian, Xu Zhang 0053
IJCNN3
2020 Intra-Correlation Encoding for Chinese Sentence Intention Matching
abstract
Sentence intention matching is vital for natural language understanding.Especially for Chinese sentence intention matching task, due to the ambiguity of Chinese words, semantic missing or semantic confusion are more likely to occur in the encoding process.Although the existing methods have enriched text representation through pre-trained word embedding to solve this problem, due to the particularity of Chinese text, different granularities of pre-trained word embedding will affect the semantic description of a piece of text.In this paper, we propose an effective approach that combines charactergranularity and word-granularity features to perform sentence intention matching, and we utilize soft alignment attention to enhance the local information of sentences on the corresponding levels.The proposed method can capture sentence feature information from multiple perspectives and correlation information between different levels of sentences.By evaluating on BQ and LCQMC datasets, our model has achieved remarkable results, and demonstrates better or comparable performance with BERT-based models.
Xu Zhang 0053, Yifeng Li 0001, Wenpeng Lu, Ping Jian, Guoqiang Zhang 0003
COLING2
2020 Anomaly Detection Based on Unsupervised Disentangled Representation Learning in Combination with Manifold Learning
abstract
Identifying anomalous samples from highly complex and unstructured data is a crucial but challenging task in a variety of intelligent systems. In this paper, we present a novel deep anomaly detection framework named AnoDM (standing for Anomaly detection based on unsupervised Disentangled representation learning and Manifold learning). The disentanglement learning is currently implemented by β-VAE for automatically discovering interpretable factorized latent representations in a completely unsupervised manner. The manifold learning is realized by t-SNE for projecting the latent representations to a 2D map. We define a new anomaly score function by combining β-VAE's reconstruction error in the raw feature space and local density estimation in the t-SNE space. AnoDM was evaluated on both image and time-series data and achieved better results than models that use just one of the two measures and other deep learning methods.
Iluju Kiringa, Tet Hin Yeap, Xiaodan Zhu 0001, Yifeng Li 0001
IJCNN5
2020 Capsule Deep Generative Model That Forms Parse Trees
abstract
Supervised capsule networks are theoretically advantageous over convolutional neural networks, because they aim to model a range of transformations of local physical or abstract objects and part-whole relationships among them. However, it remains unclear how to use the concept of capsules in deep generative models. In this study, to address this challenge, we present a statistical modelling of capsules in deep generative models where distributions are formulated in the exponential family. The major contribution of this unsupervised method is that parse trees as representations of part-whole relationships can be dynamically learned from the data.
Yifeng Li 0001, Xiaodan Zhu 0001, Richard Naud, Pengcheng Xi
IJCNN1
2019 Capsule Generative Models
Yifeng Li 0001, Xiaodan Zhu 0001
ICANN (1)1
2019 Personalized prediction of genes with tumor-causing somatic mutations based on multi-modal deep Boltzmann machine
Yifeng Li 0001, François Fauteux, Jinfeng Zou, André Nantel, Youlian Pan
Neurocomputing1
2019 Multiclass Nonnegative Matrix Factorization for Comprehensive Feature Pattern Discovery
abstract
In this big data era, interpretable machine learning models are strongly demanded for the comprehensive analytics of large-scale multiclass data. Characterizing all features from such data is a key but challenging step to understand the complexity. However, existing feature selection methods do not meet this need. In this paper, to address this problem, we propose a Bayesian multiclass nonnegative matrix factorization (MC-NMF) model with structured sparsity that is able to discover ubiquitous and class-specific features. Variational update rules were derived for efficient decomposition. In order to relieve the need of model selection and stably describe feature patterns, we further propose MC-NMF with stability selection, an ensemble method that collectively detects feature patterns from many runs of MC-NMF using different hyperparameter values and training subsets. We assessed our models on both simulated count data and multitumor ribonucleic acid-seq data. The experiments revealed that our models were able to recover predefined feature patterns from the simulated data and identify biologically meaningful patterns from the pan-cancer data.
Yifeng Li 0001, Youlian Pan, Ziying Liu
IEEE Trans. Neural Networks Learn. Syst.1
2018 Exponential Family Restricted Boltzmann Machines and Annealed Importance Sampling
abstract
In this paper, we investigate restricted Boltzmann machines (RBMs) from the exponential family perspective, en-abling the visible units to follow any suitable distributions from the exponential family. We derive a unified view to compute the free energy function for exponential family RBMs (exp-RBMs). Based on that, annealed important sampling (AIS) is generalized to the entire exponential family, allowing for estimating the log-partition function and log-likelihood. Our experiments on a document processing task demonstrate that the generalized free energy functions and AIS estimation perform well in helping capture useful knowledge from the data; the estimated log-partition functions are stable. The appropriate instances of exp-RBMs can generate novel and meaningful samples and can be applied to classification tasks.
Yifeng Li 0001, Xiaodan Zhu 0001
IJCNN1
2018 A review on machine learning principles for multi-view biological data integration
abstract
Driven by high-throughput sequencing techniques, modern genomic and clinical studies are in a strong need of integrative machine learning models for better use of vast volumes of heterogeneous information in the deep understanding of biological systems and the development of predictive models. How data from multiple sources (called multi-view data) are incorporated in a learning system is a key step for successful analysis. In this article, we provide a comprehensive review on omics and clinical data integration techniques, from a machine learning perspective, for various analyses such as prediction, clustering, dimension reduction and association. We shall show that Bayesian models are able to use prior information and model measurements with various distributions; tree-based methods can either build a tree with all features or collectively make a final decision based on trees learned from each view; kernel methods fuse the similarity matrices learned from individual views together for a final similarity matrix or learning model; network-based fusion methods are capable of inferring direct and indirect associations in a heterogeneous network; matrix factorization models have potential to learn interactions among features from different views; and a range of deep neural networks can be integrated in multi-modal learning for capturing the complex mechanism of biological systems.
Yifeng Li 0001, Fang-Xiang Wu, Alioune Ngom
Briefings Bioinform.1
2018 Genome-wide prediction of cis-regulatory regions using supervised deep learning methods
abstract
BACKGROUND: In the human genome, 98% of DNA sequences are non-protein-coding regions that were previously disregarded as junk DNA. In fact, non-coding regions host a variety of cis-regulatory regions which precisely control the expression of genes. Thus, Identifying active cis-regulatory regions in the human genome is critical for understanding gene regulation and assessing the impact of genetic variation on phenotype. The developments of high-throughput sequencing and machine learning technologies make it possible to predict cis-regulatory regions genome wide. RESULTS: Based on rich data resources such as the Encyclopedia of DNA Elements (ENCODE) and the Functional Annotation of the Mammalian Genome (FANTOM) projects, we introduce DECRES based on supervised deep learning approaches for the identification of enhancer and promoter regions in the human genome. Due to their ability to discover patterns in large and complex data, the introduction of deep learning methods enables a significant advance in our knowledge of the genomic locations of cis-regulatory regions. Using models for well-characterized cell lines, we identify key experimental features that contribute to the predictive performance. Applying DECRES, we delineate locations of 300,000 candidate enhancers genome wide (6.8% of the genome, of which 40,000 are supported by bidirectional transcription data), and 26,000 candidate promoters (0.6% of the genome). CONCLUSION: The predicted annotations of cis-regulatory regions will provide broad utility for genome interpretation from functional genomics to clinical applications. The DECRES model demonstrates potentials of deep learning technologies when combined with high-throughput sequencing data, and inspires the development of other advanced neural network models for further improvement of genome annotations.
Yifeng Li 0001, Wenqiang Shi, Wyeth W. Wasserman
BMC Bioinform.1
2016 The Max-Min High-Order Dynamic Bayesian Network for Learning Gene Regulatory Networks with Time-Delayed Regulations
abstract
Accurately reconstructing gene regulatory network (GRN) from gene expression data is a challenging task in systems biology. Although some progresses have been made, the performance of GRN reconstruction still has much room for improvement. Because many regulatory events are asynchronous, learning gene interactions with multiple time delays is an effective way to improve the accuracy of GRN reconstruction. Here, we propose a new approach, called Max-Min high-order dynamic Bayesian network (MMHO-DBN) by extending the Max-Min hill-climbing Bayesian network technique originally devised for learning a Bayesian network's structure from static data. Our MMHO-DBN can explicitly model the time lags between regulators and targets in an efficient manner. It first uses constraint-based ideas to limit the space of potential structures, and then applies search-and-score ideas to search for an optimal HO-DBN structure. The performance of MMHO-DBN to GRN reconstruction was evaluated using both synthetic and real gene expression time-series data. Results show that MMHO-DBN is more accurate than current time-delayed GRN learning methods, and has an intermediate computing performance. Furthermore, it is able to learn long time-delayed relationships between genes. We applied sensitivity analysis on our model to study the performance variation along different parameter settings. The result provides hints on the setting of parameters of MMHO-DBN.
Yifeng Li 0001, Haifen Chen, Jie Zheng 0002, Alioune Ngom
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 Data integration in machine learning
abstract
Modern data generated in many fields are in a strong need of integrative machine learning models in order to better make use of heterogeneous information in decision making and knowledge discovery. How data from multiple sources are incorporated in a learning system is key step for a successful analysis. In this paper, we provide a comprehensive review on data integration techniques from a machine learning perspective.
Yifeng Li 0001, Alioune Ngom
BIBM1
2015 Deep Feature Selection: Theory and Application to Identify Enhancers and Promoters
Yifeng Li 0001, Wyeth W. Wasserman
RECOMB1
2015 Pattern classification using a new border identification paradigm: The nearest border technique
Yifeng Li 0001, B. John Oommen, Alioune Ngom, Luis Rueda 0001
Neurocomputing1
2014 A decomposition method for large-scale sparse coding in representation learning
abstract
In representation learning, sparse representation is a parsimonious principle that a sample can be approximated by a sparse superposition of dictionary atoms. Sparse coding is the core of this technique. Since the dictionary is often redundant, the dictionary size can be very large. Many optimization methods have been proposed in the literature for sparse coding. However, the efficiency of the optimization for a tremendous number of dictionary atoms is still a bottleneck. In this paper, we propose to use decomposition method for large-scale sparse coding models. Our experimental results show that our method is very efficient.
Yifeng Li 0001, Richard J. Caron, Alioune Ngom
IJCNN1
2014 Versatile sparse matrix factorization: Theory and applications
Yifeng Li 0001, Alioune Ngom
Neurocomputing1
2013 The max-min high-order dynamic Bayesian network learning for identifying gene regulatory networks from time-series microarray data
abstract
We propose a new high-order dynamic Bayesian network (HO-DBN) learning approach, called Max-Min High-Order DBN (MMHO-DBN), for discrete time-series data. MMHO-DBN explicitly models the time lags between parents and target in an efficient manner. It extends the Max-Min Hill-Climbing Bayesian network (MMHC-BN) technique which was originally devised for learning a BN's structure from static data. Both Max-Min approaches are hybrid local learning methods which fuse concepts from both constraint-based Bayesian techniques and search-and-score Bayesian methods. The MMHO-DBN first uses constraint-based ideas to limit the space of potential structure and then applies search-and-score ideas to search for an optimal HO-DBN structure. We evaluated the ability of our MMHO-DBN approach to identify genetic regulatory networks (GRN's) from gene expression time-series data. Preliminary results on artificial and real gene expression time-series are encouraging and show that it is able to learn (long) time-delayed relationships between genes, and faster than current HO-DBN learning methods.
Yifeng Li 0001, Alioune Ngom
CIBCB1
2013 Classification approach based on non-negative least squares
Yifeng Li 0001, Alioune Ngom
Neurocomputing1
2013 Nonnegative Least-Squares Methods for the Classification of High-Dimensional Biological Data
abstract
Microarray data can be used to detect diseases and predict responses to therapies through classification models. However, the high dimensionality and low sample size of such data result in many computational problems such as reduced prediction accuracy and slow classification speed. In this paper, we propose a novel family of nonnegative least-squares classifiers for high-dimensional microarray gene expression and comparative genomic hybridization data. Our approaches are based on combining the advantages of using local learning, transductive learning, and ensemble learning, for better prediction performance. To study the performances of our methods, we performed computational experiments on 17 well-known data sets with diverse characteristics. We have also performed statistical comparisons with many classification techniques including the well-performing SVM approach and two related but recent methods proposed in literature. Experimental results show that our approaches are faster and achieve generally a better prediction performance over compared methods.
Yifeng Li 0001, Alioune Ngom
IEEE ACM Trans. Comput. Biol. Bioinform.1
2012 Fast sparse representation approaches for the classification of high-dimensional biological data
abstract
Classifying genomic and proteomic data is very important to predict diseases in a very early stage and investigate signaling pathways. However, this poses many computationally challenging problems, such as curse of dimensionality, noise, redundancy and so on. The principle of sparse representation has been applied to analyzing high-dimensional biological data within the frameworks of clustering, classification, and dimension reduction approaches. However, the existing sparse representation approaches are either inefficient or have the difficulty of kernelization. In this paper, we propose fast active-set-based sparse coding approach and a dictionary learning framework for classifying high-dimensional biological data. We show that they can be easily kernelized. Experimental results show that our approaches are very efficient, and satisfactory accuracy can be obtained compared with existing approaches.
Yifeng Li 0001, Alioune Ngom
BIBM1
2012 A new Kernel non-negative matrix factorization and its application in microarray data analysis
abstract
Non-negative factorization (NMF) has been a popular machine learning method for analyzing microarray data. Kernel approaches can capture more non-linear discriminative features than linear ones. In this paper, we propose a novel kernel NMF (KNMF) approach for feature extraction and classification of microarray data. Our approach is also generalized to kernel high-order NMF (HONMF). Extensive experiments on eight microarray datasets show that our approach generally outperforms the traditional NMF and existing KNMFs. Preliminary experiment on a high-order microarray data shows that our KHONMF is a promising approach given a suitable kernel function.
Yifeng Li 0001, Alioune Ngom
CIBCB1
2012 Fast Kernel Sparse Representation Approaches for Classification
abstract
Sparse representation involves two relevant procedures - sparse coding and dictionary learning. Learning a dictionary from data provides a concise knowledge representation. Learning a dictionary in a higher feature space might allow a better representation of a signal. However, it is usually computationally expensive to learn a dictionary if the numbers of training data and(or) dimensions are very large using existing algorithms. In this paper, we propose a kernel dictionary learning framework for three models. We reveal that the optimization has dimension-free and parallel properties. We devise fast active-set algorithms for this framework. We investigated their performance on classification. Experimental results show that our kernel sparse representation approaches can obtain better accuracy than their linear counterparts. Furthermore, our active-set algorithms are faster than the existing interior-point and proximal algorithms.
Yifeng Li 0001, Alioune Ngom
ICDM1
2012 Supervised Dictionary Learning via Non-negative Matrix Factorization for Classification
abstract
Sparse representation (SR) has been being applied as a state-of-the-art machine learning approach. Sparse representation classification (SRC1) approaches based on l1norm regularization and non-negative-least-squares (NNLS) classification approach based on non-negativity have been proposed to be powerful and robust. However, these approaches are extremely slow when the size of training samples is very large, because both of them use the whole training set as dictionary. In this paper, we briefly survey the existing SR techniques for classification, and then propose a fast approach which uses non-negative matrix factorization as supervised dictionary learning method and NNLS as non-negative sparse coding method. Experiment shows that our approach can obtain comparable accuracy with the benchmark approaches and can dramatically speed up the computation particularly in the case of large sample size and many classes.
Yifeng Li 0001, Alioune Ngom
ICMLA (1)1
2010 Non-negative matrix and tensor factorization based classification of clinical microarray gene expression data
abstract
Non-negative information can benefit the analysis of microarray data. This paper investigates the classification performance of non-negative matrix factorization (NMF) over gene-sample data. We also extends it to higher-order version for classification of clinical time-series data represented by tensor. Experiments show that NMF and the higher-order NMF can achieve at least comparable prediction performance.
Yifeng Li 0001, Alioune Ngom
BIBM1
2010 Alignment versus variation methods for clustering microarray time-series data
abstract
In the past few years, it has been shown that traditional clustering methods do not necessarily perform well on time-series data because of the temporal relationships involved in such data - this makes it a particularly difficult problem. In this paper, we compare two clustering methods that have been introduced recently, especially for gene expression time-series data, namely, multiple-alignment (MA) clustering and variation-based co-expression detection (VCD) clustering approaches. Both approaches are based on a transformation of the data that takes into account the temporal relationships, and have been shown to effectively detect groups of co-expressed genes. We investigate the performances of the MA and VCD approaches on two microarray time-series data sets and discuss their strengths and weaknesses. Our experiments show the superior accuracy of MA over VCD when finding groups of co-expressed genes.
Numanul Subhani, Yifeng Li 0001, Alioune Ngom, Luis Rueda 0001
IEEE Congress on Evolutionary Computation2
2010 Missing value imputation methods for gene-sample-time microarray data analysis
abstract
With the recent advances in microarray technology, the expression levels of genes with respect to the samples can be monitored synchronically over a series of time points. Such three-dimensional microarray data, termed gene-sample-time microarray data or GST data for short, may contain missing values. Current microarray analysis methods require complete data sets, and thus, either each row, column or tube containing missing values must be removed from the original GST data, or these missing values must be estimated before analysis. Imputation of missing values is, however, more recommended than removal of data in order to increase the effectiveness of analysis algorithms. In this paper, we extend automated imputation methods, devised for two-dimensional microarray data, to GST data. We implemented imputation methods for GST data based on Singular Value Decomposition (3SVDimpute), K-Nearest Neighbor (3KNNimpute), and gene and sample average methods (3Aimpute), and show that methods based on KNN yield the best results with the lowest normalized root mean squared error.
Yifeng Li 0001, Alioune Ngom, Luis Rueda 0001
CIBCB1