Wai-Ki Ching

dblp:98/4717 · DBLP profile ↗
← Back
65ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0002-5785-3210ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 18 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Finite-Queue Bitcoin Transaction-Confirmation Time Modelling
Jing Yee Tan, Yuqiao Zhao, Wai-Ki Ching
ICBC3
2026 Prognostic biomarker discovery via a connected network-constrained Cox proportional hazards model
abstract
Abstract Biomarker discovery in biomedical sciences can be framed as feature selection in machine learning [1]. However, existing methods often overlook gene co-localization within regulatory interaction networks, leading to the identification of isolated biomarkers with limited biological interpretability [2]. Here, we present the Connected Network-regularized Cox proportional hazards model (CNet-Cox), which incorporates network connectivity constraints into sparse regularization to identify prognostic biomarkers for breast cancer (BRCA) on the discovery dataset from TCGA (1,092 patients), while explicitly accounting for patient survival time. CNet-Cox reveals the network structures of prognostic genes, evaluated in the internal validation dataset with a concordance index of 0.913, surpassing traditional regularized Cox methods. CNet-Cox shifts biomarker recognition from isolated to connected features within biomolecular networks and offers new biological insights. Furthermore, we established a six-gene BRCA prognostic risk scoring (PRS) metric and validated its robustness across six independent external validation datasets comprising 1,829 patients, and one spatial transcriptomic dataset containing 4,992 spots. The PRS score consistently demonstrated superior performance in patient/sample stratification across extensive and diverse validation datasets. Overall, our comprehensive downstream analyses underscore that CNet-Cox offers a novel approach for embedding network topology into feature selection, enabling the systematic discovery of key connected prognostic biomarkers. This significantly advances early detection and prognosis prediction, facilitating precision medicine for BRCA. References 1. Li L, Liu Z P. “Biomarker discovery from high-throughput data by connected network-constrained support vector machine.” Expert Systems with Applications 2023; 226: 120179. 2. Hartman E, Scott A M, Karlsson C, et al. “Interpreting biologically informed neural networks for enhanced proteomic biomarker discovery and pathway analysis.” Nature Communications 2023; 14(1): 5359.
Wai-Ki Ching, Zhi-Ping Liu
Briefings Bioinform.2
2026 SpaConTDS: A multimodal contrastive learning framework for identifying spatial domains by applying tuple disturbing strategy
abstract
The rational utilization of multimodal spatial transcriptomics (ST) data enables accurate identification of spatial domains, which is essential for investigating cellular structure and functions. In this study, we proposed SpaConTDS, a novel framework that integrates reinforcement learning with self-supervised multimodal contrastive learning. SpaConTDS generates positive and negative samples through data augmentation and a pseudo-label tuple perturbation strategy, enabling the learning of fused representations that capture global semantics and cross-modal interactions. The model's hyper-parameters are dynamically optimized using reinforcement learning. Extensive experiments across various resolutions and platforms demonstrate that SpaConTDS achieves state-of-the-art accuracy in spatial domain identification and outperforms existing methods in downstream tasks such as denoising, trajectory inference, and UMAP visualization. Moreover, SpaConTDS effectively integrates multiple tissue sections and corrects batch effects without requiring prior alignment. Compared to existing approaches, SpaConTDS offers more robust fused representations of multimodal data, providing researchers with a flexible and powerful tool for a wide range of spatial transcriptomics analyses.
Ruiwen Xu, Xiaoqing Cheng, Wai-Ki Ching, Siyao Wu, Yuanben Zhang
PLoS Comput. Biol.3
2026 On the Number of Control Nodes in Boolean Networks With Degree Constraints
abstract
In this study, we analyze the minimum control node set problem for Boolean networks (BNs) with degree constraints. Our major contribution is the derivation of nontrivial lower and upper bounds on the size of the minimum control node set through combinatorial analysis of four types of BNs (i.e., $k$ - $k$ -XOR-BNs, simple $k$ - $k$ -AND-BNs, $k$ - $k$ -AND-BNs with negation, and $k$ - $k$ -NC-BNs, where the indegree and outdegree of each node are both $k$ , and the $k$ - $k$ -AND-BN with negation is an extension of the simple $k$ - $k$ -AND-BN that considers the occurrence of negation and NC means nested canalyzing). More specifically, four bounds for the size of the minimum control node set: general lower bound, best case upper bound, worst-case lower bound, and general upper bound are analyzed. By dividing nodes into three disjoint sets, extending the time to reach the target state, and utilizing necessary conditions for controllability, these bounds are obtained. Further, meaningful results and phenomena are discovered. Notably, all of the above results involving the AND function also apply to the OR function.
Liangjie Sun, Wai-Ki Ching, Tatsuya Akutsu
IEEE Trans. Cybern.2
2025 scRECL: representative ensembles with contrastive learning for scRNA-seq data clustering analysis
abstract
Single-cell transcriptomics characterizes gene expression profiles at the single-cell level, offering an unprecedented opportunity to understand cellular systems. As a fundamental task in single-cell data analysis, cell clustering significantly contributes to identifying cellular heterogeneity, thereby affecting downstream analyses. A number of deep learning methods have been proposed for clustering single-cell RNA sequencing (scRNA-seq) data. However, the large parameter space makes these methods sensitive to parameter settings. To leverage the strong capabilities of deep learning in capturing complex structures in single-cell data while ensuring algorithmic robustness, we propose a contrastive ensemble learning method named scRECL for scRNA-seq data clustering. In our approach, Siamese neural networks are trained under various $k$-nearest neighbors partitions to obtain low-dimensional embeddings of the scRNA-seq data. Multiplex graphs in representative element selection help filter out noisy and redundant cells. Consequently, contrastive ensemble learning is performed for efficient and effective latent embedding, as well as robust analysis of cellular heterogeneity in scRNA-seq data.
Hao Jiang 0009, Wai-Ki Ching, Dong Shen 0002
Briefings Bioinform.3
2025 Adjustable cash inflows based online investment decision making
Benmeng Lyu, Sini Guo, Jia-Wen Gu, Wai-Ki Ching
Expert Syst. Appl.4
2025 A discrete perspective towards the construction of sparse probabilistic Boolean networks
Christopher H. Fok, Chi-Wing Wong, Wai-Ki Ching
Inf. Sci.3
2025 On the Compressive Power of Autoencoders With Linear and ReLU Activation Functions
abstract
In this article, we mainly study the depth and width of autoencoders consisting of rectified linear unit (ReLU) activation functions. An autoencoder is a layered neural network consisting of an encoder, which compresses an input vector to a lower-dimensional vector, and a decoder, which transforms the low-dimensional vector back to the original input vector exactly (or approximately). In a previous study, Melkman et al. (2023) studied the depth and width of autoencoders using linear threshold activation functions with binary input and output vectors. We show that similar theoretical results hold if autoencoders using ReLU activation functions with real input and output vectors are used. Furthermore, we show that it is possible to compress input vectors to one-dimensional vectors using ReLU activation functions, although the size of compressed vectors is trivially Ω(log n) for autoencoders with linear threshold activation functions, where n is the number of input vectors. We also study the cases of linear activation functions. The results suggest that the compressive power of autoencoders using linear activation functions is considerably limited compared with those using ReLU activation functions.
Liangjie Sun, Chenyao Wu, Wai-Ki Ching, Tatsuya Akutsu
Neural Comput.3
2025 APDCA: An accurate and effective method for predicting associations between RBPs and AS-events during epithelial-mesenchymal transition
abstract
MOTIVATION: Epithelial-mesenchymal transition (EMT) plays a key role in cancer metastasis by promoting changes in adhesion and motility. RNA-binding proteins (RBPs) regulate alternative splicing (AS) during EMT, enabling a single gene to produce multiple protein isoforms that affect tumor progression. Disruption of RBP-AS interactions may disrupt the progress of diseases like cancer. Despite the importance of RBP-AS relationships in EMT, few computational methods predict these associations. Existing models struggle in sparse settings with limited known associations. To improve performance, we incorporate both sparsity constraints and heterogeneous biological data to infer RBP-AS associations. RESULT: We propose a new method based on Accelerated Proximal DC Algorithm (APDCA) for predicting RBP-AS associations. In particular, APDCA combines sparse low-rank matrix factorization with a Difference-of-Convex (DC) optimization framework and uses extrapolation to improve convergence. A key feature of APDCA is the use of a sparsity constraint, which filters out noise and highlights key associations. In addition, integrating multiple related data sources with direct or indirect relationships can help in reaching a more comprehensive view of RBPs and AS events and to reduce the impact of false positives associated with individual data sources. we prove that our proposed algorithm is convergent under some conditions and the experimental results have illustrated that APDCA outperforms six baseline methods in both AUC and AUPR. A case study on the RBP QKI shows that the top predictions are verified by the OncoSplicing database. Thus, APDCA provides a fast, interpretable, and scalable tool for discovering post-transcriptional regulatory interactions.
Yangsong He, Zheng-Jian Bai, Wai-Ki Ching, Quan Zou 0001, Yushan Qiu
PLoS Comput. Biol.3
2025 PLNMFG: Pseudo-label guided non-negative matrix factorization model with graph constraint for single-cell multi-omics data clustering
abstract
The development of single-cell multi-omics sequencing technologies has enabled the simultaneous analysis of multi-omics data within the same cell. Accurate clustering of these cells is crucial for downstream analyses of complex biological functions. Despite significant advances in multi-omics integration approaches, current methodologies exhibit two major limitations. First, they inadequately incorporate prior biological knowledge from various omic layers. Second, these methods often conduct independent dimensionality reduction on individual omic datasets, thereby failing to capture the intrinsic complementary information and potentially overlooking crucial cross-platform interactions. Motivated by these, this study investigates a non-negative matrix factorization model called PLNMFG, which integrates the unified latent representation learning that retains the features between and within omics and the cluster structure learning that retains the intrinsic structure of the data into one joint framework. Specially, PLNMFG performs adaptive imputation to handle dropout events and uses prior pseudo-labels as constraints during the process of collective non-negative matrix factorization, as a result, a more robust latent representation that preserves the double similarity information is obtained. Graph Laplacian constraint is applied during clustering which further preserves structure characteristic of multi-omics data. In addition, the weight of each omic is adaptively learned based on the omic contribution. A series of experiments on 8 benchmark datasets show that our model performs well in terms of clustering accuracy and computational efficiency.
Mingzhu Liu, Yushan Qiu, Wai-Ki Ching, Quan Zou 0001
PLoS Comput. Biol.4
2025 NetWalkRank: Cancer Driver Gene Prioritization in Multiplex Gene Regulatory Networks by a Random Walk Approach
abstract
Finding and prioritizing cancer driver genes (CDGs) that disrupt normal cell functionality and contribute to cancer occurrence and development is a significant challenge in oncology. Integrating multiple information pertaining to the characteristics of each gene at different stages of the disease and incorporating multiple steps as individual layers in the model provides a more comprehensive understanding of each node or gene. Thus, it is reasonable to organize them into multiplex gene regulatory networks (GRNs). In this work, we present a network-based framework called NetWalkRank, for prioritizing CDGs in the multiplex GRNs with gene expression profiling data. The framework applies the concept of network propagation to calculate the relative impact of each gene in spreading abnormality throughout the multiplex GRNs. It was employed to give priority to the driver genes of hepatocellular carcinoma (HCC) in humans. The performance of NetWalkRank was demonstrated through the ranks and classifications assigned to the known CDGs, which validated its effectiveness. To showcase the predictive capabilities of our proposed framework, we trained a random forest model that utilizes the obtained scores to accurately predict CDGs. We compared the advantage and efficiency of our method with other well-known driver gene ranking methods through numerical experiments. The findings show that the usage of GRNs across various steps of multiplex networks in prioritizing and predicting CDGs is significant, as demonstrated by the efficiency and effectiveness of NetWalkRank.
Fateme Keikha, Wai-Ki Ching, Zhi-Ping Liu
IEEE Trans. Comput. Biol. Bioinform.3
2025 Finite-Time Stabilizers for Large-Scale Stochastic Boolean Networks
abstract
This article presents a distributed pinning control strategy aimed at achieving global stabilization of Markovian jump Boolean control networks. The strategy relies on network matrix information to choose controlled nodes and adopts the algebraic state space representation approach for designing pinning controllers. Initially, a sufficient criterion is established to verify the global stability of a given Markovian jump Boolean network (MJBN) with probability one at a specific state within finite time. To stabilize an unstable MJBN at a predetermined state, the selection of pinned nodes involves removing the minimal number of entries, ensuring that the network matrix transforms into a strictly lower (or upper) triangular form. For each pinned node, two types of state feedback controllers are developed: 1) mode-dependent and 2) mode-independent, with a focus on designing a minimally updating controller. The choice of controller type is determined by the feasibility condition of the mode-dependent pinning controller, which is articulated through the solvability of matrix equations. Finally, the theoretical results are illustrated by studying the T cell large granular lymphocyte survival signaling network consisting of 54 genes and 6 stimuli.
Lin Lin 0012, James Lam, Wai-Ki Ching, Liangjie Sun, Bo Min
IEEE Trans. Cybern.3
2024 scEWE: high-order element-wise weighted ensemble clustering for heterogeneity analysis of single-cell RNA-sequencing data
abstract
With the emergence of large amount of single-cell RNA sequencing (scRNA-seq) data, the exploration of computational methods has become critical in revealing biological mechanisms. Clustering is a representative for deciphering cellular heterogeneity embedded in scRNA-seq data. However, due to the diversity of datasets, none of the existing single-cell clustering methods shows overwhelming performance on all datasets. Weighted ensemble methods are proposed to integrate multiple results to improve heterogeneity analysis performance. These methods are usually weighted by considering the reliability of the base clustering results, ignoring the performance difference of the same base clustering on different cells. In this paper, we propose a high-order element-wise weighting strategy based self-representative ensemble learning framework: scEWE. By assigning different base clustering weights to individual cells, we construct and optimize the consensus matrix in a careful and exquisite way. In addition, we extracted the high-order information between cells, which enhanced the ability to represent the similarity relationship between cells. scEWE is experimentally shown to significantly outperform the state-of-the-art methods, which strongly demonstrates the effectiveness of the method and supports the potential applications in complex single-cell data analytical problems.
Hao Jiang 0009, Wai-Ki Ching
Briefings Bioinform.3
2024 BANMF-S: a blockwise accelerated non-negative matrix factorization framework with structural network constraints for single cell imputation
abstract
MOTIVATION: Single cell RNA sequencing (scRNA-seq) technique enables the transcriptome profiling of hundreds to ten thousands of cells at the unprecedented individual level and provides new insights to study cell heterogeneity. However, its advantages are hampered by dropout events. To address this problem, we propose a Blockwise Accelerated Non-negative Matrix Factorization framework with Structural network constraints (BANMF-S) to impute those technical zeros. RESULTS: BANMF-S constructs a gene-gene similarity network to integrate prior information from the external PPI network by the Triadic Closure Principle and a cell-cell similarity network to capture the neighborhood structure and temporal information through a Minimum-Spanning Tree. By collaboratively employing these two networks as regularizations, BANMF-S encourages the coherence of similar gene and cell pairs in the latent space, enhancing the potential to recover the underlying features. Besides, BANMF-S adopts a blocklization strategy to solve the traditional NMF problem through distributed Stochastic Gradient Descent method in a parallel way to accelerate the optimization. Numerical experiments on simulations and real datasets verify that BANMF-S can improve the accuracy of downstream clustering and pseudo-trajectory inference, and its performance is superior to seven state-of-the-art algorithms. AVAILABILITY: All data used in this work are downloaded from publicly available data sources, and their corresponding accession numbers or source URLs are provided in Supplementary File Section 5.1 Dataset Information. The source codes are publicly available in Github repository https://github.com/jiayingzhao/BANMF-S.
Jiaying Zhao, Wai-Ki Ching, Chi-Wing Wong, Xiaoqing Cheng
Briefings Bioinform.2
2024 AGML: Adaptive Graph-Based Multi-Label Learning for Prediction of RBP and as Event Associations During EMT
abstract
Increasing evidence has indicated that RNA-binding proteins (RBPs) play an essential role in mediating alternative splicing (AS) events during epithelial-mesenchymal transition (EMT). However, due to the substantial cost and complexity of biological experiments, how AS events are regulated and influenced remains largely unknown. Thus, it is important to construct effective models for inferring hidden RBP-AS event associations during EMT process. In this paper, a novel and efficient model was developed to identify AS event-related candidate RBPs based on Adaptive Graph-based Multi-Label learning (AGML). In particular, we propose to adaptively learn a new affinity graph to capture the intrinsic structure of data for both RBPs and AS events. Multi-view similarity matrices are employed for maintaining the intrinsic structure and guiding the adaptive graph learning. We then simultaneously update the RBP and AS event associations that are predicted from both spaces by applying multi-label learning. The experimental results have shown that our AGML achieved AUC values of 0.9521 and 0.9873 by 5-fold and leave-one-out cross-validations, respectively, indicating the superiority and effectiveness of our proposed model. Furthermore, AGML can serve as an efficient and reliable tool for uncovering novel AS events-associated RBPs and is applicable for predicting the associations between other biological entities.
Yushan Qiu, Wai-Ki Ching, Hongmin Cai, Hao Jiang 0009, Quan Zou 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 The Construction of Sparse Probabilistic Boolean Networks: A Discrete Perspective
abstract
Boolean Network (BN) and its extension Probabilistic Boolean Network (PBN) are popular mathematical models for studying genetic regulatory networks. Apart from applications in genetic networks, BNs and PBNs also find many other applications in modeling financial risk, manufacturing systems and healthcare service systems. In this paper, we propose a novel Greedy Entry Removal (GER) algorithm for constructing sparse PBNs from rational transition probability matrices. We present theoretical upper bounds for both existing algorithms and the GER algorithm. Furthermore, we are the first to study and provide the lower bound of the captured problem under some simple condition. Our numerical experiments based on both synthetic and practical data demonstrate that GER gives the best performance among state-of-the-art sparse PBN construction algorithms.
Christopher H. Fok, Wai-Ki Ching, Chi-Wing Wong
BIBM2
2023 NG-SEM: an effective non-Gaussian structural equation modeling framework for gene regulatory network inference from single-cell RNA-seq data
abstract
Inference of gene regulatory network (GRN) from gene expression profiles has been a central problem in systems biology and bioinformatics in the past decades. The tremendous emergency of single-cell RNA sequencing (scRNA-seq) data brings new opportunities and challenges for GRN inference: the extensive dropouts and complicated noise structure may also degrade the performance of contemporary gene regulatory models. Thus, there is an urgent need to develop more accurate methods for gene regulatory network inference in single-cell data while considering the noise structure at the same time. In this paper, we extend the traditional structural equation modeling (SEM) framework by considering a flexible noise modeling strategy, namely we use the Gaussian mixtures to approximate the complex stochastic nature of a biological system, since the Gaussian mixture framework can be arguably served as a universal approximation for any continuous distributions. The proposed non-Gaussian SEM framework is called NG-SEM, which can be optimized by iteratively performing Expectation-Maximization algorithm and weighted least-squares method. Moreover, the Akaike Information Criteria is adopted to select the number of components of the Gaussian mixture. To probe the accuracy and stability of our proposed method, we design a comprehensive variate of control experiments to systematically investigate the performance of NG-SEM under various conditions, including simulations and real biological data sets. Results on synthetic data demonstrate that this strategy can improve the performance of traditional Gaussian SEM model and results on real biological data sets verify that NG-SEM outperforms other five state-of-the-art methods.
Jiaying Zhao, Chi-Wing Wong, Wai-Ki Ching, Xiaoqing Cheng
Briefings Bioinform.3
2023 Robust joint clustering of multi-omics single-cell data via multi-modal high-order neighborhood Laplacian matrix optimization
abstract
MOTIVATION: Simultaneous profiling of multi-omics single-cell data represents exciting technological advancements for understanding cellular states and heterogeneity. Cellular indexing of transcriptomes and epitopes by sequencing allowed for parallel quantification of cell-surface protein expression and transcriptome profiling in the same cells; methylome and transcriptome sequencing from single cells allows for analysis of transcriptomic and epigenomic profiling in the same individual cells. However, effective integration method for mining the heterogeneity of cells over the noisy, sparse, and complex multi-modal data is in growing need. RESULTS: In this article, we propose a multi-modal high-order neighborhood Laplacian matrix optimization framework for integrating the multi-omics single-cell data: scHoML. Hierarchical clustering method was presented for analyzing the optimal embedding representation and identifying cell clusters in a robust manner. This novel method by integrating high-order and multi-modal Laplacian matrices would robustly represent the complex data structures and allow for systematic analysis at the multi-omics single-cell level, thus promoting further biological discoveries. AVAILABILITY AND IMPLEMENTATION: Matlab code is available at https://github.com/jianghruc/scHoML.
Hao Jiang 0009, Senwen Zhan, Wai-Ki Ching, Luonan Chen
Bioinform.3
2023 LogBTF: gene regulatory network inference using Boolean threshold network model from single-cell gene expression data
abstract
MOTIVATION: From a systematic perspective, it is crucial to infer and analyze gene regulatory network (GRN) from high-throughput single-cell RNA sequencing data. However, most existing GRN inference methods mainly focus on the network topology, only few of them consider how to explicitly describe the updated logic rules of regulation in GRNs to obtain their dynamics. Moreover, some inference methods also fail to deal with the over-fitting problem caused by the noise in time series data. RESULTS: In this article, we propose a novel embedded Boolean threshold network method called LogBTF, which effectively infers GRN by integrating regularized logistic regression and Boolean threshold function. First, the continuous gene expression values are converted into Boolean values and the elastic net regression model is adopted to fit the binarized time series data. Then, the estimated regression coefficients are applied to represent the unknown Boolean threshold function of the candidate Boolean threshold network as the dynamical equations. To overcome the multi-collinearity and over-fitting problems, a new and effective approach is designed to optimize the network topology by adding a perturbation design matrix to the input data and thereafter setting sufficiently small elements of the output coefficient vector to zeros. In addition, the cross-validation procedure is implemented into the Boolean threshold network model framework to strengthen the inference capability. Finally, extensive experiments on one simulated Boolean value dataset, dozens of simulation datasets, and three real single-cell RNA sequencing datasets demonstrate that the LogBTF method can infer GRNs from time series data more accurately than some other alternative methods for GRN inference. AVAILABILITY AND IMPLEMENTATION: The source data and code are available at https://github.com/zpliulab/LogBTF.
Liangjie Sun, Chi-Wing Wong, Wai-Ki Ching, Zhi-Ping Liu
Bioinform.5
2023 PCB: A pseudotemporal causality-based Bayesian approach to identify EMT-associated regulatory relationships of AS events and RBPs during breast cancer progression
abstract
During breast cancer metastasis, the developmental process epithelial-mesenchymal (EM) transition is abnormally activated. Transcriptional regulatory networks controlling EM transition are well-studied; however, alternative RNA splicing also plays a critical regulatory role during this process. Alternative splicing was proved to control the EM transition process, and RNA-binding proteins were determined to regulate alternative splicing. A comprehensive understanding of alternative splicing and the RNA-binding proteins that regulate it during EM transition and their dynamic impact on breast cancer remains largely unknown. To accurately study the dynamic regulatory relationships, time-series data of the EM transition process are essential. However, only cross-sectional data of epithelial and mesenchymal specimens are available. Therefore, we developed a pseudotemporal causality-based Bayesian (PCB) approach to infer the dynamic regulatory relationships between alternative splicing events and RNA-binding proteins. Our study sheds light on facilitating the regulatory network-based approach to identify key RNA-binding proteins or target alternative splicing events for the diagnosis or treatment of cancers. The data and code for PCB are available at: http://hkumath.hku.hk/~wkc/PCB(data+code).zip.
Liangjie Sun, Yushan Qiu, Wai-Ki Ching, Quan Zou 0001
PLoS Comput. Biol.3
2023 On Synchronization Design and State Observer Design of (Singular) Boolean Networks
abstract
This paper first investigates the synchronization design problem for both Boolean networks (BNs) and singular BNs. According to the complete family of reachable sets, all possible synchronized response BNs and all possible weakly synchronized response singular BNs can be constructed. Subsequently, it is noted that there is similarity between synchronized response (singular) BNs and state observers. The synchronized response (singular) BN is constructed based on the state of a given drive (singular) BN, making its state gradually tend to be consistent with the state of the given drive (singular) BN. While the state observer is designed based on the output of a given real (singular) BN, making its state (state estimation) gradually tend to be consistent with the state of the given real (singular) BN. Therefore, consider applying the results for constructing synchronized response (singular) BNs to designing state observers for (singular) BNs. Specifically, the real (singular) BN to be observed can be regarded as a drive (singular) BN. Consequently, the problem for designing state observers can be transformed into the problem for constructing response (singular) BNs that achieve synchronization with a given real (singular) BN, and then the gain matrix of the designed state observer can be obtained by solving a matrix equation. The method for designing state observers is superior to other methods in that the accurate estimated time for designed state observers can be determined and adjusted to some certain extent, and all possible state observers in the considered form can be designed.
Liangjie Sun, Wai-Ki Ching, Shiyong Zhu, Jianquan Lu
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 On the Compressive Power of Boolean Threshold Autoencoders
abstract
An autoencoder is a layered neural network whose structure can be viewed as consisting of an encoder, which compresses an input vector to a lower dimensional vector, and a decoder, which transforms the low-dimensional vector back to the original input vector (or one that is very similar). In this article, we explore the compressive power of autoencoders that are Boolean threshold networks by studying the numbers of nodes and layers that are required to ensure that each vector in a given set of distinct input binary vectors is transformed back to its original. We show that for any set of n distinct vectors there exists a seven-layer autoencoder with the optimal compression ratio, (i.e., the size of the middle layer is logarithmic in n ), but that there is a set of n vectors for which there is no three-layer autoencoder with a middle layer of logarithmic size. In addition, we present a kind of tradeoff: if the compression ratio is allowed to be considerably larger than the optimal, then there is a five-layer autoencoder. We also study the numbers of nodes and layers required only for encoding, and the results suggest that the decoding part is the bottleneck of autoencoding. For example, there always is a three-layer Boolean threshold encoder that compresses n vectors into a dimension that is twice the logarithm of n .
Avraham A. Melkman, Sini Guo, Wai-Ki Ching, Pengyu Liu 0002, Tatsuya Akutsu
IEEE Trans. Neural Networks Learn. Syst.3
2022 Modeling Long-Range Travelling Times with Big Railway Data
Wenya Sun, Tobias Grubenmann, Reynold Cheng, Ben Kao, Wai-Ki Ching
DASFAA (3)5
2022 A high-order norm-product regularized multiple kernel learning framework for kernel optimization
Hao Jiang 0009, Dong Shen 0002, Wai-Ki Ching, Yushan Qiu
Inf. Sci.3
2022 On the Distribution of Successor States in Boolean Threshold Networks
abstract
We study the distribution of successor states in Boolean networks (BNs). The state vector${\mathbf{y}}$is called a successor of${\mathbf{x}}$if${\mathbf{y}}= \textbf {F}({\mathbf{x}})$holds, where${\mathbf{x}}, {\mathbf{y}}\in \{0,1\}^{n}$are state vectors and$\textbf {F}$is an ordered set of Boolean functions describing the state transitions. This problem is motivated by analyzing how information propagates via hidden layers in Boolean threshold networks (discrete model of neural networks) and is kept or lost during time evolution in BNs. In this article, we measure the distribution via entropy and study how entropy changes via the transition from${\mathbf{x}}$to${\mathbf{y}}$, assuming that${\mathbf{x}}$is given uniformly at random. We focus on BNs consisting of exclusive OR (XOR) functions, canalyzing functions, and threshold functions. As a main result, we show that there exists a BN consisting of$d$-ary XOR functions, which preserves the entropy if$d$is odd and$n > d$, whereas there does not exist such a BN if$d$is even. We also show that there exists a specific BN consisting of$d$-ary threshold functions, which preserves the entropy if$n \mod d = 0$. Furthermore, we theoretically analyze the upper and lower bounds of the entropy for BNs consisting of canalyzing functions and perform computational experiments using BN models of real biological networks.
Sini Guo, Pengyu Liu 0002, Wai-Ki Ching, Tatsuya Akutsu
IEEE Trans. Neural Networks Learn. Syst.3
2021 Matrix factorization-based data fusion for the prediction of RNA-binding proteins and alternative splicing event associations during epithelial-mesenchymal transition
abstract
MOTIVATION: The epithelial-mesenchymal transition (EMT) is a cellular-developmental process activated during tumor metastasis. Transcriptional regulatory networks controlling EMT are well studied; however, alternative RNA splicing also plays a critical regulatory role during this process. Unfortunately, a comprehensive understanding of alternative splicing (AS) and the RNA-binding proteins (RBPs) that regulate it during EMT remains largely unknown. Therefore, a great need exists to develop effective computational methods for predicting associations of RBPs and AS events. Dramatically increasing data sources that have direct and indirect information associated with RBPs and AS events have provided an ideal platform for inferring these associations. RESULTS: In this study, we propose a novel method for RBP-AS target prediction based on weighted data fusion with sparse matrix tri-factorization (WDFSMF in short) that simultaneously decomposes heterogeneous data source matrices into low-rank matrices to reveal hidden associations. WDFSMF can select and integrate data sources by assigning different weights to those sources, and these weights can be assigned automatically. In addition, WDFSMF can identify significant RBP complexes regulating AS events and eliminate noise and outliers from the data. Our proposed method achieves an area under the receiver operating characteristic curve (AUC) of $90.78\%$, which shows that WDFSMF can effectively predict RBP-AS event associations with higher accuracy compared with previous methods. Furthermore, this study identifies significant RBPs as complexes for AS events during EMT and provides solid ground for further investigation into RNA regulation during EMT and metastasis. WDFSMF is a general data fusion framework, and as such it can also be adapted to predict associations between other biological entities.
Yushan Qiu, Wai-Ki Ching, Quan Zou 0001
Briefings Bioinform.2
2021 Prediction of RNA-binding protein and alternative splicing event associations during epithelial-mesenchymal transition based on inductive matrix completion
abstract
MOTIVATION: The developmental process of epithelial-mesenchymal transition (EMT) is abnormally activated during breast cancer metastasis. Transcriptional regulatory networks that control EMT have been well studied; however, alternative RNA splicing plays a vital regulatory role during this process and the regulating mechanism needs further exploration. Because of the huge cost and complexity of biological experiments, the underlying mechanisms of alternative splicing (AS) and associated RNA-binding proteins (RBPs) that regulate the EMT process remain largely unknown. Thus, there is an urgent need to develop computational methods for predicting potential RBP-AS event associations during EMT. RESULTS: We developed a novel model for RBP-AS target prediction during EMT that is based on inductive matrix completion (RAIMC). Integrated RBP similarities were calculated based on RBP regulating similarity, and RBP Gaussian interaction profile (GIP) kernel similarity, while integrated AS event similarities were computed based on AS event module similarity and AS event GIP kernel similarity. Our primary objective was to complete missing or unknown RBP-AS event associations based on known associations and on integrated RBP and AS event similarities. In this paper, we identify significant RBPs for AS events during EMT and discuss potential regulating mechanisms. Our computational results confirm the effectiveness and superiority of our model over other state-of-the-art methods. Our RAIMC model achieved AUC values of 0.9587 and 0.9765 based on leave-one-out cross-validation (CV) and 5-fold CV, respectively, which are larger than the AUC values from the previous models. RAIMC is a general matrix completion framework that can be adopted to predict associations between other biological entities. We further validated the prediction performance of RAIMC on the genes CD44 and MAP3K7. RAIMC can identify the related regulating RBPs for isoforms of these two genes. AVAILABILITY AND IMPLEMENTATION: The source code for RAIMC is available at https://github.com/yushanqiu/RAIMC. CONTACT: [email protected] online.
Yushan Qiu, Wai-Ki Ching, Quan Zou 0001
Briefings Bioinform.2
2021 High-order Markov-switching portfolio selection with capital gain tax
Sini Guo, Wai-Ki Ching
Expert Syst. Appl.2
2021 Unsupervised Learning Framework With Multidimensional Scaling in Predicting Epithelial-Mesenchymal Transitions
abstract
Clustering tumor metastasis samples from gene expression data at the whole genome level remains an arduous challenge, in particular, when the number of experimental samples is small and the number of genes is huge. We focus on the prediction of the epithelial-mesenchymal transition (EMT), which is an underlying mechanism of tumor metastasis, here, rather than tumor metastasis itself, to avoid confounding effects of uncertainties derived from various factors. In this paper, we propose a novel model in predicting EMT based on multidimensional scaling (MDS) strategies and integrating entropy and random matrix detection strategies to determine the optimal reduced number of dimension in low dimensional space. We verified our proposed model with the gene expression data for EMT samples of breast cancer and the experimental results demonstrated the superiority over state-of-the-art clustering methods. Furthermore, we developed a novel feature extraction method for selecting the significant genes and predicting the tumor metastasis. The source code is available at "https://github.com/yushanqiu/yushan.qiu-szu.edu.cn".
Yushan Qiu, Hao Jiang 0009, Wai-Ki Ching
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 Fuzzy hidden Markov-switching portfolio selection with capital gain tax
Sini Guo, Wai-Ki Ching, Wai-Keung Li, Tak Kuen Siu, Zhiwen Zhang 0001
Expert Syst. Appl.2
2020 Switching-based stabilization of aperiodic sampled-data Boolean control networks with all subsystems unstable
abstract
We aim to further study the global stability of Boolean control networks (BCNs) under aperiodic sampled-data control (ASDC). According to our previous work, it is known that a BCN under ASDC can be transformed into a switched Boolean network (SBN), and further global stability of the BCN under ASDC can be obtained by studying the global stability of the transformed SBN. Unfortunately, since the major idea of our previous work is to use stable subsystems to offset the state divergence caused by unstable subsystems, the SBN considered has at least one stable subsystem. The central thought in this paper is that switching behavior also has good stabilization; i.e., the SBN can also be stable with appropriate switching laws designed, even if all subsystems are unstable. This is completely different from that in our previous work. Specifically, for this case, the dwell time (DT) should be limited within a pair of upper and lower bounds. By means of the discretized Lyapunov function and DT, a sufficient condition for global stability is obtained. Finally, the above results are demonstrated by a biological example.
Liangjie Sun, Jianquan Lu, Wai-Ki Ching
Frontiers Inf. Technol. Electron. Eng.3
2020 Drug Side-Effect Profiles Prediction: From Empirical to Structural Risk Minimization
abstract
The identification of drug side-effects is considered to be an important step in drug design, which could not only shorten the time but also reduce the cost of drug development. In this paper, we investigate the relationship between the potential side-effects of drug candidates and their chemical structures. The preliminary Regularized Regression (RR) model for drug side-effects prediction has promising features in the efficiency of model training and the existence of a closed form solution. It performs better than other state-of-the-art methods, in terms of minimum accuracy and average accuracy. In order to dig inside how drug structure will associate with side effect, we further propose weighted GTS (Generalized T-Student Kernel: WGTS) SVM model from a structural risk minimization perspective. The SVM model proposed in this paper provides a better understanding of drug side-effects in the process of drug development. The usefulness of the WGTS model lies in the superior performance in a cross validation setting on 888 approved drugs with 1385 side-effects profiling from SIDER database. This work is expected to shed light on intriguing studies that predict potential un-identifying side-effects and suggest how we can avoid drug side-effects by the removal of some distinguished chemical structures.
Hao Jiang 0009, Yushan Qiu, Wenpin Hou, Xiaoqing Cheng, Man Yi Yim, Wai-Ki Ching
IEEE ACM Trans. Comput. Biol. Bioinform.6
2019 On predicting epithelial mesenchymal transition by integrating RNA-binding proteins and correlation data via L1/2-regularization method
Yushan Qiu, Hao Jiang 0009, Wai-Ki Ching, Michael Kwok-Po Ng
Artif. Intell. Medicine3
2018 GPS trajectory data segmentation based on probabilistic logic
Sini Guo, Xiang Li 0006, Wai-Ki Ching, Dan A. Ralescu, Wai-Keung Li, Zhiwen Zhang 0001
Int. J. Approx. Reason.3
2018 Identifying a Probabilistic Boolean Threshold Network From Samples
abstract
This paper studies the problem of exactly identifying the structure of a probabilistic Boolean network (PBN) from a given set of samples, where PBNs are probabilistic extensions of Boolean networks. Cheng et al. studied the problem while focusing on PBNs consisting of pairs of AND/OR functions. This paper considers PBNs consisting of Boolean threshold functions while focusing on those threshold functions that have unit coefficients. The treatment of Boolean threshold functions, and triplets and -tuplets of such functions, necessitates a deepening of the theoretical analyses. It is shown that wide classes of PBNs with such threshold functions can be exactly identified from samples under reasonable constraints, which include: 1) PBNs in which any number of threshold functions can be assigned provided that all have the same number of input variables and 2) PBNs consisting of pairs of threshold functions with different numbers of input variables. It is also shown that the problem of deciding the equivalence of two Boolean threshold functions is solvable in pseudopolynomial time but remains co-NP complete.
Avraham A. Melkman, Xiaoqing Cheng, Wai-Ki Ching, Tatsuya Akutsu
IEEE Trans. Neural Networks Learn. Syst.3
2016 Unconstrained optimization in projection method for indefinite SVMs
abstract
Positive semi-definiteness is a critical property in Support Vector Machine (SVM) methods to ensure efficient solutions through convex quadratic programming. In this paper, we introduce a projection matrix on indefinite kernels to formulate a positive semi-definite one. The proposed model can be regarded as a generalized version of the spectrum method (denoising method and flipping method) by varying parameter λ. In particular, our suggested optimal λ under the Bregman matrix divergence theory can be obtained using unconstrained optimization. Experimental results on 4 real world data sets ranging from glycan classification to cancer prediction show that the proposed model can achieve better or competitive performance when compared to the related indefinite kernel methods. This may suggest a new way in motif extractions or cancer predictions.
Hao Jiang 0009, Wai-Ki Ching, Yushan Qiu, Xiaoqing Cheng
BIBM2
2016 A systematic framework to derive N-glycan biosynthesis process and the automated construction of glycosylation networks
abstract
BACKGROUND: Abnormalities in glycan biosynthesis have been conclusively related to various diseases, whereas the complexity of the glycosylation process has impeded the quantitative analysis of biochemical experimental data for the identification of glycoforms contributing to disease. To overcome this limitation, the automatic construction of glycosylation reaction networks in silico is a critical step. RESULTS: In this paper, a framework K2014 is developed to automatically construct N-glycosylation networks in MATLAB with the involvement of the 27 most-known enzyme reaction rules of 22 enzymes, as an extension of previous model KB2005. A toolbox named Glycosylation Network Analysis Toolbox (GNAT) is applied to define network properties systematically, including linkages, stereochemical specificity and reaction conditions of enzymes. Our network shows a strong ability to predict a wider range of glycans produced by the enzymes encountered in the Golgi Apparatus in human cell expression systems. CONCLUSIONS: Our results demonstrate a better understanding of the underlying glycosylation process and the potential of systems glycobiology tools for analyzing conventional biochemical or mass spectrometry-based experimental data quantitatively in a more realistic and practical way.
Wenpin Hou, Yushan Qiu, Nobuyuki Hashimoto, Wai-Ki Ching, Kiyoko F. Aoki-Kinoshita
BMC Bioinform.4
2016 Sufficient conditions for the ergodicity of fuzzy Markov chains
Dong-Mei Zhu, Wai-Ki Ching, Sy-Ming Guu
Fuzzy Sets Syst.2
2016 Exact Identification of the Structure of a Probabilistic Boolean Network from Samples
abstract
We study the number of samples required to uniquely determine the structure of a probabilistic Boolean network (PBN), where PBNs are probabilistic extensions of Boolean networks. We show via theoretical analysis and computational analysis that the structure of a PBN can be exactly identified with high probability from a relatively small number of samples for interesting classes of PBNs of bounded indegree. On the other hand, we also show that there exist classes of PBNs for which it is impossible to uniquely determine the structure of a PBN from samples.
Xiaoqing Cheng, Tomoya Mori, Yushan Qiu, Wai-Ki Ching, Tatsuya Akutsu
IEEE ACM Trans. Comput. Biol. Bioinform.4
2015 On observability of attractors in Boolean Networks
abstract
Boolean network (BN) is a popular mathematical model for revealing the behavior of a genetic regulatory network, and observability plays a vital role in understanding the underlying network feature. However, the observability of attractor cycles, which is an interesting and important problem, has not been addressed in the literature. In this paper, we first proposed a novel problem on attractor observability in BNs. Identification of the minimum set of consecutive nodes can be used to determine uniquely the attractor cycle from the others in the network. We then develop a linear-time algorithm to identify the desired set of nodes. The proposed approaches are demonstrated and verified by numerical examples. The computational results are given to illustrate both the efficiency and effectiveness of our proposed methods.
Yushan Qiu, Xiaoqing Cheng, Wai-Ki Ching, Hao Jiang 0009, Tatsuya Akutsu
BIBM3
2014 A hidden Markov reduced-form risk model
abstract
In this paper, we propose a reduced-form credit risk model with a hidden state process. The hidden state process is adopted to model the underlying economic environment with an observable state revealing the delayed and noisy information of the underlying economic state. Our model is a generalization of the work in Gu et al. [1]. Under this framework, we give a computational method to extract the underlying economic state and to find the distribution of multiple default times. Numerical experiment is conducted to illustrate the impact of change in observable state and the contagion effect of defaults.
Jia-Wen Gu, Wai-Ki Ching, Harry Zheng
CIFEr2
2014 On pricing and hedging basket credit derivatives with dependent structure
abstract
In this paper, we study the problem of hedging a basket credit derivatives, in particular, we are interested in basket default swaps. For the pricing of credit derivatives, we consider a factor Copula approach. Single-name credit default swaps will be chosen as the hedging instruments. The hedging mechanism is tested using simulated data with a given measure. Numerical results reveal the efficiency of our proposed hedging method.
Dong-Mei Zhu, Wai-Ki Ching, Harry Zheng
CIFEr3
2013 Credit portfolio management using two-level particle swarm optimization
Fuqiang Lu, Min Huang 0001, Wai-Ki Ching, Tak Kuen Siu
Inf. Sci.3
2012 The role of Eigen-matrix translation in classification of biological datasets
abstract
Driven by the challenge of integrating large amount of experimental data obtained from biological research, computational biology and bioinformatics are growing rapidly. Machine learning methods, especially kernel methods with Support Vector Machines (SVMs) are very popular tools. In the perspective of kernel matrix, a technique namely Eigen-matrix translation has been introduced for protein data classification. The Eigen-matrix translation strategy owns a lot of nice properties while the nature of which needs further exploration. We propose that its importance lies in the dimension reduction of predictor attributes within the data set. This can therefore serve as a novel perspective for future research in dimension reduction problems.
Hao Jiang 0009, Wai-Ki Ching
BIBM2
2012 Regularized orthogonal linear discriminant analysis
Wai-Ki Ching, Delin Chu, Li-Zhi Liao
Pattern Recognit.1
2011 A distributed decision making model for risk management of virtual enterprise
Min Huang 0001, Fuqiang Lu, Wai-Ki Ching, Tak Kuen Siu
Expert Syst. Appl.3
2010 Finding optimal control policy in Probabilistic Boolean Networks with hard constraints by using integer programming and dynamic programming
abstract
In this paper, we study control problems of Boolean Networks (BNs) and Probabilistic Boolean Networks (PBNs). For BN CONTROL, by applying external control, we propose to derive the network to the desired state within a few time steps. For PBN CONTROL, we propose to find a control sequence such that the network will terminate in the desired state with a maximum probability. Also, we propose to minimize the maximum cost of the terminal state to which the network will enter. Integer linear programming and dynamic programming in conjunction with hard constraints are then employed to solve the above problems. Numerical experiments are given to demonstrate the effectiveness of our algorithms. We also present a hardness result suggesting that PBN CONTROL is harder than BN CONTROL.
Tatsuya Akutsu, Takeyuki Tamura, Wai-Ki Ching
BIBM4
2010 A new multiple regression approach for the construction of genetic regulatory networks
Shuqin Zhang, Wai-Ki Ching, Nam-Kiu Tsing, Ho-Yin Leung, Dianjing Guo
Artif. Intell. Medicine2
2010 A weighted q-gram method for glycan structure classification
abstract
BACKGROUND: Glycobiology pertains to the study of carbohydrate sugar chains, or glycans, in a particular cell or organism. Many computational approaches have been proposed for analyzing these complex glycan structures, which are chains of monosaccharides. The monosaccharides are linked to one another by glycosidic bonds, which can take on a variety of comformations, thus forming branches and resulting in complex tree structures. The q-gram method is one of these recent methods used to understand glycan function based on the classification of their tree structures. This q-gram method assumes that for a certain q, different q-grams share no similarity among themselves. That is, that if two structures have completely different components, then they are completely different. However, from a biological standpoint, this is not the case. In this paper, we propose a weighted q-gram method to measure the similarity among glycans by incorporating the similarity of the geometric structures, monosaccharides and glycosidic bonds among q-grams. In contrast to the traditional q-gram method, our weighted q-gram method admits similarity among q-grams for a certain q. Thus our new kernels for glycan structure were developed and then applied in SVMs to classify glycans. RESULTS: Two glycan datasets were used to compare the weighted q-gram method and the original q-gram method. The results show that the incorporation of q-gram similarity improves the classification performance for all of the important glycan classes tested. CONCLUSION: The results in this paper indicate that similarity among q-grams obtained from geometric structure, monosaccharides and glycosidic linkage contributes to the glycan function classification. This is a big step towards the understanding of glycan function based on their complex structures.
Wai-Ki Ching, Takako Yamaguchi, Kiyoko F. Aoki-Kinoshita
BMC Bioinform.2
2010 Predicting enzyme targets for cancer drugs by profiling human Metabolic reactions in NCI-60 cell lines
abstract
BACKGROUND: Drugs can influence the whole metabolic system by targeting enzymes which catalyze metabolic reactions. The existence of interactions between drugs and metabolic reactions suggests a potential way to discover drug targets. RESULTS: In this paper, we present a computational method to predict new targets for approved anti-cancer drugs by exploring drug-reaction interactions. We construct a Drug-Reaction Network to provide a global view of drug-reaction interactions and drug-pathway interactions. The recent reconstruction of the human metabolic network and development of flux analysis approaches make it possible to predict each metabolic reaction's cell line-specific flux state based on the cell line-specific gene expressions. We first profile each reaction by its flux states in NCI-60 cancer cell lines, and then propose a kernel k-nearest neighbor model to predict related metabolic reactions and enzyme targets for approved cancer drugs. We also integrate the target structure data with reaction flux profiles to predict drug targets and the area under curves can reach 0.92. CONCLUSIONS: The cross validations using the methods with and without metabolic network indicate that the former method is significantly better than the latter. Further experiments show the synergism of reaction flux profiles and target structure for drug target prediction. It also implies the significant contribution of metabolic network to predict drug targets. Finally, we apply our method to predict new reactions and possible enzyme targets for cancer drugs.
Xiaobo Zhou 0001, Wai-Ki Ching
BMC Bioinform.3
2010 Modeling default risk via a hidden Markov model of multiple sequences
Wai-Ki Ching, Ho-Yin Leung, Hao Jiang 0009
Frontiers Comput. Sci. China1
2010 Generating probabilistic Boolean networks from a prescribed stationary distribution
Shuqin Zhang, Wai-Ki Ching, Nam-Kiu Tsing
Inf. Sci.2
2008 Efficient Reconstruction of Piecewise Constant Images Using Nonsmooth Nonconvex Minimization
abstract
We consider the restoration of piecewise constant images where the number of the regions and their values are not fixed in advance, with a good difference of piecewise constant values between neighboring regions, from noisy data obtained at the output of a linear operator (e.g., a blurring kernel or a Radon transform). Thus we also address the generic problem of unsupervised segmentation in the context of linear inverse problems. The segmentation and the restoration tasks are solved jointly by minimizing an objective function (an energy) composed of a quadratic data-fidelity term and a nonsmooth nonconvex regularization term. The pertinence of such an energy is ensured by the analytical properties of its minimizers. However, its practical interest used to be limited by the difficulty of the computational stage which requires a nonsmooth nonconvex minimization. Indeed, the existing methods are unsatisfactory since they (implicitly or explicitly) involve a smooth approximation of the regularization term and often get stuck in shallow local minima. The goal of this paper is to design a method that efficiently handles the nonsmooth nonconvex minimization. More precisely, we propose a continuation method where one tracks the minimizers along a sequence of approximate nonsmooth energies $\{J_\eps\}$, the first of which being strictly convex and the last one the original energy to minimize. Knowing the importance of the nonsmoothness of the regularization term for the segmentation task, each $J_\eps$ is nonsmooth and is expressed as the sum of an $\ell_1$ regularization term and a smooth nonconvex function. Furthermore, the local minimization of each $J_{\eps}$ is reformulated as the minimization of a smooth function subject to a set of linear constraints. The latter problem is solved by the modified primal-dual interior point method, which guarantees the descent direction at each step. Experimental results are presented and show the effectiveness and the efficiency of the proposed method. Comparison with simulated annealing methods further shows the advantage of our method.
Mila Nikolova, Michael Kwok-Po Ng, Shuqin Zhang, Wai-Ki Ching
SIAM J. Imaging Sci.4
2007 An approximation method for solving the steady-state probability distribution of probabilistic Boolean networks
abstract
MOTIVATION: Probabilistic Boolean networks (PBNs) have been proposed to model genetic regulatory interactions. The steady-state probability distribution of a PBN gives important information about the captured genetic network. The computation of the steady-state probability distribution usually includes construction of the transition probability matrix and computation of the steady-state probability distribution. The size of the transition probability matrix is 2(n)-by-2(n) where n is the number of genes in the genetic network. Therefore, the computational costs of these two steps are very expensive and it is essential to develop a fast approximation method. RESULTS: In this article, we propose an approximation method for computing the steady-state probability distribution of a PBN based on neglecting some Boolean networks (BNs) with very small probabilities during the construction of the transition probability matrix. An error analysis of this approximation method is given and theoretical result on the distribution of BNs in a PBN with at most two Boolean functions for one gene is also presented. These give a foundation and support for the approximation method. Numerical experiments based on a genetic network are given to demonstrate the efficiency of the proposed method.
Wai-Ki Ching, Shuqin Zhang, Michael Kwok-Po Ng, Tatsuya Akutsu
Bioinform.1
2007 A semi-supervised regression model for mixed numerical and categorical variables
Michael Kwok-Po Ng, Elaine Y. Chan, Mee Chi So, Wai-Ki Ching
Pattern Recognit.4
2006 On the Complexity of Finding Control Strategies for Boolean Networks
Tatsuya Akutsu, Morihiro Hayashida, Wai-Ki Ching, Michael Kwok-Po Ng
APBC3
2005 A Quantity-Time-Based Dispatching Policy for a VMI System
Wai-Ki Ching, Allen H. Tai
ICCSA (4)1
2005 Fast Algorithms for l1 Norm/Mixed l1 and l2 Norms for Image Restoration
Haoying Fu, Michael Kwok-Po Ng, Mila Nikolova, Jesse L. Barlow, Wai-Ki Ching
ICCSA (4)5
2005 On construction of stochastic genetic networks based on gene expression sequences
abstract
Reconstruction of genetic regulatory networks from time series data of gene expression patterns is an important research topic in bioinformatics. Probabilistic Boolean Networks (PBNs) have been proposed as an effective model for gene regulatory networks. PBNs are able to cope with uncertainty, corporate rule-based dependencies between genes and discover the sensitivity of genes in their interactions with other genes. However, PBNs are unlikely to use directly in practice because of huge amount of computational cost for obtaining predictors and their corresponding probabilities. In this paper, we propose a multivariate Markov model for approximating PBNs and describing the dynamics of a genetic network for gene expression sequences. The main contribution of the new model is to preserve the strength of PBNs and reduce the complexity of the networks. The number of parameters of our proposed model is O(n2) where n is the number of genes involved. We also develop efficient estimation methods for solving the model parameters. Numerical examples on synthetic data sets and practical yeast data sequences are given to demonstrate the effectiveness of the proposed model.
Wai-Ki Ching, Michael Kwok-Po Ng, Eric S. Fung, Tatsuya Akutsu
Int. J. Neural Syst.1
2005 A note on the paper: Optimizing web servers using page rank prefetching for clustered accesses
Wai-Ki Ching
Inf. Sci.1
2004 Building Genetic Networks for Gene Expression Patterns
Wai-Ki Ching, Eric S. Fung, Michael Kwok-Po Ng
IDEAL1
2004 An optimization algorithm for clustering using weighted dissimilarity measures
Elaine Y. Chan, Wai-Ki Ching, Michael Kwok-Po Ng, Joshua Zhexue Huang
Pattern Recognit.2
2004 Fast inversion of triangular Toeplitz matrices
Fu-Rong Lin, Wai-Ki Ching, Michael Kwok-Po Ng
Theor. Comput. Sci.2
2003 A Direct Method for Block-Toeplitz Systems with Applications to Re-manufacturing Systems
Wai-Ki Ching, Michael Kwok-Po Ng, Wai-On Yuen
ICCSA (1)1
2003 Higher-Order Hidden Markov Models with Applications to DNA Sequences
Wai-Ki Ching, Eric S. Fung, Michael Kwok-Po Ng
IDEAL1