Chengbin Peng 0001

dblp:11/9915 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0002-7445-2638ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 An Adaptive Neuro-fuzzy Framework for Stock Price Forecasting
Pengju Ren, Chengbin Peng 0001, Baisong Liu, Xiaoqin Fan
ICIC (7)2
2026 Defending against link prediction by residual path entropy maximization
Ru Yuan, Pietro Liò, Xu Shen 0002, Chengbin Peng 0001
Expert Syst. Appl.4
2026 Adaptive meta-path-based neural network architecture search for heterogeneous graphs
Xinsheng Li, Pietro Liò, Lintao Yang, Zhigang Ye, Chengbin Peng 0001
Inf. Sci.5
2026 AutoHGNN: Robust and efficient neural architecture search for hypergraph neural networks
Pietro Liò, Xinsheng Li, Baisong Liu, Chengbin Peng 0001
Knowl. Based Syst.5
2026 Exploiting Minority Pseudo-Labels for Semi-Supervised Fine-Grained Road Scene Understanding
Yuting Hong, Yongkang Wu, Hui Xiao 0005, Huazheng Hao, Xiaojie Qiu, Baochen Yao, Chengbin Peng 0001
IEEE Trans. Intell. Transp. Syst.7
2025 DDintensity: Addressing imbalanced drug-drug interaction risk levels using pre-trained deep learning model embeddings
Weidun Xie, Xingjian Chen, Zetian Zheng, Ruoxuan Zhang, Chengbin Peng 0001, Monika Gullerova, Ka-Chun Wong
Artif. Intell. Medicine9
2025 Active learning with joint probabilistic modeling for point cloud semantic segmentation
Baochen Yao, Dongjie Zhang 0001, Chengbin Peng 0001
Knowl. Based Syst.5
2025 An efficient and scalable semi-supervised framework for semantic segmentation
Huazheng Hao, Hui Xiao 0005, Li Dong 0006, Diqun Yan, Dongtai Liang, Jiayan Zhuang, Chengbin Peng 0001
Neural Comput. Appl.8
2024 A multi-view consistency framework with semi-supervised domain adaptation
Yuting Hong, Li Dong 0006, Xiaojie Qiu, Hui Xiao 0005, Baochen Yao, Siming Zheng, Chengbin Peng 0001
Eng. Appl. Artif. Intell.7
2024 Adaptive multi-scale Graph Neural Architecture Search framework
Lintao Yang, Pietro Liò, Xu Shen 0002, Chengbin Peng 0001
Neurocomputing5
2024 C²F²: Cross-Task Cross-Domain Feature Fusion for Semi-Supervised Change Detection
abstract
Semi-supervised learning for change detection (CD), which significantly reduces the labor costs associated with data annotation, has recently garnered substantial attention. In this study, we propose to enhance traditional semi-supervised learning frameworks by leveraging cross-task cross-domain (CTCD) models, which generate complementary features that differ from standard hidden features. The procedure is as follows. First, the standard features obtained from a traditional encoding–decoding structure are fused with attention-augmented complementary features. Second, a secondary decoder maps the fused heterogeneous features into the label space to obtain high-quality pseudo-labels, offering more precise guidance for semi-supervised learning on traditional structures. This approach improves pseudo-labels by leveraging the strength of CTCD models, including large pretrained models, to enhance the semi-supervised learning process of domain-specific and task-specific models. Experimental results on benchmark datasets demonstrate that our proposed approach surpasses state-of-the-art methods.
Dongjie Zhang 0001, Yuting Hong, Xiaojie Qiu, Li Dong 0006, Diqun Yan, Chengbin Peng 0001
IEEE Geosci. Remote. Sens. Lett.6
2024 Mixed-Bit Sampling Graphic: When Watermarking Meets Copy Detection Pattern
abstract
Copy Detection Pattern (CDP) is a high-density random noise-alike image that exhibits a different noise pattern after physical copying, and is thus treated as a promising anti-counterfeiting solution. However, CDP cannot convey any message, and it is often used in combination with additional carriers, such as QR codes. In this letter, we take the first step towards extending CDP with watermarking functionality. Specifically, we devise a scheme called Mixed-bit Sampling Graphic (MSG), which could realize invisible watermarking and anti-counterfeiting simultaneously. Compared with conventional CDP, the noise pattern generation of MSG is controlled by the portions of sampling over two bit templates. We formulate this mixed-bit sampling process as an optimization problem and solve it using a block coordinate descent sampling algorithm. Experimental results validate that the proposed MSG can effectively communicate watermark bits while retaining the anti-counterfeiting capability of CDP.
Li Dong 0006, Rangding Wang, Diqun Yan, Chengbin Peng 0001
IEEE Signal Process. Lett.5
2024 Point Cloud Semantic Segmentation by Adaptively Fusing Information With Varying Distances
abstract
Point clouds provide rich geometric representations, and point cloud semantic segmentation is essential in many applications. As the data scale of point clouds is usually quite large, some approaches propose constructing superpoint graphs from point clouds to reduce the time and space complexity during analysis. However, traditional superpoint-based graph neural network approaches for point cloud analysis typically aggregate features of adjacent superpoints and consider the most prominent feature only within each receptive field. In this work, we argue that adaptive varying-distance feature aggregation and discrimination can improve the effect of point cloud semantic segmentation. The proposed approach consists of three steps. First, we agglomerate points into superpoints and construct a superpoint graph as many traditional approaches. Second, we propose a novel varying-distance autoencoder to help each superpoint adaptively assimilate information from different distances. Third, we propose a discrimination loss to constrain the embedding space so that superpoints belonging to the same semantic class can get closer and vice versa. Regarding mIoU, our method outperforms the baseline by at least 8.1% for the S3DIS dataset and 3.1% in mIoU for the vKITTI dataset.
Zefeng Jiang, Baochen Yao, Kangkang Song, Xiaojie Qiu, Chengbin Peng 0001
IEEE Signal Process. Lett.5
2024 Self-Enhanced Feature Fusion for RGB-D Semantic Segmentation
abstract
Effectively fusing depth and RGB information to fully leverage their complementary strengths is essential for advancing RGB-D semantic segmentation. However, when fusing with RGB information, traditional methods often overlook noises in depth data, presuming that they are of high accuracy. To resolve this issue, we propose a self-enhanced feature fusion network (SEFnet) for RGB-D semantic segmentation in this work. It mainly comprises three steps. Firstly, RGB and depth embeddings from the initial layers of the network are fused together. Secondly, the fused features are enhanced by pure RGB embeddings and are progressively guided by semantic edge labels to suppress irrelevant features. Finally, the enhanced features are combined with high-level RGB features and are fed into a normalizing flow decoder to obtain segmentation results. Experimental results demonstrate that the proposed approach can provide accurate predictions, outperforming state-of-the-art methods on benchmark datasets.
Pengcheng Xiang, Baochen Yao, Zefeng Jiang, Chengbin Peng 0001
IEEE Signal Process. Lett.4
2024 Core-Periphery Detection Based on Masked Bayesian Nonnegative Matrix Factorization
abstract
Core–periphery structure is an essential mesoscale feature in complex networks. Previous researches mostly focus on discriminative approaches, while in this work we propose a generative model called masked Bayesian nonnegative matrix factorization. We build the model using two pair affiliation matrices to indicate core–periphery pair associations and using a mask matrix to highlight connections to core nodes. We propose an approach to infer the model parameters and prove the convergence of variables with our approach. Besides the abilities as traditional approaches, it is able to identify core scores with overlapping core–periphery pairs. We verify the effectiveness of our method using randomly generated networks and real-world networks. Experimental results demonstrate that the proposed method outperforms traditional approaches.
Zhonghao Wang 0002, Ru Yuan, Jiaye Fu, Ka-Chun Wong, Chengbin Peng 0001
IEEE Trans. Comput. Soc. Syst.5
2024 Uncertainty-Guided Contrastive Learning for Weakly Supervised Point Cloud Segmentation
abstract
Three-dimensional point cloud data are widely used in many fields, as they can be easily obtained and contain rich semantic information. Recently, weakly supervised segmentation has attracted lots of attention, because it only requires very few labels, thus reducing time-consuming and expensive data annotation efforts for huge amounts of point cloud data. The existing approaches typically adopt softmax scores from the last layer as the confidence for selecting high-confident point predictions. However, such approaches can ignore the potential value of a large number of low-confidence point predictions under traditional metrics. In this work, we propose an uncertainty-guided contrastive learning (UCL) framework for weakly supervised point cloud segmentation. A novel uncertainty metric based on prototype entropy (PE) is presented to estimate the reliability of model predictions. With this metric, we propose a negative contrastive learning module exploiting negative pseudo-labels of predictions with low reliability and an active contrastive learning module enhancing feature learning of segmentation models by predictions with high reliability. We also propose a generic multiscale feature perturbation method to expand a wider perturbation space. Extensive experimental results on indoor and outdoor point cloud datasets demonstrate that the proposed method achieves competitive performance.
Baochen Yao, Li Dong 0006, Xiaojie Qiu, Kangkang Song, Diqun Yan, Chengbin Peng 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Graph Rewiring and Preprocessing for Graph Neural Networks Based on Effective Resistance
abstract
Graph neural networks (GNNs) are powerful models for processing graph data and have demonstrated state-of-the-art performance on many downstream tasks. However, existing GNNs can generally suffer from two limitations: over-smoothing and over-squashing, which can significantly undermine their learning ability for large graphs. To overcome these issues simultaneously, by utilizing the concept of effective resistances, we focus on minimizing total constrained resistance while identifying problematic edges using topological redundancy and bottleneck sparsity coefficients. We introduce a novel graph rewiring and preprocessing method guided by effective resistance (GPER), capable of edge addition or removal. Theoretical analysis validates our method's efficacy in mitigating over-smoothing and over-squashing. In the experiments, we conduct node and graph classifications on the benchmark datasets and can achieve an average improvement of 7.8% and 2.0%, respectively. We also conduct scalability analysis on large graphs with GCN and demonstrate that the proposed preprocess approach can reduce graph size by over 50% while improve the performance.
Xu Shen 0002, Pietro Liò, Lintao Yang, Ru Yuan, Chengbin Peng 0001
IEEE Trans. Knowl. Data Eng.6
2024 Multi-Level Label Correction by Distilling Proximate Patterns for Semi-Supervised Semantic Segmentation
abstract
Semi-supervised semantic segmentation relieves the reliance on large-scale labeled data by leveraging unlabeled data. Recent semi-supervised semantic segmentation approaches mainly resort to pseudo-labeling methods to exploit unlabeled data. However, unreliable pseudo-labeling can undermine the semi-supervision processes. In this paper, we propose an algorithm called Multi-Level Label Correction (MLLC), which aims to use graph neural networks to capture structural relationships in Semantic-Level Graphs (SLGs) and Class-Level Graphs (CLGs) to rectify erroneous pseudo-labels. Specifically, SLGs represent semantic affinities between pairs of pixel features, and CLGs describe classification consistencies between pairs of pixel labels. With the support of proximate pattern information from graphs, MLLC can rectify incorrectly predicted pseudo-labels and can facilitate discriminative feature representations. We design an end-to-end network to train and perform this effective label corrections mechanism. Experiments demonstrate that MLLC can significantly improve supervised baselines and outperforms state-of-the-art approaches in different scenarios on Cityscapes and PASCAL VOC 2012 datasets. Specifically, MLLC improves the supervised baseline by at least 5% and 2% with DeepLabV2 and DeepLabV3+ respectively under different partition protocols.
Hui Xiao 0005, Yuting Hong, Li Dong 0006, Diqun Yan, Jiayan Zhuang, Dongtai Liang, Chengbin Peng 0001
IEEE Trans. Multim.8
2023 A Pseudo-Dual Self-Rectification Framework for Semantic Segmentation
abstract
Semantic segmentation has achieved remarkable success in various applications. However, the training process for such techniques necessitates a significant amount of labeled data. Although semi-supervised frameworks can alleviate this issue, traditional approaches typically require multiple baseline models to form a dual model. To allow a semi-supervised semantic segmentation framework to be used in robotic systems with precious computation and memory resources, we propose a framework utilizing a single baseline model only. The overall framework is composed of three parts: an encoder, a shallow decoder, and a deep decoder. It distills knowledge from the ensemble of two decoders to improve the encoder, which can implicitly form a pseudo-dual model. It also calculates class-wise likelihoods according to the similarity between features and class prototypes learned from different decoders and rectifies low-confidence pseudo-labels. Our framework outperforms state-of-the-art frameworks on benchmark datasets with a significant amount of decrease in using computing resources.
Huazheng Hao, Hui Xiao 0005, Li Dong 0006, Diqun Yan, Dongtai Liang, Jiayan Zhuang, Chengbin Peng 0001
ICME7
2023 Corrigendum to "Semi-supervised semantic segmentation with cross teacher training" [Neurocomputing 508 (2022) 36-46]
Hui Xiao 0005, Li Dong 0006, Shuibo Fu, Diqun Yan, Kangkang Song, Chengbin Peng 0001
Neurocomputing7
2023 Semi-supervised learning with pseudo-negative labels for image classification
Hui Xiao 0005, Huazheng Hao, Li Dong 0006, Xiaojie Qiu, Chengbin Peng 0001
Knowl. Based Syst.6
2022 Watermark-Preserving Keypoint Enhancement for Screen-Shooting Resilient Watermarking
abstract
Screen-shooting resilient (SSR) watermark is a special kind of robust watermarking. One can extract the watermark message even the embedded image communicates via a physical screen to the camera channel. The keypoint-based SSR watermarking is one promising solution to realize such screen-to-camera communication. The enhanced keypoints were used to locate the embedding region and then perform watermark embedding. However, the keypoint-based SSR watermarking treats the critical two steps, keypoint enhancement and watermark embedding, independently, neglecting their inter-play. This work proposes a watermark-preserving keypoint enhancement algorithm for SSR watermarking. Specifically, we resort to a convex constrained optimization framework to unify keypoint enhancement and watermark embedding. Multiple constraints are imposed to simultaneously ensure the watermark validity and blind synchronization of embedding regions. Our method enables jointly optimizing the watermarking distortion and keypoint enhancement. The proposed method achieves superior watermark extraction accuracy while retaining better watermarked image quality when compared with previous works.
Li Dong 0006, Chengbin Peng 0001, Yuanman Li, Weiwei Sun 0009
ICME3
2022 Semi-supervised semantic segmentation with cross teacher training
Hui Xiao 0005, Li Dong 0006, Shuibo Fu, Diqun Yan, Kangkang Song, Chengbin Peng 0001
Neurocomputing7
2020 Deleterious Non-Synonymous Single Nucleotide Polymorphism Predictions on Human Transcription Factors
abstract
Transcription factors (TFs) are the major components of human gene regulation. In particular, they bind onto specific DNA sequences and regulate neighborhood genes in different tissues at different developmental stages. Non-synonymous single nucleotide polymorphisms on its protein-coding sequences could result in undesired consequences in human. Therefore, it is necessary to develop methods for predicting any abnormality among those non-synonymous single nucleotide polymorphisms. To address it, we have developed and compared different strategies to predict deleterious non-synonymous single nucleotide polymorphisms (also known as missense mutations) on the protein-coding sequences of human TFs. Taking advantage of evolutionary conservation signals, we have developed and compared different classifiers with different feature sets as computed from different evolutionarily related sequence collections. The results indicate that the classic ensemble algorithm, Adaboost with decision stumps, with orthologous sequence collection, has performed the best (namely, TFmedic). We have further compared TFmedic with other state-of-the-arts methods (i.e., PolyPhen-2 and SIFT) on PolyPhen-2's own datasets, demonstrating that TFmedic can outperform the others. As applications, we have further applied TFmedic to all possible missense mutations on all human transcription factors; the proteome-wide results reveal interesting insights, consistent with the existing physiochemical knowledge. A case study with the actual 3D structure is conducted, revealing how TFmedic can be contributed to protein-DNA binding complex studies.
Ka-Chun Wong, Shankai Yan, Qiuzhen Lin, Xiangtao Li, Chengbin Peng 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2018 A scalable community detection algorithm for large graphs using stochastic block models
Chengbin Peng 0001, Ka-Chun Wong, Xiangliang Zhang 0001, David E. Keyes
Intell. Data Anal.1
2017 A scalable community detection algorithm for large graphs using stochastic block models
abstract
Community detection in graphs is widely used in social and biological networks, and the stochastic block model is a powerful probabilistic tool for describing graphs with community structures. However, in the era of “big data”, traditional inference algorithms for such a model are increasingly limi ted due to their high time complexity and poor scalability. In this paper, we propose a multi-stage maximum likelihood approach to recover the latent parameters of the stochastic block model, in time linear with respect to the number of edges. We also propose a parallel algorithm based on message passing. Our algorithm can overlap communication and computation, providing speedup without compromising accuracy as the number of processors grows. For example, to process a real-world graph with about 1.3 million nodes and 10 million edges, our algorithm requires about 6 seconds on 64 cores of a contemporary commodity Linux cluster. Experiments demonstrate that the algorithm can produce high quality results on both benchmark and real-world graphs. An example of finding more meaningful communities is illustrated consequently in comparison with a popular modularity maximization algorithm.
Chengbin Peng 0001, Ka-Chun Wong, Xiangliang Zhang 0001, David E. Keyes
Intell. Data Anal.1
2017 Evolving Transcription Factor Binding Site Models From Protein Binding Microarray Data
abstract
Protein binding microarray (PBM) is a high-throughput platform that can measure the DNA binding preference of a protein in a comprehensive and unbiased manner. In this paper, we describe the PBM motif model building problem. We apply several evolutionary computation methods and compare their performance with the interior point method, demonstrating their performance advantages. In addition, given the PBM domain knowledge, we propose and describe a novel method called kmerGA which makes domain-specific assumptions to exploit PBM data properties to build more accurate models than the other models built. The effectiveness and robustness of kmerGA is supported by comprehensive performance benchmarking on more than 200 datasets, time complexity analysis, convergence analysis, parameter analysis, and case studies. To demonstrate its utility further, kmerGA is applied to two real world applications: 1) PBM rotation testing and 2) ChIP-Seq peak sequence prediction. The results support the biological relevance of the models learned by kmerGA, and thus its real world applicability.
Ka-Chun Wong, Chengbin Peng 0001, Yue Li 0017
IEEE Trans. Cybern.2
2016 Identification of coupling DNA motif pairs on long-range chromatin interactions in human K562 cells
abstract
MOTIVATION: The protein-DNA interactions between transcription factors (TFs) and transcription factor binding sites (TFBSs, also known as DNA motifs) are critical activities in gene transcription. The identification of the DNA motifs is a vital task for downstream analysis. Unfortunately, the long-range coupling information between different DNA motifs is still lacking. To fill the void, as the first-of-its-kind study, we have identified the coupling DNA motif pairs on long-range chromatin interactions in human. RESULTS: The coupling DNA motif pairs exhibit substantially higher DNase accessibility than the background sequences. Half of the DNA motifs involved are matched to the existing motif databases, although nearly all of them are enriched with at least one gene ontology term. Their motif instances are also found statistically enriched on the promoter and enhancer regions. Especially, we introduce a novel measurement called motif pairing multiplicity which is defined as the number of motifs that are paired with a given motif on chromatin interactions. Interestingly, we observe that motif pairing multiplicity is linked to several characteristics such as regulatory region type, motif sequence degeneracy, DNase accessibility and pairing genomic distance. Taken into account together, we believe the coupling DNA motif pairs identified in this study can shed lights on the gene transcription mechanism under long-range chromatin interactions. AVAILABILITY AND IMPLEMENTATION: The identified motif pair data is compressed and available in the supplementary materials associated with this manuscript. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ka-Chun Wong, Yue Li 0017, Chengbin Peng 0001
Bioinform.3
2016 A Comparison Study for DNA Motif Modeling on Protein Binding Microarray
abstract
Transcription factor binding sites (TFBSs) are relatively short (5-15 bp) and degenerate. Identifying them is a computationally challenging task. In particular, protein binding microarray (PBM) is a high-throughput platform that can measure the DNA binding preference of a protein in a comprehensive and unbiased manner; for instance, a typical PBM experiment can measure binding signal intensities of a protein to all possible DNA k-mers (k = 8∼10). Since proteins can often bind to DNA with different binding intensities, one of the major challenges is to build TFBS (also known as DNA motif) models which can fully capture the quantitative binding affinity data. To learn DNA motif models from the non-convex objective function landscape, several optimization methods are compared and applied to the PBM motif model building problem. In particular, representative methods from different optimization paradigms have been chosen for modeling performance comparison on hundreds of PBM datasets. The results suggest that the multimodal optimization methods are very effective for capturing the binding preference information from PBM data. In particular, we observe a general performance improvement if choosing di-nucleotide modeling over mono-nucleotide modeling. In addition, the models learned by the best-performing method are applied to two independent applications: PBM probe rotation testing and ChIP-Seq peak sequence prediction, demonstrating its biological applicability.
Ka-Chun Wong, Yue Li 0017, Chengbin Peng 0001, Hau-San Wong
IEEE ACM Trans. Comput. Biol. Bioinform.3
2015 A Scalable Community Detection Algorithm for Large Graphs Using Stochastic Block Models
Chengbin Peng 0001, Ka-Chun Wong, Xiangliang Zhang 0001, David E. Keyes
IJCAI1
2015 SignalSpider: probabilistic pattern discovery on multiple normalized ChIP-Seq signal profiles
abstract
MOTIVATION: Chromatin immunoprecipitation (ChIP) followed by high-throughput sequencing (ChIP-Seq) measures the genome-wide occupancy of transcription factors in vivo. Different combinations of DNA-binding protein occupancies may result in a gene being expressed in different tissues or at different developmental stages. To fully understand the functions of genes, it is essential to develop probabilistic models on multiple ChIP-Seq profiles to decipher the combinatorial regulatory mechanisms by multiple transcription factors. RESULTS: In this work, we describe a probabilistic model (SignalSpider) to decipher the combinatorial binding events of multiple transcription factors. Comparing with similar existing methods, we found SignalSpider performs better in clustering promoter and enhancer regions. Notably, SignalSpider can learn higher-order combinatorial patterns from multiple ChIP-Seq profiles. We have applied SignalSpider on the normalized ChIP-Seq profiles from the ENCODE consortium and learned model instances. We observed different higher-order enrichment and depletion patterns across sets of proteins. Those clustering patterns are supported by Gene Ontology (GO) enrichment, evolutionary conservation and chromatin interaction enrichment, offering biological insights for further focused studies. We also proposed a specific enrichment map visualization method to reveal the genome-wide transcription factor combinatorial patterns from the models built, which extend our existing fine-scale knowledge on gene regulation to a genome-wide level. AVAILABILITY AND IMPLEMENTATION: The matrix-algebra-optimized executables and source codes are available at the authors' websites: http://www.cs.toronto.edu/∼wkc/SignalSpider.
Ka-Chun Wong, Yue Li 0017, Chengbin Peng 0001, Zhaolei Zhang
Bioinform.3
2015 Probabilistic Inference on Multiple Normalized Signal Profiles from Next Generation Sequencing: Transcription Factor Binding Sites
abstract
With the prevalence of chromatin immunoprecipitation (ChIP) with sequencing (ChIP-Seq) technology, massive ChIP-Seq data has been accumulated. The ChIP-Seq technology measures the genome-wide occupancy of DNA-binding proteins in vivo. It is well-known that different DNA-binding protein occupancies may result in a gene being regulated in different conditions (e.g. different cell types). To fully understand a gene's function, it is essential to develop probabilistic models on multiple ChIP-Seq profiles for deciphering the gene transcription causalities. In this work, we propose and describe two probabilistic models. Assuming the conditional independence of different DNA-binding proteins' occupancies, the first method (SignalRanker) is developed as an intuitive method for ChIP-Seq genome-wide signal profile inference. Unfortunately, such an assumption may not always hold in some gene regulation cases. Thus, we propose and describe another method (FullSignalRanker) which does not make the conditional independence assumption. The proposed methods are compared with other existing methods on ENCODE ChIP-Seq datasets, demonstrating its regression and classification ability. The results suggest that FullSignalRanker is the best-performing method for recovering the signal ranks on the promoter and enhancer regions. In addition, FullSignalRanker is also the best-performing method for peak sequence classification. We envision that SignalRanker and FullSignalRanker will become important in the era of next generation sequencing. FullSignalRanker program is available on the following website: http://www.cs.toronto.edu/~wkc/FullSignalRanker/.
Ka-Chun Wong, Chengbin Peng 0001, Yue Li 0017
IEEE ACM Trans. Comput. Biol. Bioinform.2
2012 Multiplicative Algorithms for Constrained Non-negative Matrix Factorization
abstract
Non-negative matrix factorization (NMF) provides the advantage of parts-based data representation through additive only combinations. It has been widely adopted in areas like item recommending, text mining, data clustering, speech denoising, etc. In this paper, we provide an algorithm that allows the factorization to have linear or approximately linear constraints with respect to each factor. We prove that if the constraint function is linear, algorithms within our multiplicative framework will converge. This theory supports a large variety of equality and inequality constraints, and can facilitate application of NMF to a much larger domain. Taking the recommender system as an example, we demonstrate how a specialized weighted and constrained NMF algorithm can be developed to fit exactly for the problem, and the tests justify that our constraints improve the performance for both weighted and unweighted NMF algorithms under several different metrics. In particular, on the Movie lens data with 94% of items, the Constrained NMF improves recall rate 3% compared to SVD50 and 45% compared to SVD150, which were reported as the best two in the top-N metric.
Chengbin Peng 0001, Ka-Chun Wong, Alyn P. Rockwood, Xiangliang Zhang 0001, Jinling Jiang, David E. Keyes
ICDM1
2012 Evolutionary multimodal optimization using the principle of locality
Ka-Chun Wong, Chun-Ho Wu, Ricky K. P. Mok, Chengbin Peng 0001, Zhaolei Zhang
Inf. Sci.4
2011 Generalizing and learning protein-DNA binding sequence representations by an evolutionary algorithm
Ka-Chun Wong, Chengbin Peng 0001, Man Hon Wong 0001, Kwong-Sak Leung
Soft Comput.2