VLDB 2026 Research / reviewers in the wild / expert
Yi-Ping Phoebe Chen
dblp:c/YPPhoebeChen · also Phoebe Chen, Yiping Phoebe Chen
· DBLP profile ↗
182ranked-venue papers
13as first author
80since 2021 · last 2026
0000-0002-4122-3767ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 61 · 2 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 61 · 4 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 3 first-author · 19 since 2021Databases, data management, data science and information retrieval · 31 · 5 first-author · 8 since 2021Computer networks · 7 · 4 since 2021Security and privacy · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-Aware Logical Reasoning via a Semiotic FrameworkabstractYunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang, Zikai Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034, Zikai Song |
ACL (1) | 6 |
| 2026 | Scalable User Admission Control in Large-Scale Cell-Free Massive MIMO
Yijie Gao, Peng Cheng 0002, Zhuo Chen 0001, Khoa Tran Phan, Wei Xiang 0001, Yi-Ping Phoebe Chen |
ICC | 6 |
| 2026 | Fine-grained food image classification with multi-modal collaboration
Linyi Lan, Zhongjie Xiao, Jianzhang Chen, Jiaxiong Lu, Huanrong Wang, Yi-Ping Phoebe Chen |
Neurocomputing | 7 |
| 2026 | EvoCap: Enhancing video captioning via self-evolving video-LLMs with knowledge consolidation
Yangliu Hu, Minye Wu, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034 |
Knowl. Based Syst. | 4 |
| 2026 | MF-PEAR-net: a multi-scale feature guided project-&-excite affine registration network for CT images
Ronald Bbosa, Kafui Efio-Akolly, Omar A. M. Salem, Yi-Ping Phoebe Chen |
Multim. Syst. | 5 |
| 2026 | TSC-Nodule: a tied non-local and spatial-channel reconstructed faster R-CNN for pulmonary nodule detection
Ronald Bbosa, Kafui Efio-Akolly, Omar A. M. Salem, Yi-Ping Phoebe Chen |
Neural Comput. Appl. | 5 |
| 2026 | Deep Imputation Bi-Stochastic Graph Regularized Matrix Factorization for Clustering Single-Cell RNA-Sequencing DataabstractBy generating massive gene transcriptome data and analyzing transcriptomic variations at the cell level, single-cell RNA-sequencing (scRNA-seq) technology has provided new way to explore cellular heterogeneity and functionality. Clustering scRNA-seq data could discover the hidden diversity and complexity of cell populations, which can aid to the identification of the disease mechanisms and biomarkers. In this paper, a novel method (DSINMF) is presented for clustering single cell RNA sequencing data by using deep matrix factorization. Our proposed method comprises four steps: first, the feature selection is utilized to remove irrelevant features. Then, the dropout imputation is used to handle missing value problem. Further, the dimension reduction is employed to preserve data characteristics and reduce noise effects. Finally, the deep matrix factorization with bi-stochastic graph regularization is used to obtain cluster results from scRNA-seq data. We compare DSINMF with other state-of-the-art algorithms on nine datasets and the results show our method outperformances than other methods. The code can be downloaded from https://github.com/lanbiolab/DSINMF. Wei Lan 0001, Qingfeng Chen, Jin Liu 0012, Jianxin Wang 0001, Yi-Ping Phoebe Chen |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2026 | Guest Editorial for the 21th Asia Pacific Bioinformatics Conference
Min Li 0007, Feng Luo 0001, Yi-Ping Phoebe Chen |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2026 | Learning-Based User Admission Control for Large-Scale Cell-Free Massive MIMOabstractCell-free massive multiple-input multiple-output (CF-mMIMO) is a promising architecture for 6G wireless networks through distributed access point (AP) cooperation. In large-scale deployments where user demand exceeds system capacity, effective user admission control (UAC) is essential to select users while meeting quality-of-service (QoS) requirements. The UAC problem in CF-mMIMO is inherently challenging, involving both discrete user selection and continuous power allocation variables. To address this challenge, we propose a Graphormer-enhanced Monte Carlo Tree Search (GE-MCTS) framework that integrates a Graphormer-based neural network (NN) with Monte Carlo Tree Search (MCTS). This framework leverages the Graphormer’s capability to model the graph-structured AP–user topology and MCTS’s planning proficiency to efficiently explore the vast decision space. Furthermore, to accommodate users initially unadmitted due to system constraints, we introduce a complementary AP deployment problem. By adapting the GE-MCTS framework, we optimize the placement of additional APs to achieve full user admission with the minimal number of new APs required. Simulation results demonstrate the effectiveness of our proposed framework. For UAC, with low computational complexity, GE-MCTS consistently admits 26.3–41.7% more users compared to baseline methods across various network scales. For AP deployment, our framework requires 45–73% fewer additional APs to achieve full user admission, highlighting its efficiency and scalability. Yijie Gao, Peng Cheng 0002, Zhuo Chen 0001, Khoa Tran Phan, Wei Xiang 0001, Yi-Ping Phoebe Chen |
IEEE Trans. Commun. | 6 |
| 2026 | DeDiff-4DGS: Fusing Temporal Correlations and Diffusion Priors for Dynamic 3D ScenesabstractReconstructing dynamic 3D (4D) scenes is challenging due to complex temporal dynamics and viewpoint sparsity in monocular videos. Existing extensions of 3D Gaussian Splatting (3D-GS) with its temporal modeling often fail to capture temporal correlations across frames, leading to redundant 3D Gaussians and reduced efficiency. To address this limitation, we propose DeDiff-4DGS, a framework that integrates temporal correlations and diffusion priors through two novel modules. The Temporal 3D Gaussian Latent Fusion (T3DLF) module fuses temporal information from sparse reference frames to promote spatio-temporal coherence and reduce the number of required 3D Gaussians. The Latent Diffusion Converter for 3D Gaussians (LDC3D) module enriches reference frames with semantic priors, complementing T3DLF under sparse-view conditions. Experimental results on standard benchmarks demonstrate that DeDiff-4DGS delivers higher reconstruction quality and improved efficiency over current state-of-the-art approaches. Hoang Nguyen Nguyen, Wei Xiang 0001, Kang Han, Phu Lai, Tianyu Chen 0004, Yi-Ping Phoebe Chen |
IEEE Trans. Multim. | 7 |
| 2025 | Temporal Coherent Object Flow for Multi-Object TrackingabstractMulti-object tracking is a challenging vision task that requires simultaneous reasoning about object detection and object association. Conventional solutions use frame as the basic unit and typically rely on a motion predictor that exploits the appearance features to associate detected candidates, leading to insufficient adaptability to long-term associations. In this study, we propose a section-based multi-object tracking approach that integrates a temporal coherent Object Flow Tracker (OFTrack), capable of achieving simultaneous multi-frame tracking by treating multiple consecutive frames as the basic processing unit, denoted as a “section”. Our OFTrack boosts the optical flow to the object flow by employing object perception and section-based motion estimation strategies. Object perception adopts object-aware sampling and scale-aware correlation to enable precise target discrimination. Motion estimation models the correlation of different objects in multi-frames via specialized temporal-spatial attention to achieve robust association in very long videos. Additionally, to address the oscillation of unpredictable trajectories in multi-frame estimation, we have designed temporal coherent enhancement including the trajectory masking pre-training and the smoothing constraint on trajectory curves. Comprehensive experiments on several widely used benchmarks demonstrate the superior performance of our approach. Zikai Song, Run Luo, Lintao Ma, Yi-Ping Phoebe Chen, Junqing Yu, Wei Yang 0034 |
AAAI | 5 |
| 2025 | SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained UnderstandingabstractVideo-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have demonstrated proficiency in providing the overall description of videos, they struggle with fine-grained understanding, particularly in aspects such as visual dynamics and video details inquiries. To tackle these shortcomings, we find that fine-tuning Video-LLMs on self-supervised fragment tasks, greatly improve their fine-grained video understanding abilities. Hence we propose two key contributions: (1) Self-Supervised Fragment Fine-Tuning (SF2T), a novel effortless fine-tuning method, employs the rich inherent characteristics of videos for training, while unlocking more fine-grained understanding ability of Video-LLMs. Moreover, it relieves researchers from labor-intensive annotations and smartly circumvents the limitations of natural language, which often fails to capture the complex spatiotemporal variations in videos; (2) A novel benchmark dataset, namely FineVidBench, for rigorously assessing Video-LLMs’ performance at both the scene and fragment levels, offering a comprehensive evaluation of their capabilities. We assessed multiple models and validated the effectiveness of SF2T on them. Experimental results reveal that our approach improves their ability to capture and interpret spatiotemporal details. Yangliu Hu, Zikai Song, Na Feng, Yawei Luo, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034 |
CVPR | 6 |
| 2025 | H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexingabstractsponsorship: This work was supported by a research grant from MeitY, Govt. of India, under the project BHASHINI. (MeitY, Govt. of India, under the project BHASHINI) Yi-Ping Phoebe Chen, Vipul Arora 0001 |
INTERSPEECH | 2 |
| 2025 | MVP: Winning Solution to SMP Challenge 2025 Video TrackabstractSocial media platforms serve as central hubs for content dissemination, opinion expression, and public engagement across diverse modalities. Accurately predicting the popularity of social media videos enables valuable applications in content recommendation, trend detection, and audience engagement. In this paper, we present Multimodal Video Predictor (MVP), our winning solution to the Video Track of the SMP Challenge 2025. MVP constructs expressive post representations by integrating deep video features extracted from pretrained models with user metadata and contextual information. The framework applies systematic preprocessing techniques, including log-transformations and outlier removal, to improve model robustness. A gradient-boosted regression model is trained to capture complex patterns across modalities. Our approach ranked first in the official evaluation of the Video Track, demonstrating its effectiveness and reliability for multimodal video popularity prediction on social platforms. The source code is available at https://github.com/yllhwa/SMPDVideo. Liliang Ye, Yunyao Zhang, Yafeng Wu, Yi-Ping Phoebe Chen, Junqing Yu, Wei Yang 0034, Zikai Song |
ACM Multimedia | 4 |
| 2025 | Establishing the Asia & Pacific Bioinformatics Joint Congress: a historic milestone in regional bioinformatics collaborationabstractIn response to the need for greater cohesion among regional conferences, the Asia Pacific Bioinformatics Network (APBioNET) set out in 2015 to realize a long-held aspiration-a single, unifying bioinformatics "super conference" for the Asia & Pacific community. Nearly a decade of persistence, coordination, and coalition-building led to the inaugural Asia & Pacific Bioinformatics Joint Congress (APBJC2024) in Okinawa, Japan. Now established as a triennial event, APBJC stands as a testament to the power of collective vision and shared purpose, offering a unifying platform for regional collaboration and scientific exchange. Tagline: Bringing a Region Together: The Making of APBJC. Asif M. Khan, Susumu Goto, Kenta Nakai, Limsoon Wong, Diane E. Kovats, Shinya Ikematsu, Yoshihiro Yamanishi, Nurul Salwanie Che Wahid, Pradeep Eranti, Yi-Ping Phoebe Chen, Tae-Min Kim, Shinn-Ying Ho, Jessica Cara Mar, Wataru Iwasaki 0001, Jayaraman Valadi, Prashanth Suravajhala, Christian Schönbach, Tin Wee Tan, Shoba Ranganathan, Kiyoko F. Aoki-Kinoshita |
Briefings Bioinform. | 12 |
| 2025 | Agri-LLM: Prompt-Based Large Language Model for Emission Data Analytics in Smart AgricultureabstractMassive emissions of greenhouse gases (GHGs) have a negative impact on the development of sustainable agriculture. While techniques of imputation and forecasting facilitate the observation of GHG emissions with improved accuracy, there is a lack of an integrated model for both GHG emission data imputation and forecasting, particularly in few-shot learning scenarios. To address this issue, this paper proposes a pre-trained large language model dubbed Agri-LLM for GHG emission data imputation and forecasting in smart agriculture. Notably, this model develops an information fusion embedding layer that fuses missing patterns, temporal irregularities and incomplete time series into multi-level patched tokens. A global temporal similarity informed prompting module is further elaborated on to generate suitable prompts for target time series, based on similar temporal characteristics captured from other nodes. Finally, the model aligns the pre-trained knowledge language with multi-level integrated tokens directly without altering the large language model’s backbone. The experimental studies demonstrate that our model outperforms state-of-the-art baselines in both tasks of imputation and forecasting using full-sample training. Extensive experiments also confirm that the Agri-LLM exhibits superior performance in few-shot learning scenarios and the effectiveness of each proposed model component. Le Fang 0001, Wei Xiang 0001, Jiong Jin, Kewen Liao, Chang Liu 0003, Yu Han 0003, Flora D. Salim, Yi-Ping Phoebe Chen |
IEEE Internet Things J. | 8 |
| 2025 | Spatiotemporal Pretrained Large Language Model for Forecasting With Missing ValuesabstractSpatiotemporal data collected by sensors within an urban Internet of Things (IoT) system inevitably contains some missing values, which significantly affects the accuracy of spatiotemporal data forecasting. However, existing techniques, including those based on Large Language Models (LLMs), show limited effectiveness in forecasting with missing values, especially in scenarios involving high-dimensional sensor data. In this article, we propose a novel spatiotemporal pre-trained large language model dubbed SPLLM for forecasting with missing values. In this network, we seamlessly integrate a specialized spatiotemporal fusion Graph Convolutional Network (GCN) module that extracts intricate spatiotemporal and graph-based information, for generating suitable inputs to the SPLLM. Furthermore, we propose a Feed-Forward Network (FFN) fine-tuning strategy within the LLM and a final fusion layer to enable the model to leverage the pre-trained foundational knowledge of the LLM and adapt to new incomplete data simultaneously. The experimental results indicate that SPLLM outperforms state-of-the-art models on real-world public datasets. Notably, SPLLM exhibits a superior performance in tackling incomplete sensory data with a variety of missing rates. A comprehensive ablation study of key components is conducted to demonstrate their efficiency. Le Fang 0001, Wei Xiang 0001, Shirui Pan, Flora D. Salim, Yi-Ping Phoebe Chen |
IEEE Internet Things J. | 5 |
| 2025 | TimeJudge: empowering video-LLMs as zero-shot judges for temporal consistency in video captionsabstractVideo large language models (video-LLMs) have demonstrated impressive capabilities in multimodal understanding, but their potential as zero-shot evaluators for temporal consistency in video captions remains underexplored. Existing methods notably underperform in detecting critical temporal errors, such as missing, hallucinated, or misordered actions. To address this gap, we introduce two key contributions. (1) TimeJudge: a novel zero-shot framework that recasts temporal error detection as answering calibrated binary question pairs. It incorporates modality-sensitive confidence calibration and uses consistency-weighted voting for robust prediction aggregation. (2) TEDBench: a rigorously constructed benchmark featuring videos across four distinct complexity levels, specifically designed with fine-grained temporal error annotations to evaluate video-LLM performance on this task. Through a comprehensive evaluation of multiple state-of-the-art video-LLMs on TEDBench, we demonstrate that TimeJudge consistently yields substantial gains in terms of recall and F1-score without requiring any task-specific fine-tuning. Our approach provides a generalizable, scalable, and training-free solution for enhancing the temporal error detection capabilities of video-LLMs. Yangliu Hu, Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2025 | The Large Language Models on Biomedical Data Analysis: A SurveyabstractWith the rapid development of Large Language Model (LLM) technology, it has become an indispensable force in biomedical data analysis research. However, biomedical researchers currently have limited knowledge about LLM. Therefore, there is an urgent need for a summary of LLM applications in biomedical data analysis. Herein, we propose this review by summarizing the latest research work on LLM in biomedicine. In this review, LLM techniques are first outlined. We then discuss biomedical datasets and frameworks for biomedical data analysis, followed by a detailed analysis of LLM applications in genomics, proteomics, transcriptomics, radiomics, single-cell analysis, medical texts and drug discovery. Finally, the challenges of LLM in biomedical data analysis are discussed. In summary, this review is intended for researchers interested in LLM technology and aims to help them understand and apply LLM in biomedical data analysis research. Wei Lan 0001, Zhentao Tang, Qingfeng Chen, Wei Peng 0004, Yi-Ping Phoebe Chen, Yi Pan 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | SCS-Voxel2Mesh: a Self-Calibrated Separable 3D Mesh Generation Network for Lung Nodule Spikes Classification and Malignancy PredictionabstractRadiologists use the standardized Lung-RADS clinical scoring criteria to assess and report spiculations/lobulations and sharp/curved spikes on the surface of lung nodules, because they are good predictors of lung cancer. Manual spiculation/lobulation annotation and classification is a tedious task for radiologists due to the nodule’s 3D geometry and 2D slice-by-slice assessment. This work presents SCS-Voxel2Mesh, a multi-class end-to-end deep learning model that segments pulmonary nodules (while preserving spikes), classifies spikes (sharp/spiculation and curved/lobulation) and performs malignancy prediction. Comprehensive experiments were conducted to evaluate and compare the performance of SCS-Voxel2Mesh with other segmentation and 3D mesh generation methods using the CIR dataset. The results demonstrate competitive performance of SCS-Voxel2Mesh on both the nodule spike classification and malignancy prediction tasks. Ronald Bbosa, Jia Ni Zou, Kafui Efio-Akolly, Yi-Ping Phoebe Chen, Wen Cai Huang |
BIBM | 4 |
| 2024 | PN-Quant: An Automated Pulmonary Nodule Quantification Method for Full-Size CT ScansabstractPulmonary nodule quantification is essential in forecasting and diagnosing potential malignant nodules, providing critical information for early intervention and treatment planning. However, most existing assistant diagnosis techniques primarily focus on the localization of lung nodules without comprehensive quantitative analysis, limiting their utility in clinical practice. To address this significant limitation, we present a novel and robust pulmonary nodule quantification framework named PN-Quant. It integrates a detection module, a segmentation module, and a quantification module to enable automated identification and precise measurement of lung nodules in full-size Computer Tomography (CT) scans, which facilitates the extraction of geometric characteristics, including volume, surface area, mass, sphericity, compactness, and elongation, offering valuable quantitative data for accurate nodule assessment. This study evaluates multiple PN-Quant pipelines with diverse configurations using datasets LIDC-IDRI, LNDb-19, and MSD-lung. Notably, the pipeline combining SANet and 3D UX-Net demonstrated superior performance, yielding low relative errors of 16.8%, 39.5%, and 24.4% on the respective datasets. These results underscore the effectiveness of the automated pipeline based on PN-Quant in efficiently and accurately quantifying pulmonary nodules across diverse datasets. The findings from this research highlight the potential of PN-Quant as a valuable tool for enhancing the precision and reliability of pulmonary nodule analysis in clinical settings, ultimately contributing to improved patient outcomes and clinical decision-making. Our source code is available at https://github.com/Xinkai-Tang/PN-Quant. Xinkai Tang, Shengjuan Guo, Yi-Ping Phoebe Chen, Wencai Huang, Jia Ni Zou |
BIBM | 5 |
| 2024 | Modeling Single-Cell ATAC-Seq Data Based on Contrastive Learning
Wei Lan 0001, Weihao Zhou, Qingfeng Chen, Ruiqing Zheng, Yi Pan 0001, Yi-Ping Phoebe Chen |
ISBRA (1) | 6 |
| 2024 | Diversity in MultimediaabstractMultimedia is a highly diverse discipline, a encompass of data types, research applications, research problems, approaches are integrated and implicated. Researchers in the domain get from very different specialties, applications ranging from digital health, transportation, energy, prediction, and discovery from different application domains. I will talk about how my research group could apply recent multimedia technologies into these importance of diversities fields. Yi-Ping Phoebe Chen |
ICMR | 1 |
| 2024 | Generative AI in Multimedia: Challenges and Opportunities for Academic and Industrial ImpactabstractGenerative AI has revolutionized multimedia, leading to groundbreaking developments in content creation, interactive experiences, and personalized media. This panel delves into the transformative potential of generative AI in academic and industrial sectors, exploring its future applications and connections to emerging techniques. Additionally, the panel will address newly identified opportunities and challenges from both technical and ethical perspectives, highlighting the importance of responsible AI development. Bringing together leading experts from universities, research institutions, and industry, this panel aims to foster discussion and debate among participants. We invite everyone to join and contribute to this critical and promising area of research in the multimedia community. Zi Huang, Yi-Ping Phoebe Chen, Shuicheng Yan |
ACM Multimedia | 2 |
| 2024 | Autogenic Language Embedding for Coherent Point TrackingabstractPoint tracking is a challenging task in computer vision, aiming to establish point-wise correspondence across long video sequences. Recent advancements have primarily focused on temporal modeling techniques to improve local feature similarity, often overlooking the valuable semantic consistency inherent in tracked points. In this paper, we introduce a novel approach leveraging language embeddings to enhance the coherence of frame-wise visual features related to the same object. Our proposed method, termed autogenic language embedding for visual feature enhancement, strengthens point correspondence in long-term sequences. Unlike existing visual-language schemes, our approach learns text embeddings from visual features through a dedicated mapping network, enabling seamless adaptation to various tracking tasks without explicit text annotations. Additionally, we introduce a consistency decoder that efficiently integrates text tokens into visual features with minimal computational overhead. Through enhanced visual consistency, our approach significantly improves tracking trajectories in lengthy videos with substantial appearance variations. Extensive experiments on widely-used tracking benchmarks demonstrate the superior performance of our method, showcasing notable enhancements compared to trackers relying solely on visual cues. The code will be available at https://github.com/SkyeSong38/ALTrack. Zikai Song, Run Luo, Lintao Ma, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034 |
ACM Multimedia | 6 |
| 2024 | DeepKEGG: a multi-omics data integration framework with biological insights for cancer recurrence prediction and biomarker discoveryabstractDeep learning-based multi-omics data integration methods have the capability to reveal the mechanisms of cancer development, discover cancer biomarkers and identify pathogenic targets. However, current methods ignore the potential correlations between samples in integrating multi-omics data. In addition, providing accurate biological explanations still poses significant challenges due to the complexity of deep learning models. Therefore, there is an urgent need for a deep learning-based multi-omics integration method to explore the potential correlations between samples and provide model interpretability. Herein, we propose a novel interpretable multi-omics data integration method (DeepKEGG) for cancer recurrence prediction and biomarker discovery. In DeepKEGG, a biological hierarchical module is designed for local connections of neuron nodes and model interpretability based on the biological relationship between genes/miRNAs and pathways. In addition, a pathway self-attention module is constructed to explore the correlation between different samples and generate the potential pathway feature representation for enhancing the prediction performance of the model. Lastly, an attribution-based feature importance calculation method is utilized to discover biomarkers related to cancer recurrence and provide a biological interpretation of the model. Experimental results demonstrate that DeepKEGG outperforms other state-of-the-art methods in 5-fold cross validation. Furthermore, case studies also indicate that DeepKEGG serves as an effective tool for biomarker discovery. The code is available at https://github.com/lanbiolab/DeepKEGG. Wei Lan 0001, Haibo Liao, Qingfeng Chen, Lingzhi Zhu, Yi Pan 0001, Yi-Ping Phoebe Chen |
Briefings Bioinform. | 6 |
| 2024 | The integration of knowledge graph convolution network with denoising autoencoderabstractThe knowledge graph convolution network (KGCN) is a recommendation model that provides a set of top recommendations based on knowledge graph developed between users, items, and their attributes. In this study, we integrate the KGCN model with denoising autoencoder (DAE) to improve its recommendation performance. A trained DAE is used to sample K-dimensional latent representation for each user, which then transforms that representation to generate a probability distribution over items. The relationship between acquired latent representation and the meta features is modelled using multivariate multiple regression (MMR) kernel. As a result, without the need for new configuration assessments, performance estimation of new data is pursued directly through MMR and the decoder of DAE. Empirically, we demonstrate that on real-world datasets, the proposed method substantially outperforms other state-of-the-art baselines. Movie-Lens 100K (ML-100K) and Movie-Lens 1M (ML-1M), two common MovieLens datasets, are used to verify the accuracy of the proposed approach. The results from experiments show significant improvement of 41.17% when the proposed method is applied on KGCN model. The proposed framework outperforms other state-of-the-art frameworks on Recall@K and normalized discounted cumulative gain (NDCG@K) metrics by achieving higher scores for Recall@5, Recall@10, NDCG@1, and NDCG@10. • The KGCN is a recommendation model that provides a set of top recommendations. • Our framework (KGCN-DAE) integrates the KGCN model with denoising autoencoder. • The KGCN-DAE improves the performance and efficiency of the KGCN model. Gurinder Kaur, Fei Liu 0003, Yi-Ping Phoebe Chen |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Validating Siamese embedded neural networks with identical representations for efficient model convergenceabstractDeep learning, neural architecture search, reinforcement learning, and embedded learning use scaled data, hyperspace, reward, and similarity to produce efficient converging neural networks. However, all these state-of-the-art frameworks require training to evaluate whether the examples are identical. Therefore, we propose our Probabilistic Asymmetric Convergence Network with Validation and Transform-Learning (PACNVT) framework that learns with fewer data, reduces the hyperspace with our validation technique and modified tangent activation function, reinforces the learning with our transform learning algorithms, and ascertains similarity independent of spatial and mask for transformational consistency. Furthermore, our framework generates neural networks that yield state-of-the-art accuracies on MNIST, OMNIGLOT, and CIFAR datasets. Moreover, representing regression as a similarity problem unveils previously unseen residual fluid intelligence patterns, yielding 99.92% accuracy with as little as 20 pairs. Mathias Hoy Talbo, Haishuai Wang, Lianhua Chi, Yi-Ping Phoebe Chen |
Knowl. Based Syst. | 4 |
| 2024 | Guest Editorial Guest Editorial for the 20th Asia Pacific Bioinformatics ConferenceabstractThe four papers in this special section were presented at the 20th Asia Pacific Bioinformatics Conference (APBC), which was held in Malaysia 26-28 April 2022. Su Datt Lam, Wai Keat Yam, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | LGCDA: Predicting CircRNA-Disease Association Based on Fusion of Local and Global FeaturesabstractCircRNA has been shown to be involved in the occurrence of many diseases. Several computational frameworks have been proposed to identify circRNA-disease associations. Despite the existing computational methods have obtained considerable successes, these methods still require to be improved as their performance may degrade due to the sparsity of the data and the problem of memory overflow. We develop a novel computational framework called LGCDA to predict circRNA-disease associations by fusing local and global features to solve the above mentioned problems. First, we construct closed local subgraphs by using k-hop closed subgraph and label the subgraphs to obtain rich graph pattern information. Then, the local features are extracted by using graph neural network (GNN). In addition, we fuse Gaussian interaction profile (GIP) kernel and cosine similarity to obtain global features. Finally, the score of circRNA-disease associations is predicted by using the multilayer perceptron (MLP) based on local and global features. We perform five-fold cross validation on five datasets for model evaluation and our model surpasses other advanced methods. Wei Lan 0001, Qingfeng Chen, Ning Yu 0004, Yi Pan 0001, Yu Zheng 0013, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2024 | Semantic Distance Adversarial Learning for Text-to-Image SynthesisabstractText-to-Image (T2I) synthesis is a cross-modality task that requires a text description as input to generate a realistic and semantically consistent image. To guarantee semantic consistency, previous studies regenerate text descriptions from synthetic images and align them with the given descriptions. However, the existing redescription modules lack explicit modeling of their training objectives, which is crucial for reliable measurement of semantic distance between redescriptions and given text inputs. Consequently, the aligned text redescriptions suffer from training bias caused by the emergence of adversarial image samples, unseen semantics, and mistaken contents from low-quality synthesized images. To this end, we propose a SEMantic distance Adversarial learning (SEMA) framework for Text-to-Image synthesis which strengthens semantic consistency from two aspects: 1) We introduce adversarial learning between the image generator and the text redescription module to mutually promote or demote the quality of generated image or text instances. This learning model ensures accurate redescription of image contents, thus diminishing the generation of adversarial image samples. 2) We introduce two-fold semantic distance discrimination (SEM distance) to characterize semantic relevance between matching text or image pairs. The unseen semantics and mistaken contents will be penalized with a large SEM distance. The proposed discrimination method also simplifies the model training process with no need to optimize multiple discriminators. Experimental results on CUB Birds 200 and MS-COCO datasets show that the proposed model outperforms the state-of-the-art methods. Yefei Sheng, Bing-Kun Bao, Yi-Ping Phoebe Chen, Changsheng Xu |
IEEE Trans. Multim. | 4 |
| 2024 | Correlation-Aware Spatial-Temporal Graph Learning for Multivariate Time-Series Anomaly DetectionabstractMultivariate time-series anomaly detection is critically important in many applications, including retail, transportation, power grid, and water treatment plants. Existing approaches for this problem mostly employ either statistical models which cannot capture the nonlinear relations well or conventional deep learning (DL) models e.g., convolutional neural network (CNN) and long short-term memory (LSTM) that do not explicitly learn the pairwise correlations among variables. To overcome these limitations, we propose a novel method, correlation-aware spatial-temporal graph learning (termed ), for time-series anomaly detection. explicitly captures the pairwise correlations via a correlation learning (MTCL) module based on which a spatial-temporal graph neural network (STGNN) can be developed. Then, by employing a graph convolution network (GCN) that exploits one-and multihop neighbor information, our STGNN component can encode rich spatial information from complex pairwise dependencies between variables. With a temporal module that consists of dilated convolutional functions, the STGNN can further capture long-range dependence over time. A novel anomaly scoring component is further integrated into to estimate the degree of an anomaly in a purely unsupervised manner. Experimental results demonstrate that can detect and diagnose anomalies effectively in general settings as well as enable early detection across different time delays. Our code is available at https://github.com/huankoh/CST-GL. Yu Zheng 0013, Huan Yee Koh, Ming Jin 0005, Lianhua Chi, Khoa Tran Phan, Shirui Pan, Yi-Ping Phoebe Chen, Wei Xiang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Compact Transformer Tracker with Correlative Masked ModelingabstractTransformer framework has been showing superior performances in visual object tracking for its great strength in information aggregation across the template and search image with the well-known attention mechanism. Most recent advances focus on exploring attention mechanism variants for better information aggregation. We find these schemes are equivalent to or even just a subset of the basic self-attention mechanism. In this paper, we prove that the vanilla self-attention structure is sufficient for information aggregation, and structural adaption is unnecessary. The key is not the attention structure, but how to extract the discriminative feature for tracking and enhance the communication between the target and search image. Based on this finding, we adopt the basic vision transformer (ViT) architecture as our main tracker and concatenate the template and search image for feature embedding. To guide the encoder to capture the invariant feature for tracking, we attach a lightweight correlative masked decoder which reconstructs the original template and search image from the corresponding masked tokens. The correlative masked decoder serves as a plugin for the compact transformer tracker and is skipped in inference. Our compact tracker uses the most simple structure which only consists of a ViT backbone and a box head, and can run at 40 fps. Extensive experiments show the proposed compact transform tracker outperforms existing approaches, including advanced attention variants, and demonstrates the sufficiency of self-attention in tracking tasks. Our method achieves state-of-the-art performance on five challenging datasets, along with the VOT2020, UAV123, LaSOT, TrackingNet, and GOT-10k benchmarks. Our project is available at https://github.com/HUSTDML/CTTrack. Zikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034 |
AAAI | 4 |
| 2023 | SoK: Systematizing Attack Studies in Federated Learning - From Sparseness to CompletenessabstractFederated Learning (FL) is a machine learning technique that enables multiple parties to collaboratively train a model using their private datasets. Given its decentralized nature, FL has inherent vulnerabilities that make it susceptible to adversarial attacks. The success of an attack on FL depends upon several (latent) factors, including the adversary’s strength, the chosen attack strategy, and the effectiveness of the defense measures in place. There is a growing body of literature on empirical attack studies on FL, but no systematic way to compare and evaluate the completeness of these studies, which raises questions about their validity. To address this problem, we introduce a causal model that captures the relationship between the different (latent) factors, and their reflexive indicators, that can impact the success of an attack on FL. The proposed model, inspired by structural equation modeling, helps systematize the existing literature on FL attack studies and provides a way to compare and contrast their completeness. We validate the model and demonstrate its utility through experimental evaluation of select attack studies. Our aim is to help researchers in the FL domain design more complete attack studies and improve the understanding of FL vulnerabilities. Geetanjli Sharma, Mahawaga Arachchige Pathum Chamikara, Mohan Baruwal Chhetri, Yi-Ping Phoebe Chen |
AsiaCCS | 4 |
| 2023 | DFNodule: a Novel Deformable Faster R-CNN for Lung Nodule DetectionabstractIn computer-aided diagnosis systems, lung nodule detection plays a crucial role in the overall framework. In this work, we propose a new three-dimensional deformable convolutional neural network (dcnn) method for lung nodule detection based on the Faster R-CNN framework. We incorporate deformable convolutions to design a hybrid convolutional module, which enhances feature extraction in the lung nodule detection model. By leveraging the deformable convolutions’ characteristics, the network is capable of capturing the diverse morphological variations of lung nodules, addressing challenges such as large morphological variations and the inability to capture unified image features. This improves the accuracy of the lung nodule detection algorithm. Additionally, we employ a second-stage network to further discriminate suspected nodules, which enhances the recognition of non-nodule tissues and reduces false positive nodules. To comprehensively evaluate the performance of various lung nodule detection models, we conducted experiments using the publicly available LUNA16 dataset. Our method surpasses other detection algorithms in terms of CPM, achieving a 1.5% improvement. Particularly, the nodule recognition rate is significantly improved at lower false positive rates. In addition, in other metrics such as F1-score, AP, we also achieved 0.6%, 2% improvement. GuoWei Tao, Fu Zhou, Hao Gui, Fei Luo 0004, Wen Cai Huang, Jia Ni Zou, Yi-Ping Phoebe Chen |
BIBM | 8 |
| 2023 | Thematic relations outperform taxonomic relations in a cued recall task
Weijia Cao, Omri Raccah, Yi-Ping Phoebe Chen, David Poeppel |
CogSci | 3 |
| 2023 | DeepMNF: Deep Multimodal Neuroimaging Framework for Diagnosing Autism Spectrum Disorder
Syed Qasim Abbas, Lianhua Chi, Yi-Ping Phoebe Chen |
Artif. Intell. Medicine | 3 |
| 2023 | Benchmarking of computational methods for predicting circRNA-disease associationsabstractAccumulating evidences demonstrate that circular RNA (circRNA) plays an important role in human diseases. Identification of circRNA-disease associations can help for the diagnosis of human diseases, while the traditional method based on biological experiments is time-consuming. In order to address the limitation, a series of computational methods have been proposed in recent years. However, few works have summarized these methods or compared the performance of them. In this paper, we divided the existing methods into three categories: information propagation, traditional machine learning and deep learning. Then, the baseline methods in each category are introduced in detail. Further, 5 different datasets are collected, and 14 representative methods of each category are selected and compared in the 5-fold, 10-fold cross-validation and the de novo experiment. In order to further evaluate the effectiveness of these methods, six common cancers are selected to compare the number of correctly identified circRNA-disease associations in the top-10, top-20, top-50, top-100 and top-200. In addition, according to the results, the observation about the robustness and the character of these methods are concluded. Finally, the future directions and challenges are discussed. Wei Lan 0001, Qingfeng Chen, Jin Liu 0012, Jianxin Wang 0001, Yi-Ping Phoebe Chen |
Briefings Bioinform. | 8 |
| 2023 | Applying staged event-driven access control to combat ransomwareabstractThe advancement of modern Operating Systems (OSs), and the popularity of personal computing devices with Internet connectivity, have facilitated the proliferation of ransomware attacks. Ransomware has evolved from executable programs encrypting user files, to novel attack vectors including fileless command scripts, information exfiltration and human-operated ransomware. Many anti-ransomware studies have been published, but many of them assumed newer ransomware variants only performed file encryption, were similar to existing variants, and often did not consider those novel attack vectors. We have defined an updated ransomware threat model to include those novel attack vectors, and redefined false positives and false negatives in the context of ransomware mitigation. We proposed to apply both program-centric and user-centric access control to combat ransomware, but only delegate access control decisions that users are capable of making to users, while enforcing non-negotiable access control decisions by OS and software developers. We have designed a Staged Event-Driven Access Control (SEDAC) approach to incorporate both program-centric and user-centric access control measures, and demonstrated a prototype on Windows OS. Our prototype was able to intercept more types of ransomware attack vectors than existing proposals. We hope to convince OS and software architects to incorporate our design to better combat ransomware. Timothy R. McIntosh, A. S. M. Kayes, Yi-Ping Phoebe Chen, Alex Ng, Paul A. Watters |
Comput. Secur. | 3 |
| 2023 | Transformed domain convolutional neural network for Alzheimer's disease diagnosis using structural MRI
Syed Qasim Abbas, Lianhua Chi, Yi-Ping Phoebe Chen |
Pattern Recognit. | 3 |
| 2023 | eX-ViT: A Novel explainable vision transformer for weakly supervised semantic segmentationabstractRecently vision transformer models have become prominent models for a multitude of vision tasks. These models, however, are usually opaque with weak feature interpretability, making their predictions inaccessible to the users. While there has been a surge of interest in the development of post-hoc solutions that explain model decisions, these methods can not be broadly applied to different transformer architectures, as rules for interpretability have to change accordingly based on the heterogeneity of data and model structures. Moreover, there is no method currently built for an intrinsically interpretable transformer, which is able to explain its reasoning process and provide a faithful explanation. To close these crucial gaps, we propose a novel vision transformer dubbed the eXplainable Vision Transformer (eX-ViT), an intrinsically interpretable transformer model that is able to jointly discover robust interpretable features and perform the prediction. Specifically, eX-ViT is composed of the Explainable Multi-Head Attention (E-MHA) module, the Attribute-guided Explainer (AttE) module with the self-supervised attribute-guided loss. The E-MHA tailors explainable attention weights that are able to learn semantically interpretable representations from tokens in terms of model decisions with noise robustness. Meanwhile, AttE is proposed to encode discriminative attribute features for the target object through diverse attribute discovery, which constitutes faithful evidence for the model predictions. Additionally, we have developed a self-supervised attribute-guided loss for our eX-ViT architecture, which utilizes both the attribute discriminability mechanism and the attribute diversity mechanism to enhance the quality of learned representations. As a result, the proposed eX-ViT model can produce faithful and robust interpretations with a variety of learned attributes. To verify and evaluate our method, we apply the eX-ViT to several weakly supervised semantic segmentation (WSSS) tasks, since these tasks typically rely on accurate visual explanations to extract object localization maps. Particularly, the explanation results obtained via eX-ViT are regarded as pseudo segmentation labels to train WSSS models. Comprehensive simulation results illustrate that our proposed eX-ViT model achieves comparable performance to supervised baselines, while surpassing the accuracy and interpretability of state-of-the-art black-box methods using only image-level labels. Wei Xiang 0001, Juan Fang 0004, Yi-Ping Phoebe Chen, Lianhua Chi |
Pattern Recognit. | 4 |
| 2023 | A Fast and More Accurate Seed-and-Extension Density-Based Clustering AlgorithmabstractClustering algorithms have been widely studied in many scientific areas, such as data mining, knowledge discovery, bioinformatics and machine learning. A density-based clustering algorithm, called density peaks (DP), which was proposed by Rodriguez and Laio, outperform almost all other approaches. Although the DP algorithm performs well in many cases, there is still room for improvement in the precision of its output clusters as well as the quality of the selected centers. In this study, we propose a more accurate clustering algorithm, seed-and-extension-based density peaks (SDP). SDP selects the centers that hold the features of their clusters while building a spanning forest, and meanwhile, constructs the output clusters in a seed-and-extension manner. Experiment results demonstrate the effectiveness of SDP, especially when dealing with clusters with relatively high densities. Precisely, we show that SDP is more accurate than the DP algorithm as well as other state-of-the-art clustering approaches concerning the quality of both output clusters and cluster centers while maintaining similar running time of the DP algorithm, particularly for a variety of time-series (i.e. non-metric) data. Moreover, SDP outperforms DP in the dynamic model in which data point insertion and deletion are allowed. From a practical perspective, the proposed SDP algorithm is obviously helpful to many application problems. Ming-Hao Tung, Yi-Ping Phoebe Chen, Chen-Yu Liu, Chung-Shou Liao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Generative and Contrastive Self-Supervised Learning for Graph Anomaly DetectionabstractAnomaly detection from graph data has drawn much attention due to its practical significance in many critical applications including cybersecurity, finance, and social networks. Existing data mining and machine learning methods are either shallow methods that could not effectively capture the complex interdependency of graph data or graph autoencoder methods that could not fully exploit the contextual information as supervision signals for effective anomaly detection. To overcome these challenges, in this paper, we propose a novel method, Self-Supervised Learning for Graph Anomaly Detection (SL-GAD). Our method constructs different contextual subgraphs (views) based on a target node and employs two modules,generative attribute regressionandmulti-view contrastive learningfor anomaly detection. While thegenerative attribute regressionmodule allows us to capture the anomalies in the attribute space, themulti-view contrastive learningmodule can exploit richer structure information from multiple subgraphs, thus abling to capture the anomalies in the structure space, mixing of structure, and attribute information. We conduct extensive experiments on six benchmark datasets and the results demonstrate that our method outperforms state-of-the-art methods by a large margin. Yu Zheng 0013, Ming Jin 0005, Yixin Liu 0001, Lianhua Chi, Khoa Tran Phan, Yi-Ping Phoebe Chen |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | Transformer Tracking with Cyclic Shifting Window AttentionabstractTransformer architecture has been showing its great strength in visual object tracking, for its effective attention mechanism. Existing transformer-based approaches adopt the pixel-to-pixel attention strategy on flattened image features and unavoidably ignore the integrity of ob-jects. In this paper, we propose a new transformer ar-chitecture with multi-scale cyclic shifting window attention for visual object tracking, elevating the attention from pixel to window level. The cross-window multi-scale at-tention has the advantage of aggregating attention at dif-ferent scales and generates the best fine-scale match for the target object. Furthermore, the cyclic shifting strat-egy brings greater accuracy by expanding the window sam-ples with positional information, and at the same time saves huge amounts of computational power by removing redun-dant calculations. Extensive experiments demonstrate the superior performance of our method, which also sets the new state-of-the-art records on five challenging datasets, along with the VOT2020, UAV123, LaSOT, TrackingNet, and GOT-lOk benchmarks. Our project is available at https://github.com/SkyeSong38/CSWinTT. Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034 |
CVPR | 3 |
| 2022 | Balanced Contrastive Learning for Long-Tailed Visual RecognitionabstractReal-world data typically follow a long-tailed distribution, where a few majority categories occupy most of the data while most minority categories contain a limited number of samples. Classification models minimizing crossentropy struggle to represent and classify the tail classes. Although the problem of learning unbiased classifiers has been well studied, methods for representing imbalanced data are under-explored. In this paper, we focus on representation learning for imbalanced data. Recently, supervised contrastive learning has shown promising performance on balanced data recently. However, through our theoretical analysis, we find that for long-tailed data, it fails to form a regular simplex which is an ideal geometric configuration for representation learning. To correct the optimization behavior of SCL and further improve the performance of long-tailed visual recognition, we propose a novel loss for balanced contrastive learning (BCL). Compared with SCL, we have two improvements in BCL: classaveraging, which balances the gradient contribution of negative classes; class-complement, which allows all classes to appear in every mini-batch. The proposed balanced contrastive learning (BCL) method satisfies the condition of forming a regular simplex and assists the optimization of cross-entropy. Equipped with BCL, the proposed two-branch framework can obtain a stronger feature representation and achieve competitive performance on long-tailed benchmark datasets such as CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, and iNaturalist2018. Jianggang Zhu, Zheng Wang 0059, Jingjing Chen 0001, Yi-Ping Phoebe Chen, Yu-Gang Jiang 0001 |
CVPR | 4 |
| 2022 | Aggregated Bidirectional Local Binary Pattern for Robust Perceptual Image HashingabstractEasy access and speedy advancements to image modification tools have made it tough for researchers and experts working in the domain of image integrity verification. It has ever become a grueling job in present times. Robust Perceptual Image Hashing (RPIH) is one of the available approaches to verify image integrity. RPIH algorithms are developed for extracting a specific group of designated features from a query image to produce a consolidated representation, which can be utilized to verify image integrity. In this paper, a novel aggregated bidirectional local binary pattern RPIH algorithm is proposed, to compute the image hash, which constitutes fixed-length data. We combine Local binary Pattern (LBP) and Reverse Local Binary Pattern (RLBP) histograms to generate Aggregated Bidirectional Local Binary Pattern (AB-LBP) features for creating a hash vector. The experimental results demonstrate the usefulness of the proposed AB-LBP algorithm as a RPIH scheme and its superior computation efficiency. Additionally, the AB-LBP scheme conveys information to localize the tamper location, which is handy in rejecting the selective part of the tampered image. Syed Qasim Abbas, S. Jannat Shirazi, Yi-Ping Phoebe Chen |
ISM | 3 |
| 2022 | Fuzzy Information Measures Feature Selection Using Descriptive Statistics Data
Omar A. M. Salem, Yi-Ping Phoebe Chen, Xi Chen 0078 |
KSEM (3) | 4 |
| 2022 | KGANCDA: predicting circRNA-disease associations based on knowledge graph attention networkabstractIncreasing evidences have proved that circRNA plays a significant role in the development of many diseases. In addition, many researches have shown that circRNA can be considered as the potential biomarker for clinical diagnosis and treatment of disease. Some computational methods have been proposed to predict circRNA-disease associations. However, the performance of these methods is limited as the sparsity of low-order interaction information. In this paper, we propose a new computational method (KGANCDA) to predict circRNA-disease associations based on knowledge graph attention network. The circRNA-disease knowledge graphs are constructed by collecting multiple relationship data among circRNA, disease, miRNA and lncRNA. Then, the knowledge graph attention network is designed to obtain embeddings of each entity by distinguishing the importance of information from neighbors. Besides the low-order neighbor information, it can also capture high-order neighbor information from multisource associations, which alleviates the problem of data sparsity. Finally, the multilayer perceptron is applied to predict the affinity score of circRNA-disease associations based on the embeddings of circRNA and disease. The experiment results show that KGANCDA outperforms than other state-of-the-art methods in 5-fold cross validation. Furthermore, the case study demonstrates that KGANCDA is an effective tool to predict potential circRNA-disease associations. Wei Lan 0001, Qingfeng Chen, Ruiqing Zheng, Jin Liu 0012, Yi Pan 0001, Yi-Ping Phoebe Chen |
Briefings Bioinform. | 7 |
| 2022 | Fuzzy joint mutual information feature selection based on ideal vector
Omar A. M. Salem, Yi-Ping Phoebe Chen, Ahmed Hamed Attia, Xi Chen 0078 |
Expert Syst. Appl. | 3 |
| 2022 | Advanced calibration of mortality prediction on cardiovascular disease using feature-based artificial neural network
Alessio Bonti, Lianhua Chi, Mohamed Almorsy, Yi-Ping Phoebe Chen |
Expert Syst. Appl. | 5 |
| 2022 | A blockchain-based application for genomic access and variant discovery using smart contracts and homomorphic encryption
Abukari Mohammed Yakubu, Yi-Ping Phoebe Chen |
Future Gener. Comput. Syst. | 2 |
| 2022 | A comprehensive review of federated learning for COVID-19 detectionabstractThe coronavirus of 2019 (COVID-19) was declared a global pandemic by World Health Organization in March 2020. Effective testing is crucial to slow the spread of the pandemic. Artificial intelligence and machine learning techniques can help COVID-19 detection using various clinical symptom data. While deep learning (DL) approach requiring centralized data is susceptible to a high risk of data privacy breaches, federated learning (FL) approach resting on decentralized data can preserve data privacy, a critical factor in the health domain. This paper reviews recent advances in applying DL and FL techniques for COVID-19 detection with a focus on the latter. A model FL implementation use case in health systems with a COVID-19 detection using chest X-ray image data sets is studied. We have also reviewed applications of previously published FL experiments for COVID-19 research to demonstrate the applicability of FL in tackling health research issues. Last, several challenges in FL implementation in the healthcare domain are discussed in terms of potential future work. Sadaf Naz, Khoa Tran Phan, Yi-Ping Phoebe Chen |
Int. J. Intell. Syst. | 3 |
| 2022 | GANLDA: Graph attention network for lncRNA-disease associations prediction
Wei Lan 0001, Ximin Wu, Qingfeng Chen, Wei Peng 0004, Jianxin Wang 0001, Yi-Ping Phoebe Chen |
Neurocomputing | 6 |
| 2022 | Effective fuzzy joint mutual information feature selection based on uncertainty region for classification problem
Omar A. M. Salem, Yi-Ping Phoebe Chen, Ahmed Hamed Attia, Xi Chen 0078 |
Knowl. Based Syst. | 3 |
| 2022 | A novel explainable neural network for Alzheimer's disease diagnosis
Wei Xiang 0001, Juan Fang 0004, Yi-Ping Phoebe Chen, Ruifeng Zhu |
Pattern Recognit. | 4 |
| 2022 | IGNSCDA: Predicting CircRNA-Disease Associations Based on Improved Graph Convolutional Network and Negative SamplingabstractAccumulating evidences have shown that circRNA plays an important role in human diseases. It can be used as potential biomarker for diagnose and treatment of disease. Although some computational methods have been proposed to predict circRNA-disease associations, the performance still need to be improved. In this paper, we propose a new computational model based on Improved Graph convolutional network and Negative Sampling to predict CircRNA-Disease Associations. In our method, it constructs the heterogeneous network based on known circRNA-disease associations. Then, an improved graph convolutional network is designed to obtain the feature vectors of circRNA and disease. Further, the multi-layer perceptron is employed to predict circRNA-disease associations based on the feature vectors of circRNA and disease. In addition, the negative sampling method is employed to reduce the effect of the noise samples, which selects negative samples based on circRNA's expression profile similarity and Gaussian Interaction Profile kernel similarity. The 5-fold cross validation is utilized to evaluate the performance of the method. The results show that IGNSCDA outperforms than other state-of-the-art methods in the prediction performance. Moreover, the case study shows that IGNSCDA is an effective tool for predicting potential circRNA-disease associations. Wei Lan 0001, Qingfeng Chen, Jin Liu 0012, Jianxin Wang 0001, Yi-Ping Phoebe Chen, Shirui Pan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2022 | LDICDL: LncRNA-Disease Association Identification Based on Collaborative Deep LearningabstractIt has been proved that long noncoding RNA (lncRNA) plays critical roles in many human diseases. Therefore, inferring associations between lncRNAs and diseases can contribute to disease diagnosis, prognosis and treatment. To overcome the limitation of traditional experimental methods such as expensive and time-consuming, several computational methods have been proposed to predict lncRNA-disease associations by fusing different biological data. However, the prediction performance of lncRNA-disease associations identification needs to be improved. In this study, we propose a computational model (named LDICDL) to identify lncRNA-disease associations based on collaborative deep learning. It uses an automatic encoder to denoise multiple lncRNA feature information and multiple disease feature information, respectively. Then, the matrix decomposition algorithm is employed to predict the potential lncRNA-disease associations. In addition, to overcome the limitation of matrix decomposition, the hybrid model is developed to predict associations between new lncRNA (or disease) and diseases (or lncRNA). The ten-fold cross validation and de novo test are applied to evaluate the performance of method. The experimental results show LDICDL outperforms than other state-of-the-art methods in prediction performance. Wei Lan 0001, Dehuan Lai, Qingfeng Chen, Ximin Wu, Baoshan Chen, Jin Liu 0012, Jianxin Wang 0001, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2022 | EditorialabstractPresents the introductory editorial for this issue of the publication. Hsiao-Fang Sunny Sun, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Multi-Classes and Motion Properties for Concurrent Visual SLAM in Dynamic EnvironmentsabstractWorking in a dynamic environment is a challenging problem for visual simultaneous localization and mapping (visual SLAM). Most of the existing visual SLAM algorithms fail resulting in significant error or losing in tracking when moving objects dominate the scene. We found two reasons cause these issues: (i) Previous approaches use information from all regions in the image; (ii) Existing algorithms use just two groups and block all feature points from moveable objects. In this paper, we propose a novel Multi-classes and motion properties for Concurrent Visual SLAM (MCV-SLAM) algorithm, which defines classes into five categories and concurrently fuses prior knowledge and observation of moving objects with semantic segmentation to ensure visual SLAM works properly for dynamic environments in real time. We also propose an adaptive method to optimize camera pose by using more potential inlier feature points with continuous weights, while eliminating the impact of moving objects. Our experiments are performed on public datasets of both indoor and outdoor scenes with moving objects in dynamic environments. The experimental results demonstrate that our method outperforms previous works with greater robustness and smaller tracking errors, and our MCV-SLAM can deal with the situations (i.e., the dominance of moving objects, lack of matching points), which lead misestimating occurs in existing SLAMs. Bohong Yang, Wu Ran, Lin Wang 0033, Hong Lu 0001, Yi-Ping Phoebe Chen |
IEEE Trans. Multim. | 5 |
| 2021 | HAUNet-3D: a Novel Hierarchical Attention 3D UNet for Lung Nodule SegmentationabstractUNet and its extended versions are the most used networks in the lung nodule segmentation from CT images. However, current UNet-like methods still suffer from some problems: 1) The heterogeneity of lung nodules affect the segmentation performance; 2) the mixture of lung nodules and their surrounding tissues in the CT image increases the segmentation difficulty. To address these issues, we propose a novel hierarchical attention 3D UNet named HAUNet-3D. It introduces the attention mechanism at multiple scales and organizes them in a bottom-up hierarchical connection way. Such a proposition could better capture features with various sizes and guide the fusion of features from adjacent attention outputs without losing the advantages of 3D UNet. In experiment, our method has been extensively evaluated on the public LUNA16 dataset. It achieves competitive segmentation performance on dice similarity coefficient of 83.34% and average surface distance of 0.28 mm. More importantly, our method is proven to be more robust to the heterogeneous types of lung nodules and shows better segmentation performance on small lung nodules. Fu Zhou, Fei Luo 0004, Kafui Efio-Akolly, Ronald Bbosa, Wen Cai Huang, Jia Ni Zou, Yi-Ping Phoebe Chen |
BIBM | 7 |
| 2021 | Multi-Directional Convolution Networks with Spatial-Temporal Feature Pyramid Module for Action RecognitionabstractRecent attempts show that factorizing 3D convolutional filters into separate spatial and temporal components brings impressive improvement in action recognition. However, traditional temporal convolution operating along the temporal dimension will aggregate unrelated features, since the feature maps of fast-moving objects have shifted spatial positions. In this paper, we propose a novel and effective Multi-Directional Convolution (MDConv), which extracts features along different spatial-temporal orientations. Especially, MDConv has the same FLOPs and parameters as the traditional 1D temporal convolution. Also, we propose the Spatial-Temporal Feature Pyramid Module (STFPM) to fuse spatial semantics in different scales in a light-weight way. Our extensive experiments show that the models which integrate with MDConv achieve better accuracy on several large-scale action recognition benchmarks such as Kinetics, AVA and Something-Something V1&V2 datasets. Bohong Yang, Wu Ran, Hong Lu 0001, Yi-Ping Phoebe Chen |
ICASSP | 5 |
| 2021 | Concept-based Topic Attention for a Convolutional Sequence Document Summarization ModelabstractNeural network-based document summarization often suffers from the problem of summarizing irrelevant topic content regarding the main idea. One of the main reasons leading to this problem is a lack of human common-sense knowledge which generates facts that are not decipherable. We propose a document summarization framework called Document Summarization with Concept-based Topic Triple Attention (DOSCTTA). The framework incorporates concept-based topic information into a convolutional sequence document summarization model. We propose a concept-based topic model (CTM) to generate semantic topic information using conceptual information or knowledge which is retrieved from a knowledge base. We introduce a triple attention mechanism (TAM) to not only measure the importance of each topic concept and source element to the output elements but also the importance of the topic concept to the source element. TAM presents contextual information from three aspects and then combines them using a softmax activation to acquire the final probability distribution to enable the model to produce coherent and meaningful summaries with a wide range of rich vocabulary. The experimental evaluations which are conducted over the Gigaword and CNN/Daily Mail (CNN/DM) datasets reveal that DOSCTTA surpasses the various widely recognized state-of-the-art models (WSOTA) such as Seq2Seq, PGEN, CSM and TopicCSM. DOSCTTA achieves competitive results by generating coherent and informative summaries. Shirin Akther Khanam, Fei Liu 0003, Yi-Ping Phoebe Chen |
IJCNN | 3 |
| 2021 | Distractor-Aware Tracker with a Domain-Special Optimized Benchmark for Soccer Player TrackingabstractPlayer tracking in broadcast soccer videos has received widespread attention in the field of sports video analysis, however, we note that there is not a suitable tracking algorithm specifically for soccer video, and the existing benchmarks used for soccer player tracking cover few scenarios with low difficulties. From the observation of the soccer scene that interference and occlusion are knotty problems because the distractors are extremely similar to the targets, a distractor-aware player tracking algorithm and a high-quality benchmark for soccer play tracking (BSPT) have been presented. The distractor-aware player tracking algorithm is able to perceive semantic information about distracting players in the background by similarity judgment, the semantic distractor-aware information is encoded into a context vector and is constantly updated as the objects move through a video sequence. Distractor-aware information is then appended to the tracking result of the baseline tracker to improve the intra-class discriminative power. BSPT contains a total of 120 sequences with rich annotations. Each sequence covers 8 specialized frame-level attributes from soccer scenarios and the player occlusion situations are finely divided into 4 categories for a more comprehensive comparison. In the experimental section, the performance of our algorithm and the other 14 compared trackers are evaluated on BSPT with detailed analysis. Experimental results reveal the effectiveness of the proposed distractor-aware model especially under the attribute of occlusion. The BSPT benchmark and raw experimental results are available on the project page at http://media.hust.edu.cn/BSPT.htm. Zikai Song, Zhiwen Wan, Junqing Yu, Yi-Ping Phoebe Chen |
ICMR | 6 |
| 2021 | Shot Boundary Detection Through Multi-stage Deep Convolution Neural Network
Tingting Wang 0003, Na Feng, Junqing Yu, Yunfeng He, Yangliu Hu, Yi-Ping Phoebe Chen |
MMM (1) | 6 |
| 2021 | Motivating Literature and Evaluation of the Teaching Practices Game: Preparing Teaching Assistants to Promote InclusivityabstractIn the US, there are longstanding patterns of underrepresentation in computing (see Table 1). To make CS more inclusive, we can train computer science (CS) teaching assistants (TAs) to create inclusive classrooms. The Teaching Practices Game is a scenario-based card game meant to prepare CS TAs for difficult situations they may encounter and help them promote diversity and inclusion. Game participants (N=86) were surveyed from multiple institutions. The majority of survey respondents (N=86) agreed or strongly agreed that the game taught them new strategies for responding to difficult teaching situations (83%) and for discussing diversity and inclusion (69%). Additionally, respondents reported that the game made them more confident (74%) and more likely (66%) to respond to biased statements. Contrary to our hypotheses, we found that there were small and statistically insignificant differences in enjoyment, learning, or likelihood of response to biased statements between participants who do and do not identify as underrepresented in CS. A primary contribution of the research is a review of professional development practices and diversity training strategies that can inform the training of TAs, and we discuss the extent to which best practices from previous research were incorporated in the game. Audra Lane, Ruth Mekonnen, Catherine Jang, Yi-Ping Phoebe Chen, Colleen M. Lewis |
SIGCSE | 4 |
| 2021 | Tracking leukocytes in intravital time lapse images using 3D cell association learning network
Marzieh R. Moghadam, Yi-Ping Phoebe Chen |
Artif. Intell. Medicine | 2 |
| 2021 | Prediction of drug adverse events using deep learning in pharmaceutical discoveryabstractTraditional machine learning methods used to detect the side effects of drugs pose significant challenges as feature engineering processes are labor-intensive, expert-dependent, time-consuming and cost-ineffective. Moreover, these methods only focus on detecting the association between drugs and their side effects or classifying drug-drug interaction. Motivated by technological advancements and the availability of big data, we provide a review on the detection and classification of side effects using deep learning approaches. It is shown that the effective integration of heterogeneous, multidimensional drug data sources, together with the innovative deployment of deep learning approaches, helps reduce or prevent the occurrence of adverse drug reactions (ADRs). Deep learning approaches can also be exploited to find replacements for drugs which have side effects or help to diversify the utilization of drugs through drug repurposing. Chun Yen Lee, Yi-Ping Phoebe Chen |
Briefings Bioinform. | 2 |
| 2021 | Genetic source completeness of HIV-1 circulating recombinant forms (CRFs) predicted by multi-label learningabstractMOTIVATION: Infection with strains of different subtypes and the subsequent crossover reading between the two strands of genomic RNAs by host cells' reverse transcriptase are the main causes of the vast HIV-1 sequence diversity. Such inter-subtype genomic recombinants can become circulating recombinant forms (CRFs) after widespread transmissions in a population. Complete prediction of all the subtype sources of a CRF strain is a complicated machine learning problem. It is also difficult to understand whether a strain is an emerging new subtype and if so, how to accurately identify the new components of the genetic source. RESULTS: We introduce a multi-label learning algorithm for the complete prediction of multiple sources of a CRF sequence as well as the prediction of its chronological number. The prediction is strengthened by a voting of various multi-label learning methods to avoid biased decisions. In our steps, frequency and position features of the sequences are both extracted to capture signature patterns of pure subtypes and CRFs. The method was applied to 7185 HIV-1 sequences, comprising 5530 pure subtype sequences and 1655 CRF sequences. Results have demonstrated that the method can achieve very high accuracy (reaching 99%) in the prediction of the complete set of labels of HIV-1 recombinant forms. A few wrong predictions are actually incomplete predictions, very close to the complete set of genuine labels. AVAILABILITY AND IMPLEMENTATION: https://github.com/Runbin-tang/The-source-of-HIV-CRFs-prediction. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Runbin Tang, Yuanlin Ma, Yaoqun Wu, Yi-Ping Phoebe Chen, Limsoon Wong, Jinyan Li 0001 |
Bioinform. | 5 |
| 2021 | Dynamic user-centric access control for detection of ransomware attacks
Timothy R. McIntosh, A. S. M. Kayes, Yi-Ping Phoebe Chen, Alex Ng, Paul A. Watters |
Comput. Secur. | 3 |
| 2021 | Enforcing situation-aware access control to build malware-resilient file systems
Timothy R. McIntosh, Paul A. Watters, A. S. M. Kayes, Alex Ng, Yi-Ping Phoebe Chen |
Future Gener. Comput. Syst. | 5 |
| 2021 | Feature selection and threshold method based on fuzzy joint mutual information
Omar A. M. Salem, Yi-Ping Phoebe Chen, Xi Chen 0078 |
Int. J. Approx. Reason. | 3 |
| 2021 | Descriptive prediction of drug side-effects using a hybrid deep learning modelabstractIn this study, we developed a hybrid deep learning (DL) model, which is one of the first interpretable hybrid DL models with Inception modules, to give a descriptive prediction of drug side-effects. The model consists of a graph convolutional neural network (GCNN) with Inception modules to allow more efficient learning of drug molecular features and bidirectional long short-term memory (BiLSTM) recurrent neural networks to associate drug structure with its associated side effects. The outputs from the two networks (GCNN and BiLSTM) are then concatenated and a fully connected network is used to predict the side effects of drugs. Our model achieves an AUC score of 0.846 irrespective of what classification threshold is chosen. It has a precision score of 0.925 and the Bilingual Evaluation Understudy (BLEU) scores obtained were 0.973, 0.938, 0.927, and 0.318 which show significant achievements despite the fact that a small drug data set is used for adverse drug reaction (ADR) prediction. Moreover, the model is capable of accurately structuring correct words to describe drug side-effects and associates them with its drug name and molecular structure. The predicted drug structure and ADR relation will provide a reference for preclinical safety pharmacology studies and facilitate the identification of ADRs during early phases of drug development. It can also help detect unknown ADRs embedded in existing drugs, hence contributing significantly to the science of pharmacovigilance. Chun Yen Lee, Yi-Ping Phoebe Chen |
Int. J. Intell. Syst. | 2 |
| 2021 | Machine learning for medical imaging-based COVID-19 detection and diagnosisabstractThe novel coronavirus disease 2019 (COVID-19) is considered to be a significant health challenge worldwide because of its rapid human-to-human transmission, leading to a rise in the number of infected people and deaths. The detection of COVID-19 at the earliest stage is therefore of paramount importance for controlling the pandemic spread and reducing the mortality rate. The real-time reverse transcription-polymerase chain reaction, the primary method of diagnosis for coronavirus infection, has a relatively high false negative rate while detecting early stage disease. Meanwhile, the manifestations of COVID-19, as seen through medical imaging methods such as computed tomography (CT), radiograph (X-ray), and ultrasound imaging, show individual characteristics that differ from those of healthy cases or other types of pneumonia. Machine learning (ML) applications for COVID-19 diagnosis, detection, and the assessment of disease severity based on medical imaging have gained considerable attention. Herein, we review the recent progress of ML in COVID-19 detection with a particular focus on ML models using CT and X-ray images published in high-ranking journals, including a discussion of the predominant features of medical imaging in patients with COVID-19. Deep Learning algorithms, particularly convolutional neural networks, have been utilized widely for image segmentation and classification to identify patients with COVID-19 and many ML modules have achieved remarkable predictive results using datasets with limited sample sizes. Rokaya Rehouma, Michael Buchert, Yi-Ping Phoebe Chen |
Int. J. Intell. Syst. | 3 |
| 2021 | Joint knowledge-powered topic level attention for a convolutional text summarization model
Shirin Akther Khanam, Fei Liu 0003, Yi-Ping Phoebe Chen |
Knowl. Based Syst. | 3 |
| 2021 | Perceptual image hashing using transform domain noise resistant local binary pattern
Syed Qasim Abbas, Fawad Ahmed, Yi-Ping Phoebe Chen |
Multim. Tools Appl. | 3 |
| 2021 | FexRNA: Exploratory Data Analysis and Feature Selection of Non-Coding RNAabstractNon-coding RNA (ncRNA) is involved in many biological processes and diseases in all species. Many ncRNA datasets exist that provide ncRNA data in FASTA format which is well suited for biomedical purposes. However, for ncRNA analysis and classification, statistical learning methods require hidden numerical features from the data. Furthermore, in the literature, a wealth of sequence intrinsic features has been proposed for ncRNA identification. The extraction of hidden features, their analysis, and usage of a suitable set of features is crucial for the performance of any statistical learning method. To alleviate the posed challenges, we generated 96 feature datasets from ncRNA widely used features. The feature datasets are based on RNACentral and consist of species, ncRNA types, and expert databases that are available on the FexRNA platform. Additionally, the feature datasets are explored and analysed to provide statistical information, univariate, and bivariate analysis. We sought to determine which of these 17 features would be most appropriate to use in developing ncRNA classification approaches. For feature selection (FS), a two-phase hierarchical FS framework based on correlation and majority voting is proposed and evaluated on 5 species. The FexRNA platform provides information about ncRNA feature analysis and selection. Noorul Amin, Annette McGrath, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | ILDMSF: Inferring Associations Between Long Non-Coding RNA and Disease Based on Multi-Similarity FusionabstractThe dysregulation and mutation of long non-coding RNAs (lncRNAs) have been proved to result in a variety of human diseases. Identifying potential disease-related lncRNAs may benefit disease diagnosis, treatment and prognosis. A number of methods have been proposed to predict the potential lncRNA-disease relationships. However, most of them may give rise to incorrect results due to relying on single similarity measure. This article proposes a novel framework (ILDMSF) by fusing the lncRNA similarities and disease similarities, which are measured by lncRNA-related gene and known lncRNA-disease interaction and disease semantic interaction, and known lncRNA-disease interaction, respectively. Further, the support vector machine is employed to identify the potential lncRNA-disease associations based on the integrated similarity. The leave-one-out cross validation is performed to compare ILDMSF with other state of the art methods. The experimental results demonstrate our method is prospective in exploring potential correlations between lncRNA and disease. Qingfeng Chen, Dehuan Lai, Wei Lan 0001, Ximin Wu, Baoshan Chen, Jin Liu 0012, Yi-Ping Phoebe Chen, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2021 | Guest Editorial for the 17th Asia Pacific Bioinformatics ConferenceabstractThe eight papers in this special section were presented at the 17th Asia Pacific Bioinformatics Conference (APBC), which was held in Wuhan, China, 14-16 January 2019. Louxin Zhang, Shaoliang Peng, Yi-Ping Phoebe Chen, David Sankoff, Guoliang Li 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | Tracking Neutrophil Migration in Zebrafish Model Using Multi-Channel Feature LearningabstractTracking cells over time is crucial in the fields of computer vision and biomedical science. Studying neutrophils and their migratory profile is the highly topical fields in inflammation research due to determining role of these cells during immune responses. As neutrophils generally are of various shapes and motion, it remains challenging to track and describe their behaviours from multi-dimensional microscopy datasets. In this study, we propose a robust novel multi-channel feature learning (MCFL) model inspired by deep learning to extract the complex behaviour of neutrophils moved in time lapse images. In this model, the convolutional neural networks along with cell relocation distance and orientation channels learn the robust significant spatial and temporal features of an individual neutrophil. Additionally, we also proposed a new cell tracking framework to detect and track neutrophils in the original time-laps microscopy images, entails sampling, observation, and visualisation functions. Our proposed cell tracking-based-multi channel feature learning method has remarkable performance in rectifying common cell tracking problem compared with state-of the-art methods. Marzieh R. Moghadam, Yi-Ping Phoebe Chen |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | New Insights Into Drug Repurposing for COVID-19 Using Deep LearningabstractThe coronavirus disease 2019 (COVID-19) has continued to spread worldwide since late 2019. To expedite the process of providing treatment to those who have contracted the disease and to ensure the accessibility of effective drugs, numerous strategies have been implemented to find potential anti-COVID-19 drugs in a short span of time. Motivated by this critical global challenge, in this review, we detail approaches that have been used for drug repurposing for COVID-19 and suggest improvements to the existing deep learning (DL) approach to identify and repurpose drugs to treat this complex disease. By optimizing hyperparameter settings, deploying suitable activation functions, and designing optimization algorithms, the improved DL approach will be able to perform feature extraction from quality big data, turning the traditional DL approach, referred to as a "black box," which generalizes and learns the transmitted data, into a "glass box" that will have the interpretability of its rationale while maintaining a high level of prediction accuracy. When adopted for drug repurposing for COVID-19, this improved approach will create a new generation of DL approaches that can establish a cause and effect relationship as to why the repurposed drugs are suitable for treating COVID-19. Its ability can also be extended to repurpose drugs for other complex diseases, develop appropriate treatment strategies for new diseases, and provide precision medical treatment to patients, thus paving the way to discover new drugs that can potentially be effective for treating COVID-19. Chun Yen Lee, Yi-Ping Phoebe Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Fine-Grain Level Sports Video Search Engine
Zikai Song, Junqing Yu, Hengyou Cai, Yangliu Hu, Yi-Ping Phoebe Chen |
MMM (1) | 5 |
| 2020 | Ensuring privacy and security of genomic data and functionalitiesabstractIn recent times, the reduced cost of DNA sequencing has resulted in a plethora of genomic data that is being used to advance biomedical research and improve clinical procedures and healthcare delivery. These advances are revolutionizing areas in genome-wide association studies (GWASs), diagnostic testing, personalized medicine and drug discovery. This, however, comes with security and privacy challenges as the human genome is sensitive in nature and uniquely identifies an individual. In this article, we discuss the genome privacy problem and review relevant privacy attacks, classified into identity tracing, attribute disclosure and completion attacks, which have been used to breach the privacy of an individual. We then classify state-of-the-art genomic privacy-preserving solutions based on their application and computational domains (genomic aggregation, GWASs and statistical analysis, sequence comparison and genetic testing) that have been proposed to mitigate these attacks and compare them in terms of their underlining cryptographic primitives, security goals and complexities-computation and transmission overheads. Finally, we identify and discuss the open issues, research challenges and future directions in the field of genomic privacy. We believe this article will provide researchers with the current trends and insights on the importance and challenges of privacy and security issues in the area of genomics. Abukari Mohammed Yakubu, Yi-Ping Phoebe Chen |
Briefings Bioinform. | 2 |
| 2020 | SSET: a dataset for shot segmentation, event detection, player tracking in soccer videos
Na Feng, Zikai Song, Junqing Yu, Yi-Ping Phoebe Chen, Yizhu Zhao, Yunfeng He |
Multim. Tools Appl. | 4 |
| 2020 | Designing secure substitution boxes based on permutation of symmetric group
Amir Anees, Yi-Ping Phoebe Chen |
Neural Comput. Appl. | 2 |
| 2020 | Enhancing the Quality of Image Tagging Using a Visio-Textual Knowledge BaseabstractAuto-tagging of images is important for image understanding and for tag-based applications viz. image retrieval, visual question-answering, image captioning, etc. Although existing tagging methods incorporate both visual and textual information to assign/refine tags, they lag in tag-image relevance, completeness, and preciseness, thereby resulting in the unsatisfactory performance of tag-based applications. In order to bridge this gap, we propose a novel framework for tag assignment using knowledge embedding (TAKE) from a proposed external knowledge base, considering properties such as Rarity, Newness, Generality, and Naturalness (RNGN properties). These properties help in providing a rich semantic representation to images. Existing knowledge bases provide multiple types of relations extracted through only one modality, either text or visual, which is not effective in image related applications. We construct a simple yet effective Visio-Textual Knowledge Base (VTKB) with only four relations using reliable resources such as Wikipedia, thesauruses, dictionaries, etc. Our large scale experiments demonstrate that the proposed combination of TAKE and VTKB assigns a large number of high quality tags in comparison to the ConceptNet and ImageNet knowledge bases when used in conjunction with TAKE. Also, the effectiveness of knowledge embedding through VTKB is evaluated for image tagging and tag-based image retrieval (TBIR). Chandramani Chaudhary, Poonam Goyal, Dhanashree Nellayi Prasad, Yi-Ping Phoebe Chen |
IEEE Trans. Multim. | 4 |
| 2020 | Image Retrieval for Complex Queries Using Knowledge EmbeddingabstractWith the increase in popularity of image-based applications, users are retrieving images using more sophisticated and complex queries. We present three types of complex queries, namely, long, ambiguous, and abstract. Each type of query has its own characteristics/complexities and thus leads to imprecise and incomplete image retrieval. Existing methods for image retrieval are unable to deal with the high complexity of such queries. Search engines need to integrate their image retrieval process with knowledge to obtain rich semantics for effective retrieval. We propose a framework, Image Retrieval using Knowledge Embedding (ImReKE), for embedding knowledge with images and queries, allowing retrieval approaches to understand the context of queries and images in a better way. ImReKE (IR_Approach, Knowledge_Base) takes two inputs, namely, an image retrieval approach and a knowledge base. It selects quality concepts (concepts that possess properties such as rarity, newness , etc.) from the knowledge base to provide rich semantic representations for queries and images to be leveraged by the image retrieval approach. For the first time, an effective knowledge base that exploits both the visual and textual information of concepts has been developed. Our extensive experiments demonstrate that the proposed framework improves image retrieval significantly for all types of complex queries. The improvement is remarkable in the case of abstract queries, which have not yet been dealt with explicitly in the existing literature. We also compare the quality of our knowledge base with the existing text-based knowledge bases, such as ConceptNet, ImageNet, and the like. Chandramani Chaudhary, Poonam Goyal, Navneet Goyal, Yi-Ping Phoebe Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | Structural Network Embedding using Multi-modal Deep Auto-encoders for Predicting Drug-drug InteractionsabstractPredicting drug-drug interactions (DDIs) is crucial for patient safety and public health. The existing DDI prediction methods mainly fall into three categories: knowledge-based, similarity-based and network-based. Most recently, studies have demonstrated that integrating heterogeneous drug features is significantly important for developing high-accuracy prediction models, but it also brings many new challenges, i.e. heterogeneous properties, non-linear relations and incomplete data. In this paper, we propose a multi-modal deep auto-encoders based drug representation learning method for the DDI prediction, abbreviated as DDI-MDAE. The proposed method learns unified representations of drugs simultaneously from multiple drug feature networks using multi-modal deep auto-encoders. Then we adopt several operators on the learned drug embeddings to represent drug-drug pairs, and utilize the random forest to train models for the DDI prediction. Experimental results show that DDI-MDAE effectively learns the representations of drugs by fusing diverse information, and outperforms the other state-of-the-art benchmark methods. More importantly, DDI-MDAE works even for drugs without any known interaction. Shichao Liu 0002, Yi-Ping Phoebe Chen, Wen Zhang 0008 |
BIBM | 4 |
| 2019 | Detection of Cell Types from Single-cell RNA-seq Data using Similarity via Kernel Preserving Learning EmbeddingabstractThe recent advances in single-cell sequencing techniques allow us to study biological issues on cell levels. Detecting cell types from scRNA-seq data analysis is important and meaningful. However, high-level noise and the nonlinearity and sparsity of scRNA-seq data are great challenges. In this paper, we propose a cell-type detection algorithm preserving the overall cell relations named POCR to analyze scRNA-seq data. POCR utilizes a kernel embedding similarity measure to calculate cell-to-cell similarity, by minimizing the reconstruction error of a kernel matrix, rather than the reconstruction error of the original data adopted by other similarity metrics. According to the scale of scRNA-seq datasets, we select Gaussian kernel or linear kernel to calculate the embedding. We then adopt spectral clustering to detect the cell types based on the learned cell-to-cell similarity. The results are further visualized to demonstrate the effectiveness of the cell-type detection algorithm POCR. Further analysis shows that the learned similarity could improve the clustering and visualization of cell types in scRNA-seq data. Our proposed algorithm is compared with five other state-of-the-art cell subtype detection methods. The effectiveness of the algorithms is evaluated by two criteria: ARI and NMI. The experiments show that POCR achieves accurate and robust performance across different scRNA-seq data. Our python implementation of POCR is available at https://github.com/ZeMing-Liu/POCR. Zeming Liu, Chengzhi Hong, Yi-Ping Phoebe Chen, Shichao Liu 0002, Wen Zhang 0008 |
BIBM | 5 |
| 2019 | TC-GAN: Triangle Cycle-Consistent GANs for Face Frontalization with Facial Features PreservedabstractFace frontalization has always been an important field. Recently, with the introduction of generative adversarial networks (GANs), face frontalization has achieved remarkable success. A critical challenge during face frontalization is to ensure the features of the original profile image are retained. Even though some state-of-the-art methods can preserve identity features while rotating the face to the frontal view, they still have difficulty preserving facial expression features. Therefore, we propose the novel triangle cycle-consistent generative adversarial networks for the face frontalization task, termed TC-GAN. Our networks contain two generators and one discriminator. One of the generators generates the frontal contour, and the other generates the facial features. They work together to generate a photo-realistic frontal view of the face. We also introduce cycle-consistent loss to retain feature information effectively. To validate the advantages of TC-GAN, we apply it to the face frontalization task on two datasets. The experimental results demonstrate that our method can perform large-pose face frontalization while preserving the facial features (both identity and expression). To the best of our knowledge, TC-GAN outperforms the state-of-the-art methods in the preservation of facial identity and expression features during face frontalization. Juntong Cheng, Yi-Ping Phoebe Chen, Minjun Li, Yu-Gang Jiang 0001 |
ACM Multimedia | 2 |
| 2019 | Representation learning with extreme learning machines and empirical mode decomposition for wind speed forecasting methods
HaoFan Yang, Yi-Ping Phoebe Chen |
Artif. Intell. | 2 |
| 2019 | Hybrid deep learning and empirical mode decomposition model for time series applications
HaoFan Yang, Yi-Ping Phoebe Chen |
Expert Syst. Appl. | 2 |
| 2019 | A novel multimodal clustering framework for images with diverse associated text
Chandramani Chaudhary, Poonam Goyal, Siddhant Tuli, Shuchita Banthia, Navneet Goyal, Yi-Ping Phoebe Chen |
Multim. Tools Appl. | 6 |
| 2019 | Guest Editorial for the 16th Asia Pacific Bioinformatics ConferenceabstractThe eight papers in this special section were presented at the 16th Asia Pacific Bioinformatics Conference (APBC2018), which was held in Yokohama, Japan, 15-17 January 2018. The aim of this conference is to provide an international forum for researchers, professionals, and industrial practitioners to share their knowledge and ideas of how to surf the tidal wave of information in the area of bioinformatics and computational biology. Yoshihiro Yamanishi, Yasubumi Sakakibara, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2019 | EditorialabstractThis special section consists of eight papers selected from the accepted papers of the 27th International Conference on Genome Informatics (GIW2016), which was held in Shanghai, China, October 3-5, 2016. These papers cover diverse topics, including gene clustering, protein-protein interaction network inference, essential proteins identification, glycan structure identification, lncRNA function prediction, lncRNA-disease association prediction, signal transduction network construction, and parallel algorithms. Shuigeng Zhou, Yi-Ping Phoebe Chen, Hiroshi Mamitsuka |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | Optimized Configuration of Exponential Smoothing and Extreme Learning Machine for Traffic Flow ForecastingabstractTraffic flow forecasting is a useful technology applied to solve traffic congestion problems and to improve transportation mobility. Neural networks related approaches have been applied to develop traffic forecasting models for more than two decades. Since neural networks are sensitivity in parameters selection, selecting appropriate modeling configuration is essential to improve the accuracy and efficiency of traffic flow prediction. However, this is usually conducted by the trial-and-error method, which is very time consuming while involving too many design factors. Therefore, this paper utilizes a robust and systematic optimization approach, the Taguchi method, for obtaining the optimized configuration of the proposed exponential smoothing and extreme learning machine forecasting model. The developed model is applied to real-world data collected from freeways and highways in the United Kingdom and is compared with three existing forecasting models. The results indicate that the Taguchi method is efficient and capable for the forecasting model design and the proposed model with the optimized configuration has superior performance in traffic flow forecasting with approximate 91% and 88% accuracy rate in freeway and highway in both peak and nonpeak traffic periods. HaoFan Yang, Tharam S. Dillon, Elizabeth Chang 0001, Yi-Ping Phoebe Chen |
IEEE Trans. Ind. Informatics | 4 |
| 2018 | Linguistic Patterns and Cross Modality-based Image Retrieval for Complex QueriesabstractWith the rising prevalence of social media, coupled with the ease of sharing images, people with specific needs and applications such as known item search, multimedia question answering, etc., have started searching for visual content, which is expressed in terms of complex queries. A complex query consists of multiple concepts and their attributes are arranged to convey semantics. It is less effective to answer such queries by simply appending the search results gathered from individual or subsets of concepts present in the query. In this paper, we propose to exploit the query constituents and relationships among them. The proposed approach finds image-query relevance by integrating three models - the linguistic pattern-based textual model, the visual model, and the cross modality model. We extract linguistic patterns from complex queries, gather their related crawled images, and assign relevance scores to images in the corpus. The relevance scores are then used to rank the images. We experiment on more than 140k images and compare the [email protected] scores with the state-of-the-art image ranking methods for complex queries. Also, ranking of images obtained by our approach outperforms than that of obtained by a popular search engine. Chandramani Chaudhary, Poonam Goyal, Joel Ruben Antony Moniz, Navneet Goyal, Yi-Ping Phoebe Chen |
ICMR | 5 |
| 2018 | GCOTraj: A storage approach for historical trajectory data sets using grid cells ordering
Shengxun Yang, Zhen He 0002, Yi-Ping Phoebe Chen |
Inf. Sci. | 3 |
| 2018 | Discriminative binary feature learning and quantization in biometric key generation
Amir Anees, Yi-Ping Phoebe Chen |
Pattern Recognit. | 2 |
| 2018 | Effectively Identifying Compound-Protein Interactions by Learning from Positive and Unlabeled ExamplesabstractPrediction of compound-protein interactions (CPIs) is to find new compound-protein pairs where a protein is targeted by at least a compound, which is a crucial step in new drug design. Currently, a number of machine learning based methods have been developed to predict new CPIs in the literature. However, as there is not yet any publicly available set of validated negative CPIs, most existing machine learning based approaches use the unknown interactions (not validated CPIs) selected randomly as the negative examples to train classifiers for predicting new CPIs. Obviously, this is not quite reasonable and unavoidably impacts the CPI prediction performance. In this paper, we simply take the unknown CPIs as unlabeled examples, and propose a new method called PUCPI (the abbreviation of PU learning for Compound-Protein Interaction identification) that employs biased-SVM (Support Vector Machine) to predict CPIs using only positive and unlabeled examples. PU learning is a class of learning methods that leans from positive and unlabeled (PU) samples. To the best of our knowledge, this is the first work that identifies CPIs using only positive and unlabeled examples. We first collect known CPIs as positive examples and then randomly select compound-protein pairs not in the positive set as unlabeled examples. For each CPI/compound-protein pair, we extract protein domains as protein features and compound substructures as chemical features, then take the tensor product of the corresponding compound features and protein features as the feature vector of the CPI/compound-protein pair. After that, biased-SVM is employed to train classifiers on different datasets of CPIs and compound-protein pairs. Experiments over various datasets show that our method outperforms six typical classifiers, including random forest, L1- and L2-regularized logistic regression, naive Bayes, SVM and k-nearest neighbor (kNN), and three types of existing CPI prediction models. More information can be found at http://admis.fudan.edu.cn/projects/pucpi.html. Zhanzhan Cheng, Shuigeng Zhou, Yang Wang 0100, Hui Liu 0026, Jihong Guan, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2018 | Guest Editorial for the 14th Asia Pacific Bioinformatics ConferenceabstractThe eight papers in this special section were presented at the 14th Asia Pacific Bioinformatics Conference (APBC2016), which was held in San Francisco, USA, 11-13 January 2016. Jijun Tang, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | Guest Editorial for the 15th Asia Pacific Bioinformatics ConferenceabstractThe eight papers in this special section were presented at the 15th Asia Pacific Bioinformatics Conference (APBC2017), which was held in Shenzhen, China, 17-19 January 2017. Lusheng Wang 0001, Shuaicheng Li 0001, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | Exploiting visual and textual neighborhood information to improve image-tag relevanceabstractMany applications, such as image searching, image indexing, and image label recommendations, have started using tagged images to benefit from user input. However, tags tend to be imprecise, incomplete, and ambiguous. Moreover, tags are also biased towards the user's perspective which degrades the performance of tag-based systems. Most of the existing methods use visual neighborhoods and/or tags to estimate image-tag relevance. We improve image-tag relevance by combining visual neighborhood of images and textual neighborhood of tags. By doing this, we boost the ranking of informative tags of an image. Most of the image-tag relevance measures work well when large supporting data is available, which is typically not sufficient in real datasets. This problem of Void of Information (VoI) is addressed by exploiting tags of visual neighbors of the images. We also exploit external resources like Wikipedia and WordNet to strengthen the tags. The proposed approach, TVNTag (Textual Visual Neighborhood based Tag) exhibits up to 46.1% relative improvement in tag ranking and 79.5% in image ranking, with respect to the current state-of-the-art methods. The experiments are conducted for different tasks and evaluation scenarios on benchmarked social data, such as MIRFlickr, NUS-WIDE, and train10k. Chandramani Chaudhary, Poonam Goyal, Yi-Ping Phoebe Chen |
IEEE BigData | 3 |
| 2017 | A comprehensive study of RNA secondary structure alignment algorithmsabstractRNA secondary structure alignment has received more attention since the discovery of the structure-function relationships in some non-protein-encoding RNAs. However, unlike the pure sequence alignment problem, which has been solved in polynomial time, secondary structure alignment incorporates the base pairings as another information dimension in addition to the base sequence. This problem therefore becomes more challenging. In this study, we classify the selected approaches, and algorithmically illustrate how these methods address the alignment problems with different structure types. Other features such as the types of base pair edit operations supported and the time complexity are also compared. Jimmy Ka Ho Chiu, Yi-Ping Phoebe Chen |
Briefings Bioinform. | 2 |
| 2017 | Identification of protein complexes by integrating multiple alignment of protein interaction networksabstractMOTIVATION: Protein complexes are one of the keys to studying the behavior of a cell system. Many biological functions are carried out by protein complexes. During the past decade, the main strategy used to identify protein complexes from high-throughput network data has been to extract near-cliques or highly dense subgraphs from a single protein-protein interaction (PPI) network. Although experimental PPI data have increased significantly over recent years, most PPI networks still have many false positive interactions and false negative edge loss due to the limitations of high-throughput experiments. In particular, the false negative errors restrict the search space of such conventional protein complex identification approaches. Thus, it has become one of the most challenging tasks in systems biology to automatically identify protein complexes. RESULTS: In this study, we propose a new algorithm, NEOComplex ( NE CC- and O rtholog-based Complex identification by multiple network alignment), which integrates functional orthology information that can be obtained from different types of multiple network alignment (MNA) approaches to expand the search space of protein complex detection. As part of our approach, we also define a new edge clustering coefficient (NECC) to assign weights to interaction edges in PPI networks so that protein complexes can be identified more accurately. The NECC is based on the intuition that there is functional information captured in the common neighbors of the common neighbors as well. Our results show that our algorithm outperforms well-known protein complex identification tools in a balance between precision and recall on three eukaryotic species: human, yeast, and fly. As a result of MNAs of the species, the proposed approach can tolerate edge loss in PPI networks and even discover sparse protein complexes which have traditionally been a challenge to predict. AVAILABILITY AND IMPLEMENTATION: http://acolab.ie.nthu.edu.tw/bionetwork/NEOComplex. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cheng-Yu Ma, Yi-Ping Phoebe Chen, Bonnie Berger, Chung-Shou Liao |
Bioinform. | 2 |
| 2017 | Optimized Structure of the Traffic Flow Forecasting Model With a Deep Learning ApproachabstractForecasting accuracy is an important issue for successful intelligent traffic management, especially in the domain of traffic efficiency and congestion reduction. The dawning of the big data era brings opportunities to greatly improve prediction accuracy. In this paper, we propose a novel model, stacked autoencoder Levenberg-Marquardt model, which is a type of deep architecture of neural network approach aiming to improve forecasting accuracy. The proposed model is designed using the Taguchi method to develop an optimized structure and to learn traffic flow features through layer-by-layer feature granulation with a greedy layerwise unsupervised learning algorithm. It is applied to real-world data collected from the M6 freeway in the U.K. and is compared with three existing traffic predictors. To the best of our knowledge, this is the first time that an optimized structure of the traffic flow forecasting model with a deep learning approach is presented. The evaluation results demonstrate that the proposed model with an optimized structure has superior performance in traffic flow forecasting. HaoFan Yang, Tharam S. Dillon, Yi-Ping Phoebe Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Systematic review of virtual speech therapists for speech disorders
Yi-Ping Phoebe Chen, Caddi Johnson, Pooia Lalbakhsh, Terry Caelli, Guang Deng, David B. H. Tay, Shane Erickson, Philip Broadbridge, Amr El Refaie, Wendy Doubé, Meg E. Morris |
Comput. Speech Lang. | 1 |
| 2016 | PaMeCo join: A parallel main memory compact hash join
Steven Keith Begley, Zhen He 0002, Yi-Ping Phoebe Chen |
Inf. Syst. | 3 |
| 2016 | Workload-Based Ordering of Multi-Dimensional DataabstractTransforming multi-dimensional data into a one-dimensional sequence using space-filling curves such as the Hilbert curve, the Gray curve, and the Z-curve has been studied extensively. These techniques are not sensitive to data or workload skewness, however, in practice, user-access patterns and data distributions are often very skewed in high dimensional space. It is desirable to produce a one-dimensional sequence which keeps the multi-dimensional grid cells that are queried together close to each other. This generates sequences with higher spatial locality. We propose a workload-based approach to produce one-dimensional ordering from multi-dimensional data in this paper. An extensive experimental evaluation suggests that our approach produces a high quality ordering sequence which outperforms the existing state-of-the-art Hilbert curve by a factor of 4.84, the Gray curve by a factor of 6.66, and the Z-curve by a factor of 7.26 for the number of subsequences used to answer a query; and for IO time, it outperforms the Hilbert curve by a factor of 2.20, the Gray curve by a factor of 2.25, and the Z-curve by 2.38. Shengxun Yang, Zhen He 0002, Yi-Ping Phoebe Chen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2015 | Pairwise RNA secondary structure alignment with conserved stem patternabstractMOTIVATION: The regulatory functions performed by non-coding RNAs are related to their 3D structures, which are, in turn, determined by their secondary structures. Pairwise secondary structure alignment gives insight into the functional similarity between a pair of RNA sequences. Numerous exact or heuristic approaches have been proposed for computational alignment. However, the alignment becomes intractable when arbitrary pseudoknots are allowed. Also, since non-coding RNAs are, in general, more conserved in structures than sequences, it is more effective to perform alignment based on the common structural motifs discovered. RESULTS: We devised a method to approximate the true conserved stem pattern for a secondary structure pair, and constructed the alignment from it. Experimental results suggest that our method identified similar RNA secondary structures better than the existing tools, especially for large structures. It also successfully indicated the conservation of some pseudoknot features with biological significance. More importantly, even for large structures with arbitrary pseudoknots, the alignment can usually be obtained efficiently. AVAILABILITY AND IMPLEMENTATION: Our algorithm has been implemented in a tool called PSMAlign. The source code of PSMAlign is freely available at http://homepage.cs.latrobe.edu.au/ypchen/psmalign/. Jimmy Ka Ho Chiu, Yi-Ping Phoebe Chen |
Bioinform. | 2 |
| 2015 | Image based computer aided diagnosis system for cancer detection
Howard Lee, Yi-Ping Phoebe Chen |
Expert Syst. Appl. | 2 |
| 2015 | Data mining in lung cancer pathologic staging diagnosis: Correlation between clinical and pathology information
HaoFan Yang, Yi-Ping Phoebe Chen |
Expert Syst. Appl. | 2 |
| 2015 | Guest Editorial for the 13th Asia Pacific Bioinformatics ConferenceabstractPresents papers that were presented at the 13th Asia Pacific Bioinformatics Conference. Hsien-Da Huang, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2014 | Skin cancer extraction with optimum fuzzy thresholding technique
Howard Lee, Yi-Ping Phoebe Chen |
Appl. Intell. | 2 |
| 2014 | Cell cycle phase detection with cell deformation analysis
Howard Lee, Yi-Ping Phoebe Chen |
Expert Syst. Appl. | 2 |
| 2014 | Using Dead Ants to improve the robustness and adaptability of AntNet routing algorithm
Pooia Lalbakhsh, Bahram Zaeri, Yi-Ping Phoebe Chen |
J. Netw. Comput. Appl. | 3 |
| 2014 | Model-based approach to spatial-temporal sampling of video clips for video object detection by classification
Chi-Han Chuang, Shyi-Chyi Cheng, Chin-Chun Chang, Yi-Ping Phoebe Chen |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | Cell morphology based classification for red cells in blood smear images
Howard Lee, Yi-Ping Phoebe Chen |
Pattern Recognit. Lett. | 2 |
| 2014 | Guest editorial for the 12th Asia Pacific Bioinformatics ConferenceabstractThe eight papers in this special section were presented at the Twelfth Asia Pacific Bioinformatics Conference (APBC2014), which was held in Shanghai, China, 17-19 January, 2014. The papers cover diverse topics, including protein complex prediction, nucleosome positioning, biological data classification and clustering, sequence comparison, genome reconstruction, and enhancers etc. Shuigeng Zhou, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2013 | SeTPR*-tree: Efficient Buffering for Spatiotemporal Indexes Via Shared ExecutionabstractIn this paper, we study the problem of efficient spatiotemporal indexing of moving objects. In order to reduce the frequency of object location updates, a linear motion model is used to model the near future location of moving objects. A number of existing spatiotemporal indexes have already been proposed for indexing these models. However, these indexes are either designed to offer high query performance or high update performance. Therefore, they are all ill suited to handle situations where both queries and updates arrive at a high rate. In this paper, we propose the SeTPR*-tree which extends the TPR*-tree to more efficiently use a limited-sized RAM buffer for processing queries in batches and rapidly arriving updates. We provide both theoretical and empirical evidence of the effectiveness of the SeTPR*-tree in improving query and update performance. We have conducted extensive experiments using a recognized spatiotemporal benchmark on a solid state drive. The SeTPR*-tree simultaneously outperforms the best tested index optimized for queries by up to a factor of 5.6 for query I/O and outperforms the best tested index for updates by up to a factor of 11.5 for update I/O. Thi Nguyen, Zhen He 0002, Yi-Ping Phoebe Chen |
Comput. J. | 3 |
| 2013 | Computational intelligence for heart disease diagnosis: A medical knowledge driven approach
Jesmin Nahar, Tasadduq Imam, Kevin Tickle, Yi-Ping Phoebe Chen |
Expert Syst. Appl. | 4 |
| 2013 | Association rule mining to detect factors which contribute to heart disease in males and females
Jesmin Nahar, Tasadduq Imam, Kevin Tickle, Yi-Ping Phoebe Chen |
Expert Syst. Appl. | 4 |
| 2013 | GHT-based associative memory learning and its application to Human action detection and classification
Shyi-Chyi Cheng, Kwang-Yu Cheng, Yi-Ping Phoebe Chen |
Pattern Recognit. | 3 |
| 2013 | Guest Editorial: Advanced Algorithms of BioinformaticsabstractAdvanced algorithms can help to identify functional relationships of genes in a biological process, discovery sequence, and structural similarities and provide insights into the regulatory mechanism in bioinformatics. In this special section, three papers in their significantly extended versions were selected from the papers presented at the Tenth Asia Pacific Bioinformatics Conference (APBC2012). These papers have shown great cooperation and conscientious consideration throughout the challenging bioinformatics experiments. Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2012 | A Scheduling Method for Node Relay-based Webcast Considering ReconnectionabstractDue to the recent popularization of digital broadcasting systems, selective contents broadcasting depending on users' preferences with node relay-based web cast have attracted much attention. In node relay-based web cast, waiting time is reduced by receiving contents from several nodes. However, when a node that delivers contents disconnects from the network or reconnects to it while delivering contents, waiting time increases by changing a delivery schedule. In this paper, we propose a scheduling method considering reconnection on selective contents delivery with node relay-based web cast that relay data among nodes. Our proposed method reduces waiting time by restructuring the delivery schedule considering the disconnection and the reconnection of the node. Yusuke Gotoh, Tomoki Yoshihisa, Hideo Taniguchi, Masanori Kanazawa, Wenny Rahayu, Yi-Ping Phoebe Chen |
AINA | 6 |
| 2012 | MCJoin: a memory-constrained join for column-store main-memory databasesabstractThere exists a need for high performance, read-only main-memory database systems for OLAP-style application scenarios. Most of the existing works in this area are centered around the domain of column-store databases, which are particularly well suited to OLAP-style scenarios and have been shown to overcome the memory bottleneck issues that have been found to hinder the more traditional row-store database systems. One of the main database operations these systems are focused on optimizing is the JOIN operation. However, all these existing systems use join algorithms that are designed with the unrealistic assumption that there is unlimited temporary memory available to perform the join. In contrast, we propose a Memory Constrained Join algorithm (MCJoin) which is both high performing and also performs all of its operations within a tight given memory constraint. Extensive experimental results show that MCJoin outperforms a naive memory constrained version of the state-of-the-art Radix-Clustered Hash Join algorithm in all of the situations tested, with margins of up to almost 500%. Steven Keith Begley, Zhen He 0002, Yi-Ping Phoebe Chen |
SIGMOD Conference | 3 |
| 2012 | Computational intelligence for microarray data and biomedical image analysis for the early diagnosis of breast cancer
Jesmin Nahar, Tasadduq Imam, Kevin Tickle, A. B. M. Shawkat Ali, Yi-Ping Phoebe Chen |
Expert Syst. Appl. | 5 |
| 2012 | Guest Editorial: Application and Development of Bioinformatics
Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | A method to reduce waiting time for close-range broadcastingabstractDue to the recent popularization of digital webcast systems. close-range broadcasting using continuous media data, i.e. audio and video, has attracted great attention. For example, in a movie program, after a user watches interesting content such as a highlight scene, he/she will watch the main program continuously. In close-range broadcasting, the necessary bandwidth for continuously playing the two types of data increases. Conventional methods reduce the necessary bandwidth by producing an effective broadcast schedule for continuous media data. However, these methods do not consider the broadcast schedule for two types of continuous media data. When two types of continuous media data are scheduled, waiting time that occurs from finishing the highlight scene to starting the main scene may increase. In this paper, we propose a scheduling method to reduce the waiting time for close-range broadcasting. In our proposed method, by dividing two types of data and producing an effective broadcast schedule considering the available bandwidth, we can effectively reduce the waiting time. Yusuke Gotoh, Tomoki Yoshihisa, Hideo Taniguchi, Masanori Kanazawa, Wenny Rahayu, Yi-Ping Phoebe Chen |
MoMM | 6 |
| 2011 | Comparative study of computational methods to detect the correlated reaction sets in biochemical networksabstractCorrelated reaction sets (Co-Sets) are mathematically defined modules in biochemical reaction networks which facilitate the study of biological processes by decomposing complex reaction networks into conceptually simple units. According to the degree of association, Co-Sets can be classified into three types: perfect, partial and directional. Five approaches have been developed to calculate Co-Sets, including network-based pathway analysis, Monte Carlo sampling, linear optimization, enzyme subsets and hard-coupled reaction sets. However, differences in design and implementation of these methods lead to discrepancies in the resulted Co-Sets as well as in their use in biotechnology which need careful interpretation. In this paper, we provide a comparative study of the methods for Co-Sets computing in detail from four aspects: (i) sensitivity, (ii) completeness and soundness, (iii) flexibility and (iv) scalability. By applying them to Escherichia coli core metabolic network, the differences and relationships among these methods are clearly articulated which may be useful for potential users. Yanping Xi, Yi-Ping Phoebe Chen, Fei Wang 0017 |
Briefings Bioinform. | 2 |
| 2011 | Development and application of a modified dynamic time warping algorithm (DTW-S) to analyses of primate brain expression time seriesabstractBACKGROUND: Comparing biological time series data across different conditions, or different specimens, is a common but still challenging task. Algorithms aligning two time series represent a valuable tool for such comparisons. While many powerful computation tools for time series alignment have been developed, they do not provide significance estimates for time shift measurements. RESULTS: Here, we present an extended version of the original DTW algorithm that allows us to determine the significance of time shift estimates in time series alignments, the DTW-Significance (DTW-S) algorithm. The DTW-S combines important properties of the original algorithm and other published time series alignment tools: DTW-S calculates the optimal alignment for each time point of each gene, it uses interpolated time points for time shift estimation, and it does not require alignment of the time-series end points. As a new feature, we implement a simulation procedure based on parameters estimated from real time series data, on a series-by-series basis, allowing us to determine the false positive rate (FPR) and the significance of the estimated time shift values. We assess the performance of our method using simulation data and real expression time series from two published primate brain expression datasets. Our results show that this method can provide accurate and robust time shift estimates for each time point on a gene-by-gene basis. Using these estimates, we are able to uncover novel features of the biological processes underlying human brain development and maturation. CONCLUSIONS: The DTW-S provides a convenient tool for calculating accurate and robust time shift estimates at each time point for each gene, based on time series data. The estimates can be used to uncover novel biological features of the system being studied. The DTW-S is freely available as an R package TimeShift at http://www.picb.ac.cn/Comparative/data.html. Yi-Ping Phoebe Chen, Shengyu Ni, Augix Guohua Xu, Martin Vingron, Mehmet Somel, Philipp Khaitovich |
BMC Bioinform. | 2 |
| 2011 | External Sorting on Flash Memory Via Natural Page Run GenerationabstractThe increasing popularity of flash memory means more database systems will run on flash memory in the future. One of the most important database operations is the external sort. Hence, this paper is focused on studying the problem of efficient external sorting on flash memory. In contrast to most previous work, we target the situation where previously sorted data have become progressively unsorted due to data updates. Accordingly, we call this ‘partially’ sorted data. We focus on re-sorting partially sorted data by taking advantage of the partial sorted nature of the data to speed up the run generation phase of the traditional external merge sort. We do this by finding ‘naturally occurring’ page runs in the partially sorted data. Our algorithm can perform up to a factor of 1024 less write IO compared with a traditional external merge sort during the run generation phase. We map the problem of finding naturally occurring runs into the shortest distance problem in a directed acyclic graph (DAG). Accordingly, we propose an optimal solution to the problem using the well-known DAG-Shortest-Paths algorithm. However, we found that the optimal solution was too slow for even moderate-sized data sets and accordingly propose a fast heuristic solution that—we experimentally show—finds a high percentage of page runs using a minimum of computational overhead. Experiments using both real and synthetic data sets show that our heuristic algorithm can halve the external sorting time when compared with three likely competing external sorting algorithms. Zhen He 0002, Yi-Ping Phoebe Chen, Thi Nguyen |
Comput. J. | 3 |
| 2011 | Function Annotation for Pseudoknot Using Structure SimilarityabstractMany raw biological sequence data have been generated by the human genome project and related efforts. The understanding of structural information encoded by biological sequences is important to acquire knowledge of their biochemical functions but remains a fundamental challenge. Recent interest in RNA regulation has resulted in a rapid growth of deposited RNA secondary structures in varied databases. However, a functional classification and characterization of the RNA structure have only been partially addressed. This article aims to introduce a novel interval-based distance metric for structure-based RNA function assignment. The characterization of RNA structures relies on distance vectors learned from a collection of predicted structures. The distance measure considers the intersected, disjoint, and inclusion between intervals. A set of RNA pseudoknotted structures with known function are applied and the function of the query structure is determined by measuring structure similarity. This not only offers sequence distance criteria to measure the similarity of secondary structures but also aids the functional classification of RNA structures with pesudoknots. Qingfeng Chen, Yi-Ping Phoebe Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2010 | Guest Editors' IntroductionabstractGuest editors' introduction Zili Zhang 0001, Yi-Ping Phoebe Chen |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2010 | Exploring the ncRNA-ncRNA patterns based on bridging rules
Feng Chen 0014, Yi-Ping Phoebe Chen |
J. Biomed. Informatics | 2 |
| 2010 | Mining characteristic relations bind to RNA secondary structuresabstractThe identification of RNA secondary structures has been among the most exciting recent developments in biology and medical science. It has been recognized that there is an abundance of functional structures with frameshifting, regulation of translation, and splicing functions. However, the inherent signal for secondary structures is weak and generally not straightforward due to complex interleaving substrings. This makes it difficult to explore their potential functions from various structure data. Our approach, based on a collection of predicted RNA secondary structures, allows us to efficiently capture interesting characteristic relations in RNA and bring out the top-ranked rules for specified association groups. Our results not only point to a number of interesting associations and include a brief biological interpretation to them. It assists biologists in sorting out the most significant characteristic structure patterns and predicting structure-function relationships in RNA. Qingfeng Chen, Yi-Ping Phoebe Chen |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2010 | Knowledge-Discounted Event Detection in Sports VideoabstractAutomatic events annotation is an essential requirement for constructing an effective sports video summary. Researchers worldwide have actively been seeking the most robust and powerful solutions to detect and classify key events (or highlights) in different sports. Most of the current and widely used approaches have employed rules that model the typical pattern of audiovisual features within particular sport events. These rules are mainly based on manual observation and heuristic knowledge; therefore, machine learning can be used as an alternative. To bridge the gap between the two alternatives, we propose a hybrid approach, which integrates statistics into logical rule-based models during highlight detection. We have also successfully pioneered the use of play-break segment as a universal scope of detection and a standard set of features that can be applied for different sports, including soccer, basketball, and Australian football. The proposed method uses a limited amount of domain knowledge, making this method less subjective and more robust for different sports. An experiment using a large data set of sports video has demonstrated the effectiveness and robustness of the algorithms. Dian Tjondronegoro, Yi-Ping Phoebe Chen |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2009 | High Functional Coherence in k-Partite Protein Cliques of Protein Interaction NetworksabstractWe introduce a new topological concept called k-partite protein cliques to study protein interaction (PPI) networks.In particular, we examine functional coherence of proteins in k-partite protein cliques. A k-partite protein clique is a k-partite maximal clique comprising two or more nonoverlapping protein subsets between any two of which full interactions are exhibited. In the detection of PPI's k-partite maximal cliques, we propose to transform PPI networks into induced K-partite graphs with proteins as vertices where edges only exist among the graph's partites. Then, we present a k-partite maximal clique mining (MaCMik) algorithm to enumerate k-partite maximal cliques from K-partite graphs. Our MaCMik algorithm is applied to a yeast PPI network. We observe that there does exist interesting and unusually high functional coherence in k-partite proteincliques-most proteins in k-partite protein cliques, especially those in the same partites, share the same functions. Therefore, the idea of k-partite protein cliques suggests a novel approach to characterizing PPI networks, and may help function prediction for unknown proteins. Qian Liu 0014, Yi-Ping Phoebe Chen, Jinyan Li 0001 |
BIBM | 2 |
| 2009 | Spherical Harmonics and Distance Transform for Image Representation and Retrieval
Atul Sajjanhar, Guojun Lu, Dengsheng Zhang, Jingyu Hou 0001, Yi-Ping Phoebe Chen |
IDEAL | 5 |
| 2009 | Early Breast Cancer Identification: Which Way to Go? Microarray or Image Based Computer Aided Diagnosis!abstractThe goal of this research is to develop a computer aided diagnostic (CAD) system that can detect breast cancer in the early stage by using microarray and image data. We verified the performance of six well known classification algorithms with various performance matrices. Although we do not suggest a unique classifier algorithm for a CAD system, we do identify a number of algorithms whose performance is very promising. The algorithms performance was validated by 3 images dataset; two have been used for the first time in this experiment. Multidimensional image filtering is adopted for the final data extraction. The image data classification performance is compared with microarray data. Results suggest the most effective means of breast cancer identification in the early stage is a hybrid approach. Jesmin Nahar, Kevin Tickle, A. B. M. Shawkat Ali, Yi-Ping Phoebe Chen |
NSS | 4 |
| 2009 | Analysis on relationship between extreme pathways and correlated reaction setsabstractBACKGROUND: Constraint-based modeling of reconstructed genome-scale metabolic networks has been successfully applied on several microorganisms. In constraint-based modeling, in order to characterize all allowable phenotypes, network-based pathways, such as extreme pathways and elementary flux modes, are defined. However, as the scale of metabolic network rises, the number of extreme pathways and elementary flux modes increases exponentially. Uniform random sampling solves this problem to some extent to study the contents of the available phenotypes. After uniform random sampling, correlated reaction sets can be identified by the dependencies between reactions derived from sample phenotypes. In this paper, we study the relationship between extreme pathways and correlated reaction sets. RESULTS: Correlated reaction sets are identified for E. coli core, red blood cell and Saccharomyces cerevisiae metabolic networks respectively. All extreme pathways are enumerated for the former two metabolic networks. As for Saccharomyces cerevisiae metabolic network, because of the large scale, we get a set of extreme pathways by sampling the whole extreme pathway space. In most cases, an extreme pathway covers a correlated reaction set in an 'all or none' manner, which means either all reactions in a correlated reaction set or none is used by some extreme pathway. In rare cases, besides the 'all or none' manner, a correlated reaction set may be fully covered by combination of a few extreme pathways with related function, which may bring redundancy and flexibility to improve the survivability of a cell. In a word, extreme pathways show strong complementary relationship on usage of reactions in the same correlated reaction set. CONCLUSION: Both extreme pathways and correlated reaction sets are derived from the topology information of metabolic networks. The strong relationship between correlated reaction sets and extreme pathways suggests a possible mechanism: as a controllable unit, an extreme pathway is regulated by its corresponding correlated reaction sets, and a correlated reaction set is further regulated by the organism's regulatory network. Yanping Xi, Yi-Ping Phoebe Chen, Ming Cao 0005, Weirong Wang, Fei Wang 0017 |
BMC Bioinform. | 2 |
| 2009 | Acoustic feature selection for automatic emotion recognition from speech
Jia Rong, Gang Li 0009, Yi-Ping Phoebe Chen |
Inf. Process. Manag. | 3 |
| 2009 | Candidate working set strategy based SMO algorithm in support vector machine
Yi-Ping Phoebe Chen, Bin Jiang 0001 |
Inf. Process. Manag. | 3 |
| 2009 | Discovery of Structural and Functional Features in RNA PseudoknotsabstractAn RNA pseudoknot consists of nonnested double-stranded stems connected by single-stranded loops. There is increasing recognition that RNA pseudoknots are one of the most prevalent RNA structures and fulfill a diverse set of biological roles within cells, and there is an expanding rate of studies into RNA pseudoknotted structures as well as increasing allocation of function. These not only produce valuable structural data but also facilitate an understanding of structural and functional characteristics in RNA molecules. PseudoBase is a database providing structural, functional, and sequence data related to RNA pseudoknots. To capture the features of RNA pseudoknots, we present a novel framework using quantitative association rule mining to analyze the pseudoknot data. The derived rules are classified into specified association groups regarding structure, function, and category of RNA pseudoknots. The discovered association rules assist biologists in filtering out significant knowledge of structure-function and structure-category relationships. A brief biological interpretation to the relationships is presented, and their potential correlations with each other are highlighted. Qingfeng Chen, Yi-Ping Phoebe Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2008 | Interacting Amino Acid Preferences of 3D Pattern Pairs at the Binding Sites of Transient and Obligate Protein Complexes
Suryani Lukman, Kelvin Sim, Jinyan Li 0001, Yi-Ping Phoebe Chen |
APBC | 4 |
| 2008 | Medical Knowledge Discovery from a Regional Asthma Dataset
Sam Schmidt, Gang Li 0009, Yi-Ping Phoebe Chen |
ICIC (2) | 3 |
| 2008 | A scalable and extensible segment-event-object-based sports video retrieval systemabstractSport video data is growing rapidly as a result of the maturing digital technologies that support digital video capture, faster data processing, and large storage. However, (1) semi-automatic content extraction and annotation, (2) scalable indexing model, and (3) effective retrieval and browsing, still pose the most challenging problems for maximizing the usage of large video databases. This article will present the findings from a comprehensive work that proposes a scalable and extensible sports video retrieval system with two major contributions in the area of sports video indexing and retrieval. The first contribution is a new sports video indexing model that utilizes semi-schema-based indexing scheme on top of an Object-Relationship approach. This indexing model is scalable and extensible as it enables gradual index construction which is supported by ongoing development of future content extraction algorithms. The second contribution is a set of novel queries which are based on XQuery to generate dynamic and user-oriented summaries and event structures. The proposed sports video retrieval system has been fully implemented and populated with soccer, tennis, swimming, and diving video. The system has been evaluated against 20 users to demonstrate and confirm its feasibility and benefits. The experimental sports genres were specifically selected to represent the four main categories of sports domain: period-, set-point-, time (race)-, and performance-based sports. Thus, the proposed system should be generic and robust for all types of sports. Dian Tjondronegoro, Yi-Ping Phoebe Chen, Adrien Joly |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2007 | Learning Dependency Model for AMP-Activated Protein Kinase Regulation
Yi-Ping Phoebe Chen, Qiumei Qin, Qingfeng Chen |
KSEM | 1 |
| 2007 | Identifying Dependency Between Secure Messages for Protocol Analysis
Qingfeng Chen, Shichao Zhang 0001, Yi-Ping Phoebe Chen |
KSEM | 3 |
| 2007 | Detecting inconsistency in biological molecular databases using ontologies
Qingfeng Chen, Yi-Ping Phoebe Chen, Chengqi Zhang |
Data Min. Knowl. Discov. | 2 |
| 2006 | Analyzing Inconsistency Toward Enhancing Integration of Biological Molecular Databases
Yi-Ping Phoebe Chen, Qingfeng Chen |
APBC | 1 |
| 2006 | Preface
Ueng-Cheng Yang, Yi-Ping Phoebe Chen, Limsoon Wong |
APBC | 3 |
| 2006 | Detecting Collusion Attacks in Security Protocols
Qingfeng Chen, Yi-Ping Phoebe Chen, Shichao Zhang 0001, Chengqi Zhang |
APWeb | 2 |
| 2006 | Using Decision-Tree to Automatically Construct Learned-Heuristics for Events Classification in Sports VideoabstractAutomatic events classification is an essential requirement for constructing an effective sports video summary. It has become a well-known theory that the high-level semantics in sport video can be "computationally interpreted" based on the occurrences of specific audio and visual features which can be extracted automatically. State-of-the-art solutions for features-based event classification have only relied on either manual-knowledge based heuristics or machine learning. To bridge the gaps, we have successfully combined the two approaches by using learning-based heuristics. The heuristics are constructed automatically using decision tree while manual supervision is only required to check the features and highlight contained in each training segment. Thus, fully automated construction of classification system for sports video events has been achieved. A comprehensive experiment on 10 hours video dataset, with five full-match soccer and five full-match basketball videos, has demonstrated the effectiveness/robustness of our algorithms Dian Tjondronegoro, Yi-Ping Phoebe Chen |
ICME | 2 |
| 2006 | A Similarity Search Algorithm to Predict Protein Structures
Jiyuan An, Yi-Ping Phoebe Chen |
KES (2) | 2 |
| 2006 | Towards universal and statistical-driven heuristics for automatic classification of sports video eventsabstractResearchers worldwide have been actively seeking for the most robust and powerful solutions to detect and classify key events (or highlights) in various sports domains. Most approaches have employed manual heuristics that model the typical pattern of audio-visual features within particular sport events. To avoid manual observation and knowledge, machine-learning can be used as an alternative approach. To bridge the gaps between these two alternatives, an attempt is made to integrate statistics into heuristic models during highlight detection in our investigation. The models can be designed with a modest amount of domain-knowledge, making them less subjective and more robust for different sports. We have also successfully used a universal scope of detection and a standard set of features that can be applied for different sports that include soccer, basketball and Australian football. An experiment on a large dataset of sport videos, with a total of around 15 hours, has demonstrated the effectiveness and robustness of our algorithms Dian Tjondronegoro, Yi-Ping Phoebe Chen |
MMM | 2 |
| 2006 | Mining Frequent Itemsets for Protein Kinase Regulation
Qingfeng Chen, Yi-Ping Phoebe Chen, Chengqi Zhang, Lianggang Li |
PRICAI | 2 |
| 2006 | Finding Short Patterns to Classify Text DocumentsabstractMany classification methods have been proposed to find patterns in text documents. However, according to Occam's razor principle, "the explanation of any phenomenon should make as few assumptions as possible", short patterns usually have more explainable and meaningful for classifying text documents. In this paper, we propose a depth-first pattern generation algorithm, which can find out short patterns from text document more effectively, comparing with breadth-first algorithm Jiyuan An, Yi-Ping Phoebe Chen |
Web Intelligence | 2 |
| 2006 | Mining frequent patterns for AMP-activated protein kinase regulation on skeletal muscleabstractBACKGROUND: AMP-activated protein kinase (AMPK) has emerged as a significant signaling intermediary that regulates metabolisms in response to energy demand and supply. An investigation into the degree of activation and deactivation of AMPK subunits under exercise can provide valuable data for understanding AMPK. In particular, the effect of AMPK on muscle cellular energy status makes this protein a promising pharmacological target for disease treatment. As more AMPK regulation data are accumulated, data mining techniques can play an important role in identifying frequent patterns in the data. Association rule mining, which is commonly used in market basket analysis, can be applied to AMPK regulation. RESULTS: This paper proposes a framework that can identify the potential correlation, either between the state of isoforms of alpha, beta and gamma subunits of AMPK, or between stimulus factors and the state of isoforms. Our approach is to apply item constraints in the closed interpretation to the itemset generation so that a threshold is specified in terms of the amount of results, rather than a fixed threshold value for all itemsets of all sizes. The derived rules from experiments are roughly analyzed. It is found that most of the extracted association rules have biological meaning and some of them were previously unknown. They indicate direction for further research. CONCLUSION: Our findings indicate that AMPK has a great impact on most metabolic actions that are related to energy demand and supply. Those actions are adjusted via its subunit isoforms under specific physical training. Thus, there are strong co-relationships between AMPK subunit isoforms and exercises. Furthermore, the subunit isoforms are correlated with each other in some cases. The methods developed here could be used when predicting these essential relationships and enable an understanding of the functions and metabolic pathways regarding AMPK. Qingfeng Chen, Yi-Ping Phoebe Chen |
BMC Bioinform. | 2 |
| 2006 | An evolutionary learning approach for adaptive negotiation agentsabstractDeveloping effective and efficient negotiation mechanisms for real-world applications such as e-business is challenging because negotiations in such a context are characterized by combinatorially complex negotiation spaces, tough deadlines, very limited information about the opponents, and volatile negotiator preferences. Accordingly, practical negotiation systems should be empowered by effective learning mechanisms to acquire dynamic domain knowledge from the possibly changing negotiation contexts. This article illustrates our adaptive negotiation agents, which are underpinned by robust evolutionary learning mechanisms to deal with complex and dynamic negotiation contexts. Our experimental results show that GA-based adaptive negotiation agents outperform a theoretically optimal negotiation mechanism that guarantees Pareto optimal. Our research work opens the door to the development of practical negotiation systems for real-world applications. © 2006 Wiley Periodicals, Inc. Int J Int Syst 21: 41–72, 2006. Raymond Y. K. Lau, Maolin Tang, On Wong, Stephen Milliner, Yi-Ping Phoebe Chen |
Int. J. Intell. Syst. | 5 |
| 2006 | MDSM: Microarray database schema matching using the Hungarian method
Yi-Ping Phoebe Chen, Supawan Prompramote, Frédéric Maire |
Inf. Sci. | 1 |
| 2006 | Using object and trajectory analysis to facilitate indexing and retrieval of video
Carlos Lopez, Yi-Ping Phoebe Chen |
Knowl. Based Syst. | 2 |
| 2006 | Guest editors' introduction: multimedia modelling
Yi-Ping Phoebe Chen, Tat-Seng Chua |
Vis. Comput. | 1 |
| 2005 | Preface
Yi-Ping Phoebe Chen, Limsoon Wong |
APBC | 1 |
| 2005 | A New Indexing Method for High Dimensional Dataset
Jiyuan An, Yi-Ping Phoebe Chen, Qinying Xu, Xiaofang Zhou 0001 |
DASFAA | 2 |
| 2005 | Yet Another Induction Algorithm
Jiyuan An, Yi-Ping Phoebe Chen |
KES (2) | 2 |
| 2005 | Multi-level Semantic Analysis for Sports Video
Dian Tjondronegoro, Yi-Ping Phoebe Chen |
KES (2) | 2 |
| 2005 | Content-based video indexing for sports applications using integrated multi-modal approachabstractTo sustain an ongoing rapid growth of video information, there is an emerging demand for a sophisticated content-based video indexing system. However, current video indexing solutions are still immature and lack of any standard. This doctoral consists of a research work based on an integrated multi-modal approach for sports video indexing and retrieval. By combining specific features extractable from multiple audio-visual modalities, generic structure and specific events can be detected and classified. During browsing and retrieval, users will benefit from the integration of high-level semantic and some descriptive mid-level features such as whistle and close-up view of player(s). Dian Tjondronegoro, Yi-Ping Phoebe Chen, Binh Pham 0001 |
ACM Multimedia | 2 |
| 2005 | DDR: an index method for large time-series datasets
Jiyuan An, Yi-Ping Phoebe Chen, Hanxiong Chen |
Inf. Syst. | 2 |
| 2004 | A Computational Framework for Nucleic Acid Sub-Sequence IdentificationabstractIdentification of nucleic acid sub-sequences within larger background sequences is a fundamental need of the biology community. The applicability correlates to research studies looking for homologous regions, diagnostic purposes and many other related activities. This paper serves to detail the approaches taken leading to sub-sequence identification through the use of hidden Markov models and associated scoring optimisations. The investigation of techniques for locating conserved basal promoter elements correlates to promoter thus gene identification techniques. The case study centred on the TATA box basal promoter element, as such the background is a gene sequence with the TATA box the target. Outcomes from the research conducted, highlights generic algorithms for sub-sequence identification, as such these generic processes can be transposed to any case study where identification of a target sequence is required. Paths extending from the work conducted in this investigation have led to the development of a generic framework for the future applicability of hidden Markov models to biological sequence analysis in a computational context. Scott Mann, Yi-Ping Phoebe Chen, Luke Eaton |
BIBE | 2 |
| 2004 | Classification of self-consumable highlights for soccer video summariesabstractAn effective scheme for soccer summarization is significant to improve the usage of this massively growing video data. The work presents an extension to our recent work which proposed a framework to integrate highlights into play-breaks to construct more complete soccer summaries. The current focus is to demonstrate the benefits of detecting some specific audio-visual features during play-break sequences in order to classify highlights contained within them. The main purpose is to generate summaries which are self-consumable individually. To support this framework, the algorithms for shot classification and detection of near-goal and slow-motion replay scenes is described. The results of our experiment using 5 soccer videos (20 minutes each) show the performance and reliability of our framework. Dian Tjondronegoro, Yi-Ping Phoebe Chen, Binh Pham 0001 |
ICME | 2 |
| 2004 | Surface Spatial Index Structure of High-Dimensional Space
Jiyuan An, Yi-Ping Phoebe Chen, Qinying Xu |
IDEAL | 2 |
| 2004 | A Grid-Based Index Method for Time Warping Distance
Jiyuan An, Yi-Ping Phoebe Chen, Eamonn J. Keogh |
WAIM | 2 |
| 2004 | Concept Learning of Text DocumentsabstractConcept learning of text documents can be viewed as the problem of acquiring the definition of a general category of documents. To definite the category of a text document, the Conjunctive of keywords is usually be used. These keywords should be fewer and comprehensible. A naïve method is enumerating all combinations of keywords to extract suitable ones. However, because of the enormous number of keyword combinations, it is impossible to extract the most relevant keywords to describe the categories of documents by enumerating all possible combinations of keywords. Many heuristic methods are proposed, such as GA-base, immune based algorithm. In this work, we introduce pruning power technique and propose a robust enumeration-based concept learning algorithm. Experimental results show that the rules produce by our approach has more comprehensible and simplicity than by other methods. Jiyuan An, Yi-Ping Phoebe Chen |
Web Intelligence | 2 |
| 2004 | Automatic Pattern-Taxonomy Extraction for Web MiningabstractIn this paper, we propose a model for discovering frequent sequential patterns, phrases, which can be used as profile descriptors of documents. It is indubitable that we can obtain numerous phrases using data mining algorithms. However, it is difficult to use these phrases effectively for answering what users want. Therefore, we present a pattern taxonomy extraction model which performs the task of extracting descriptive frequent sequential patterns by pruning the meaningless ones. The model then is extended and tested by applying it to the information filtering system. The results of the experiment show that pattern-based methods outperform the keyword-based methods. The results also indicate that removal of meaningless patterns not only reduces the cost of computation but also improves the effectiveness of the system. Sheng-Tang Wu, Yuefeng Li 0001, Yue Xu 0001, Binh Pham 0001, Yi-Ping Phoebe Chen |
Web Intelligence | 5 |
| 2003 | XML-Based Multimedia Query System
Yi-Ping Phoebe Chen |
MMM | 1 |
| 2003 | Guest Editor'S Introduction
Yi-Ping Phoebe Chen |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2003 | Database Technologies for L-System Simulations in Virtual Plant Applications on Bioinformatics
Yi-Ping Phoebe Chen, Robert M. Colomb |
Knowl. Inf. Syst. | 1 |
| 2002 | Derivation of L-system Models from Measurements of Biological Branching Structures Using Genetic Algorithms
Bian Runqiang, Yi-Ping Phoebe Chen, Kevin Burrage, Jim Hanan, Peter Room, John A. Belward |
IEA/AIE | 2 |
| 2002 | Content-based Indexing and Retrieval Using MPEG-7 and X-Query in Video Data Management Systems
Dian Tjondronegoro, Yi-Ping Phoebe Chen |
World Wide Web | 2 |
| 2001 | A New Approach to Hybrid SOM Implementations for Text ClassificationabstractThis paper analyses several recent treatises on hybridised self-organising map (SOM) theory. Each article proposes a solution to expedite the SOM mapping process and provides more accurate results within a shorter response time via hybridisation: including utilisation of Bayesian classification techniques; an interactive associative search and exploration tool; and the use of a hierarchical organization of tiered SOM's with input derived via auto-associative feedforward neural network technology. In this paper, we propose that an amalgamation of SOM and association rule theory may hold the key to a more generic solution, less reliant on initial supervision and redundant user interaction. The results of clustering stem words from text documents could be utilised to derive association rules which designate the applicability of documents to the user. A four stage process is consequently detailed, demonstrating a generic example of how a graphical derivation of associations may be derived from a repository of text documents, or even a set of synopses of many such repositories. Paul Gunther, Yi-Ping Phoebe Chen |
FUZZ-IEEE | 2 |
| 2001 | Fuzzy Genetic Algorithms Based on Level Interval AlgorithmabstractMany decisions need to be made based on imprecise or incomplete initial information. In such cases, decision makers are generally more interested in sets of the most promising solutions rather than the best single solution. Therefore, in contrast to conventional optimisation approaches that aim to find exact optimal points, we aim to find optimal ranges with variable satisfaction degrees. The paper presents a fuzzy-set-based approach for the representation and optimisation of practical problems with imprecise properties where evolutionary computation is used for obtaining fuzzy solutions through guided searching. The representation of fuzzy sets, its initialisation, crossover, mutation, and validation, the ranking approach for fuzzy objective values, and the propagation method of fuzzy information are discussed. Several examples for illustrating the fuzzy evolutionary optimisation approach are provided. Jinglan Zhang, Binh Pham 0001, Yi-Ping Phoebe Chen |
FUZZ-IEEE | 3 |
| 1999 | DML: A Bridge between Database Systems and L-Systems for Biological ResearchabstractPresents DML (Data Model tools for L-systems). This work combines a novel application area with databases in bioinformatics. DML allows the specification of database structures and queries in terms of objects and relations that are specific to scientific L-system applications. L-systems have been used in many scientific areas. However, not much consideration has been given to persistent storage and querying of the large quantities of data resulting from these L-system simulations. Major limitations of current database products mean that they cannot support the mapping between L-systems and database systems. We find it hard to use any of the existing database products to build a system to satisfy the needs of researchers using L-systems. The research presented in this paper focuses on providing a bridge between L-systems and database systems. The work contributes a schema generator and population translator from L-systems to database systems in the scientific area of generic branching structures. Yi-Ping Phoebe Chen |
SSDBM | 1 |