EDBT 2026 Demo / reviewers in the wild / expert
Ying-Ying Xu
dblp:29/7299
· DBLP profile ↗
17ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-3185-811XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCPLoc: A weakly-supervised multi-instance learning framework for protein subcellular localization and heterogeneity profiling in single-cell immunofluorescence images
Yi-Lin Li, Ying-Yi Wang, Xi-Liang Zhu, Ying-Ying Xu |
Neurocomputing | 6 |
| 2025 | Knowledge-enhanced protein subcellular localization prediction from 3D fluorescence microscope imagesabstractMOTIVATION: Pinpointing the subcellular location of proteins is essential for studying protein function and related diseases. Advances in spatial proteomics have shown that automatic recognition of protein subcellular localization from images could highly facilitate protein translocation analysis and biomarker discovery, but existing machine-learning works have been mostly limited to processing 2D images. By contrast, 3D images have higher spatial resolution and allow researchers to observe cellular structures in their natural context, but currently, there are only a few studies of 3D image processing for protein distribution analysis due to the lack of data and complexity of modeling. RESULTS: We developed a knowledge-enhanced protein subcellular localization model, KE3DLoc, which could recognize distribution patterns in 3D fluorescence microscope images using deep learning methods. The model designs an image feature extraction module that incorporates information from 3D and 2D projected cells and implements asymmetric loss and confidence weights to address data imbalance and weak cell annotation issues. Besides, considering that the biological knowledge in the Gene Ontology (GO) database can provide valuable support for protein location understanding, the KE3DLoc model incorporates a novel knowledge enhancement module that optimizes the protein representation by related knowledge graphs derived from the GO. Since the image module and the knowledge module calculate features from different levels, KE3DLoc designs protein ID aggregation to enhance the consistency of protein features across different cells. Experimental results on three public datasets have demonstrated that the KE3DLoc significantly outperforms existing methods and provides valuable insights for spatial proteomics research. AVAILABILITY AND IMPLEMENTATION: All datasets and codes used in this study are available at GitHub: https://github.com/PRBioimages/KE3DLoc. Guo-Hua Zeng, Xing-Zheng Zhu, Hong-Rui Yang, Yong-Jia Liang, Yu-Jia Zhai, Ying-Ying Xu |
Bioinform. | 6 |
| 2023 | Automatic recognition of protein subcellular location patterns in single cells from immunofluorescence images based on deep learningabstractWith the improvement of single-cell measurement techniques, there is a growing awareness that individual differences exist among cells, and protein expression distribution can vary across cells in the same tissue or cell line. Pinpointing the protein subcellular locations in single cells is crucial for mapping functional specificity of proteins and studying related diseases. Currently, research about single-cell protein location is still in its infancy, and most studies and databases do not annotate proteins at the cell level. For example, in the human protein atlas database, an immunofluorescence image stained for a particular protein shows multiple cells, but the subcellular location annotation is for the whole image, ignoring intercellular difference. In this study, we used large-scale immunofluorescence images and image-level subcellular locations to develop a deep-learning-based pipeline that could accurately recognize protein localizations in single cells. The pipeline consisted of two deep learning models, i.e. an image-based model and a cell-based model. The former used a multi-instance learning framework to comprehensively model protein distribution in multiple cells in each image, and could give both image-level and cell-level predictions. The latter firstly used clustering and heuristics algorithms to assign pseudo-labels of subcellular locations to the segmented cell images, and then used the pseudo-labels to train a classification model. Finally, the image-based model was fused with the cell-based model at the decision level to obtain the final ensemble model for single-cell prediction. Our experimental results showed that the ensemble model could achieve higher accuracy and robustness on independent test sets than state-of-the-art methods. Xi-Liang Zhu, Lin-Xia Bao, Min-Qi Xue, Ying-Ying Xu |
Briefings Bioinform. | 4 |
| 2022 | Learning protein subcellular localization multi-view patterns from heterogeneous data of imaging, sequence and networksabstractLocation proteomics seeks to provide automated high-resolution descriptions of protein location patterns within cells. Many efforts have been undertaken in location proteomics over the past decades, thereby producing plenty of automated predictors for protein subcellular localization. However, most of these predictors are trained solely from high-throughput microscopic images or protein amino acid sequences alone. Unifying heterogeneous protein data sources has yet to be exploited. In this paper, we present a pipeline called sequence, image, network-based protein subcellular locator (SIN-Locator) that constructs a multi-view description of proteins by integrating multiple data types including images of protein expression in cells or tissues, amino acid sequences and protein-protein interaction networks, to classify the patterns of protein subcellular locations. Proteins were encoded by both handcrafted features and deep learning features, and multiple combining methods were implemented. Our experimental results indicated that optimal integrations can considerately enhance the classification accuracy, and the utility of SIN-Locator has been demonstrated through applying to new released proteins in the human protein atlas. Furthermore, we also investigate the contribution of different data sources and influence of partial absence of data. This work is anticipated to provide clues for reconciliation and combination of multi-source data for protein location analysis. Min-Qi Xue, Hong-Bin Shen, Ying-Ying Xu |
Briefings Bioinform. | 4 |
| 2022 | GraphLoc: a graph neural network model for predicting protein subcellular localization from immunohistochemistry imagesabstractMOTIVATION: Recognition of protein subcellular distribution patterns and identification of location biomarker proteins in cancer tissues are important for understanding protein functions and related diseases. Immunohistochemical (IHC) images enable visualizing the distribution of proteins at the tissue level, providing an important resource for the protein localization studies. In the past decades, several image-based protein subcellular location prediction methods have been developed, but the prediction accuracies still have much space to improve due to the complexity of protein patterns resulting from multi-label proteins and the variation of location patterns across cell types or states. RESULTS: Here, we propose a multi-label multi-instance model based on deep graph convolutional neural networks, GraphLoc, to recognize protein subcellular location patterns. GraphLoc builds a graph of multiple IHC images for one protein, learns protein-level representations by graph convolutions and predicts multi-label information by a dynamic threshold method. Our results show that GraphLoc is a promising model for image-based protein subcellular location prediction with model interpretability. Furthermore, we apply GraphLoc to the identification of candidate location biomarkers and potential members for protein networks. A large portion of the predicted results have supporting evidence from the existing literatures and the new candidates also provide guidance for further experimental screening. AVAILABILITY AND IMPLEMENTATION: The dataset and code are available at: www.csbio.sjtu.edu.cn/bioinf/GraphLoc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jin-Xian Hu, Yang Yang 0030, Ying-Ying Xu, Hong-Bin Shen |
Bioinform. | 3 |
| 2022 | DULoc: quantitatively unmixing protein subcellular location patterns in immunofluorescence images based on deep learning featuresabstractMOTIVATION: Knowledge of subcellular locations of proteins is of great significance for understanding their functions. The multi-label proteins that simultaneously reside in or move between more than one subcellular structure usually involve with complex cellular processes. Currently, the subcellular location annotations of proteins in most studies and databases are descriptive terms, which fail to capture the protein amount or fractions across different locations. This highly limits the understanding of complex spatial distribution and functional mechanism of multi-label proteins. Thus, quantitatively analyzing the multiplex location patterns of proteins is an urgent and challenging task. RESULTS: In this study, we developed a deep-learning-based pattern unmixing pipeline for protein subcellular localization (DULoc) to quantitatively estimate the fractions of proteins localizing in different subcellular compartments from immunofluorescence images. This model used a deep convolutional neural network to construct feature representations, and combined multiple nonlinear decomposing algorithms as the pattern unmixing method. Our experimental results showed that the DULoc can achieve over 0.93 correlation between estimated and true fractions on both real and synthetic datasets. In addition, we applied the DULoc method on the images in the human protein atlas database on a large scale, and showed that 70.52% of proteins can achieve consistent location orders with the database annotations. AVAILABILITY AND IMPLEMENTATION: The datasets and code are available at: https://github.com/PRBioimages/DULoc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Min-Qi Xue, Xi-Liang Zhu, Ying-Ying Xu |
Bioinform. | 4 |
| 2022 | Automated classification of protein expression levels in immunohistochemistry images to improve the detection of cancer biomarkersabstractBACKGROUND: The expression changes of some proteins are associated with cancer progression, and can be used as biomarkers in cancer diagnosis. Automated systems have been frequently applied in the large-scale detection of protein biomarkers and have provided a valuable complement for wet-laboratory experiments. For example, our previous work used an immunohistochemical image-based machine learning classifier of protein subcellular locations to screen biomarker proteins that change locations in colon cancer tissues. The tool could recognize the location of biomarkers but did not consider the effect of protein expression level changes on the screening process. RESULTS: In this study, we built an automated classification model that recognizes protein expression levels in immunohistochemical images, and used the protein expression levels in combination with subcellular locations to screen cancer biomarkers. To minimize the effect of non-informative sections on the immunohistochemical images, we employed the representative image patches as input and applied a Wasserstein distance method to determine the number of patches. For the patches and the whole images, we compared the ability of color features, characteristic curve features, and deep convolutional neural network features to distinguish different levels of protein expression and employed deep learning and conventional classification models. Experimental results showed that the best classifier can achieve an accuracy of 73.72% and an F1-score of 0.6343. In the screening of protein biomarkers, the detection accuracy improved from 63.64 to 95.45% upon the incorporation of the protein expression changes. CONCLUSIONS: Machine learning can distinguish different protein expression levels and speed up their annotation in the future. Combining information on the expression patterns and subcellular locations of protein can improve the accuracy of automatic cancer biomarker screening. This work could be useful in discovering new cancer biomarkers for clinical diagnosis and research. Zhenzhen Xue, Cheng Li 0008, Zhuo-Ming Luo, Shanshan Wang 0002, Ying-Ying Xu |
BMC Bioinform. | 5 |
| 2020 | Learning complex subcellular distribution patterns of proteins via analysis of immunohistochemistry imagesabstractMOTIVATION: Systematic and comprehensive analysis of protein subcellular location as a critical part of proteomics ('location proteomics') has been studied for many years, but annotating protein subcellular locations and understanding variation of the location patterns across various cell types and states is still challenging. RESULTS: In this work, we used immunohistochemistry images from the Human Protein Atlas as the source of subcellular location information, and built classification models for the complex protein spatial distribution in normal and cancerous tissues. The models can automatically estimate the fractions of protein in different subcellular locations, and can help to quantify the changes of protein distribution from normal to cancer tissues. In addition, we examined the extent to which different annotated protein pathways and complexes showed similarity in the locations of their member proteins, and then predicted new potential proteins for these networks. AVAILABILITY AND IMPLEMENTATION: The dataset and code are available at: www.csbio.sjtu.edu.cn/bioinf/complexsubcellularpatterns. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ying-Ying Xu, Hong-Bin Shen, Robert F. Murphy |
Bioinform. | 1 |
| 2020 | Automated classification of protein subcellular localization in immunohistochemistry images to reveal biomarkers in colon cancerabstractBACKGROUND: Protein biomarkers play important roles in cancer diagnosis. Many efforts have been made on measuring abnormal expression intensity in biological samples to identity cancer types and stages. However, the change of subcellular location of proteins, which is also critical for understanding and detecting diseases, has been rarely studied. RESULTS: In this work, we developed a machine learning model to classify protein subcellular locations based on immunohistochemistry images of human colon tissues, and validated the ability of the model to detect subcellular location changes of biomarker proteins related to colon cancer. The model uses representative image patches as inputs, and integrates feature engineering and deep learning methods. It achieves 92.69% accuracy in classification of new proteins. Two validation datasets of colon cancer biomarkers derived from published literatures and the human protein atlas database respectively are employed. It turns out that 81.82 and 65.66% of the biomarker proteins can be identified to change locations. CONCLUSIONS: Our results demonstrate that using image patches and combining predefined and deep features can improve the performance of protein subcellular localization, and our model can effectively detect biomarkers based on protein subcellular translocations. This study is anticipated to be useful in annotating unknown subcellular localization for proteins and discovering new potential location biomarkers. Zhenzhen Xue, Qing-Zu Gao, Ying-Ying Xu |
BMC Bioinform. | 5 |
| 2018 | Bioimage-based protein subcellular location prediction: a comprehensive review
Ying-Ying Xu, Li-Xiu Yao, Hong-Bin Shen |
Frontiers Comput. Sci. | 1 |
| 2018 | An Organelle Correlation-Guided Feature Selection Approach for Classifying Multi-Label Subcellular Bio-ImagesabstractNowadays, with the advances in microscopic imaging, accurate classification of bioimage-based protein subcellular location pattern has attracted as much attention as ever. One of the basic challenging problems is how to select the useful feature components among thousands of potential features to describe the images. This is not an easy task especially considering there is a high ratio of multi-location proteins. Existing feature selection methods seldom take the correlation among different cellular compartments into consideration, and thus may miss some features that will be co-important for several subcellular locations. To deal with this problem, we make use of the important structural correlation among different cellular compartments and propose an organelle structural correlation regularized feature selection method CSF (Common-Sets of Features) in this paper. We formulate the multi-label classification problem by adopting a group-sparsity regularizer to select common subsets of relevant features from different cellular compartments. In addition, we also add a cell structural correlation regularized Laplacian term, which utilizes the prior biological structural information to capture the intrinsic dependency among different cellular compartments. The CSF provides a new feature selection strategy for multi-label bio-image subcellular pattern classifications, and the experimental results also show its superiority when comparing with several existing algorithms. Wei Shao 0005, Mingxia Liu 0001, Ying-Ying Xu, Hong-Bin Shen, Daoqiang Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2016 | Internet-Based Urban Bus Travelling Data Acquisition and Missing Data RecoveryabstractSensing of urban bus travelling data based on data collection from the Internet enables researchers from different regions to focus on the same study objects and promote the development of corresponding theories and technology. In this study, seven processes running on two servers were employed to retrieve real-time bus travelling data from the official bus information inquiry website of Suzhou, China. For data acquisition, a frequent problem encountered is the missing of bus travelling data caused by GPS signal fluctuation and network failure. In order to recover the missing data, we proposed Knearest-neighbor linear regression (KNN-LR) for data interpolation. Leave-one-out cross-validation is adapted to assess the performance of the proposed approach. It is demonstrated that the proposed KNN-LR method is robust and suitable for recovering missing data of bus trajectories in practice. Changfei Tong, Ying-Ying Xu, Huaizhong Li |
MSN | 4 |
| 2016 | Incorporating organelle correlations into semi-supervised learning for protein subcellular localization predictionabstractMOTIVATION: Bioimages of subcellular protein distribution as a new data source have attracted much attention in the field of automated prediction of proteins subcellular localization. Performance of existing systems is significantly limited by the small number of high-quality images with explicit annotations, resulting in the small sample size learning problem. This limitation is more serious for the multi-location proteins that co-exist at two or more organelles, because it is difficult to accurately annotate those proteins by biological experiments or automated systems. RESULTS: In this study, we designed a new protein subcellular localization prediction pipeline aiming to deal with the small sample size learning and multi-location proteins annotation problems. Five semi-supervised algorithms that can make use of lower-quality data were integrated, and a new multi-label classification approach by incorporating the correlations among different organelles in cells was proposed. The organelle correlations were modeled by the Bayesian network, and the topology of the correlation graph was used to guide the order of binary classifiers training in the multi-label classification to reflect the label dependence relationship. The proposed protocol was applied on both immunohistochemistry and immunofluorescence images, and our experimental results demonstrated its efficiency. AVAILABILITY AND IMPLEMENTATION: The datasets and code are available at: www.csbio.sjtu.edu.cn/bioinf/CorrASemiB CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ying-Ying Xu, Hong-Bin Shen |
Bioinform. | 1 |
| 2016 | Enhancing the Prediction of Transmembrane β-Barrel Segments with Chain Learning and Feature Sparse RepresentationabstractTransmembrane β-barrels (TMBs) are one important class of membrane proteins that play crucial functions in the cell. Membrane proteins are difficult wet-lab targets of structural biology, which call for accurate computational prediction approaches. Here, we developed a novel method named MemBrain-TMB to predict the spanning segments of transmembrane β-barrel from amino acid sequence. MemBrain-TMB is a statistical machine learning-based model, which is constructed using a new chain learning algorithm with input features encoded by the image sparse representation approach. We considered the relative status information between neighboring residues for enhancing the performance, and the matrix of features was translated into feature image by sparse coding algorithm for noise and dimension reduction. To deal with the diverse loop length problem, we applied a dynamic threshold method, which is particularly useful for enhancing the recognition of short loops and tight turns. Our experiments demonstrate that the new protocol designed in MemBrain-TMB effectively helps improve prediction performance. Xi Yin 0004, Ying-Ying Xu, Hong-Bin Shen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2015 | Bioimaging-based detection of mislocalized proteins in human cancers by semi-supervised learningabstractMOTIVATION: There is a long-term interest in the challenging task of finding translocated and mislocated cancer biomarker proteins. Bioimages of subcellular protein distribution are new data sources which have attracted much attention in recent years because of their intuitive and detailed descriptions of protein distribution. However, automated methods in large-scale biomarker screening suffer significantly from the lack of subcellular location annotations for bioimages from cancer tissues. The transfer prediction idea of applying models trained on normal tissue proteins to predict the subcellular locations of cancerous ones is arbitrary because the protein distribution patterns may differ in normal and cancerous states. RESULTS: We developed a new semi-supervised protocol that can use unlabeled cancer protein data in model construction by an iterative and incremental training strategy. Our approach enables us to selectively use the low-quality images in normal states to expand the training sample space and provides a general way for dealing with the small size of annotated images used together with large unannotated ones. Experiments demonstrate that the new semi-supervised protocol can result in improved accuracy and sensitivity of subcellular location difference detection. AVAILABILITY AND IMPLEMENTATION: The data and code are available at: www.csbio.sjtu.edu.cn/bioinf/SemiBiomarker/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ying-Ying Xu, Yang Zhang 0040, Hong-Bin Shen |
Bioinform. | 1 |
| 2014 | Image-based classification of protein subcellular location patterns in human reproductive tissue by ensemble learning global and local features
Ying-Ying Xu, Shitong Wang 0001, Hong-Bin Shen |
Neurocomputing | 2 |
| 2013 | An image-based multi-label human protein subcellular localization predictor (iLocator) reveals protein mislocalizations in cancer tissuesabstractMOTIVATION: Human cells are organized into compartments of different biochemical cellular processes. Having proteins appear at the right time to the correct locations in the cellular compartments is required to conduct their functions in normal cells, whereas mislocalization of proteins can result in pathological diseases, including cancer. RESULTS: To reveal the cancer-related protein mislocalizations, we developed an image-based multi-label subcellular location predictor, iLocator, which covers seven cellular localizations. The iLocator incorporates both global and local image descriptors and generates predictions by using an ensemble multi-label classifier. The algorithm has the ability to treat both single- and multiple-location proteins. We first trained and tested iLocator on 3240 normal human tissue images that have known subcellular location information from the human protein atlas. The iLocator was then used to generate protein localization predictions for 3696 protein images from seven cancer tissues that have no location annotations in the human protein atlas. By comparing the output data from normal and cancer tissues, we detected eight potential cancer biomarker proteins that have significant localization differences with P-value < 0.01. AVAILABILITY: http://www.csbio.sjtu.edu.cn/bioinf/iLocator/ Ying-Ying Xu, Yang Zhang 0040, Hong-Bin Shen |
Bioinform. | 1 |