EDBT 2026 Demo / reviewers in the wild / expert
Tong Pan
dblp:17/357
· DBLP profile ↗
12ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CoPRA: Bridging Cross-domain Pretrained Sequence Models with Complex Structures for Protein-RNA Binding Affinity PredictionabstractAccurately measuring protein-RNA binding affinity is crucial in many biological processes and drug design. Previous computational methods for protein-RNA binding affinity prediction rely on either sequence or structure features, unable to capture the binding mechanisms comprehensively. The recent emerging pre-trained language models trained on massive unsupervised sequences of protein and RNA have shown strong representation ability for various in-domain downstream tasks, including binding site prediction. However, applying different-domain language models collaboratively for complex-level tasks remains unexplored. In this paper, we propose CoPRA to bridge pre-trained language models from different biological domains via Complex structure for Protein-RNA binding Affinity prediction. We demonstrate for the first time that cross-biological modal language models can collaborate to improve binding affinity prediction. We propose a Co-Former to combine the cross-modal sequence and structure information and a bi-scope pre-training strategy for improving Co-Former's interaction understanding. Meanwhile, we build the largest protein-RNA binding affinity dataset PRA310 for performance evaluation. We also test our model on a public dataset for mutation effect prediction. CoPRA reaches state-of-the-art performance on all the datasets. We provide extensive analyses and verify that CoPRA can (1) accurately predict the protein-RNA binding affinity; (2) understand the binding affinity change caused by mutations; and (3) benefit from scaling data and model size. Xiaohong Liu 0007, Tong Pan, Jing Xu 0008, Xiaoyu Wang 0016, Wuyang Lan, Jiangning Song, Ting Chen 0006 |
AAAI | 3 |
| 2025 | Integrating Graph Convolutional Networks for Missing Gene Expression ImputationabstractSingle-cell RNA sequencing (scRNA-seq) techniques are emerging to revolutionize modern biomedical sciences by providing a detailed landscape of individual cells. However, these methods often lack crucial spatial localization information. To address this gap, spatial transcriptomic technologies have developed, enabling gene expression profiling while mapping cells spatial information. Yet, the gene throughput in spatial transcriptomic technologies makes it challenging to characterize whole-transcriptome-level data for single cells in space. In this context, approaches for predicting the spatial distribution of genes are still under development. Here, we present GCNgene, a novel method to predict the spatial distribution of the undetected RNA transcripts, through integrating spatial and scRNA-seq datasets. GCNgene leverages a graph convolutional network to embed spatial transcriptomics data and then applies a learned rule to reconstruct gene expression by combining the reference single-cell data with the calculated cell-type proportions. Ultimately, this learned paradigm enables accurate predictions of gene expression levels. Ying Zhang 0053, Hong-Jin Yu, Zihao Yan, Tong Pan, Yan Liu 0038, Shanshan Li 0008, Yuming Guo 0001, Jiangning Song, Dongjun Yu |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | Evaluating the Generalization Ability of Spatiotemporal Model in Urban ScenarioabstractSpatiotemporal neural networks have shown great promise in urban scenarios by effectively capturing temporal and spatial correlations. However, urban environments are constantly evolving, and current model evaluations are often limited to traffic scenarios and use data mainly collected only a few weeks after training period to evaluate model performance. The generalization ability of these models remains largely unexplored. To address this, we propose a Spatiotemporal Out-of-Distribution (ST-OOD) benchmark, which comprises six urban scenario: bike-sharing, 311 services, pedestrian counts, traffic speed, traffic flow, ride-hailing demand, and bike-sharing, each with in-distribution (same year) and out-of-distribution (next years) settings. We extensively evaluate state-of-the-art spatiotemporal models and find that their performance degrades significantly in out-of-distribution settings, with most models performing even worse than a simple Multi-Layer Perceptron (MLP). Our findings suggest that current leading methods tend to over-rely on parameters to overfit training data, which may lead to good performance on in-distribution data but often results in poor generalization. We also investigated whether dropout could mitigate the negative effects of overfitting. Our results showed that a slight dropout rate could significantly improve generalization performance on most datasets, with minimal impact on in-distribution performance. However, balancing in-distribution and out-of-distribution performance remains a challenging problem. We hope that the proposed benchmark will encourage further research on this critical issue. Hongjun Wang 0007, Jiyuan Chen, Tong Pan, Zheng Dong 0006, Renhe Jiang, Xuan Song 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Enhancing low-resource cross-lingual summarization from noisy data with fine-grained reinforcement learningabstractCross-lingual summarization (CLS) is the task of generating a summary in a target language from a document in a source language. Recently, end-to-end CLS models have achieved impressive results using large-scale, high-quality datasets typically constructed by translating monolingual summary corpora into CLS corpora. However, due to the limited performance of low-resource language translation models, translation noise can seriously degrade the performance of these models. In this paper, we propose a fine-grained reinforcement learning approach to address low-resource CLS based on noisy data. We introduce the source language summary as a gold signal to alleviate the impact of the translated noisy target summary. Specifically, we design a reinforcement reward by calculating the word correlation and word missing degree between the source language summary and the generated target language summary, and combine it with cross-entropy loss to optimize the CLS model. To validate the performance of our proposed model, we construct Chinese-Vietnamese and Vietnamese-Chinese CLS datasets. Experimental results show that our proposed model outperforms the baselines in terms of both the ROUGE score and BERTScore. Yuxin Huang 0004, Huailing Gu, Zhengtao Yu 0001, Yumeng Gao, Tong Pan, Jialong Xu |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2023 | Easy Begun Is Half Done: Spatial-Temporal Graph Modeling with ST-Curriculum DropoutabstractSpatial-temporal (ST) graph modeling, such as traffic speed forecasting and taxi demand prediction, is an important task in deep learning area. However, for the nodes in the graph, their ST patterns can vary greatly in difficulties for modeling, owning to the heterogeneous nature of ST data. We argue that unveiling the nodes to the model in a meaningful order, from easy to complex, can provide performance improvements over traditional training procedure. The idea has its root in Curriculum Learning, which suggests in the early stage of training models can be sensitive to noise and difficult samples. In this paper, we propose ST-Curriculum Dropout, a novel and easy-to-implement strategy for spatial-temporal graph modeling. Specifically, we evaluate the learning difficulty of each node in high-level feature space and drop those difficult ones out to ensure the model only needs to handle fundamental ST relations at the beginning, before gradually moving to hard ones. Our strategy can be applied to any canonical deep learning architecture without extra trainable parameters, and extensive experiments on a wide range of datasets are conducted to illustrate that, by controlling the difficulty level of ST relations as the training progresses, the model is able to capture better representation of the data and thus yields better generalization. Hongjun Wang 0007, Jiyuan Chen, Tong Pan, Zipei Fan, Xuan Song 0001, Renhe Jiang, Lingyu Zhang 0001, Boyuan Zhang 0005 |
AAAI | 3 |
| 2023 | SMG: self-supervised masked graph learning for cancer gene identificationabstractCancer genomics is dedicated to elucidating the genes and pathways that contribute to cancer progression and development. Identifying cancer genes (CGs) associated with the initiation and progression of cancer is critical for characterization of molecular-level mechanism in cancer research. In recent years, the growing availability of high-throughput molecular data and advancements in deep learning technologies has enabled the modelling of complex interactions and topological information within genomic data. Nevertheless, because of the limited labelled data, pinpointing CGs from a multitude of potential mutations remains an exceptionally challenging task. To address this, we propose a novel deep learning framework, termed self-supervised masked graph learning (SMG), which comprises SMG reconstruction (pretext task) and task-specific fine-tuning (downstream task). In the pretext task, the nodes of multi-omic featured protein-protein interaction (PPI) networks are randomly substituted with a defined mask token. The PPI networks are then reconstructed using the graph neural network (GNN)-based autoencoder, which explores the node correlations in a self-prediction manner. In the downstream tasks, the pre-trained GNN encoder embeds the input networks into feature graphs, whereas a task-specific layer proceeds with the final prediction. To assess the performance of the proposed SMG method, benchmarking experiments are performed on three node-level tasks (identification of CGs, essential genes and healthy driver genes) and one graph-level task (identification of disease subnetwork) across eight PPI networks. Benchmarking experiments and performance comparison with existing state-of-the-art methods demonstrate the superiority of SMG on multi-omic feature engineering. Yan Cui 0008, Zhikang Wang, Xiaoyu Wang 0016, Ying Zhang 0053, Tong Pan, Shanshan Li 0008, Yuming Guo 0001, Tatsuya Akutsu, Jiangning Song |
Briefings Bioinform. | 6 |
| 2023 | PFresGO: an attention mechanism-based deep-learning approach for protein annotation by integrating gene ontology inter-relationshipsabstractMOTIVATION: The rapid accumulation of high-throughput sequence data demands the development of effective and efficient data-driven computational methods to functionally annotate proteins. However, most current approaches used for functional annotation simply focus on the use of protein-level information but ignore inter-relationships among annotations. RESULTS: Here, we established PFresGO, an attention-based deep-learning approach that incorporates hierarchical structures in Gene Ontology (GO) graphs and advances in natural language processing algorithms for the functional annotation of proteins. PFresGO employs a self-attention operation to capture the inter-relationships of GO terms, updates its embedding accordingly and uses a cross-attention operation to project protein representations and GO embedding into a common latent space to identify global protein sequence patterns and local functional residues. We demonstrate that PFresGO consistently achieves superior performance across GO categories when compared with 'state-of-the-art' methods. Importantly, we show that PFresGO can identify functionally important residues in protein sequences by assessing the distribution of attention weightings. PFresGO should serve as an effective tool for the accurate functional annotation of proteins and functional domains within proteins. AVAILABILITY AND IMPLEMENTATION: PFresGO is available for academic purposes at https://github.com/BioColLab/PFresGO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tong Pan, Chen Li 0021, Yue Bi, Zhikang Wang, Robin B. Gasser, Anthony W. Purcell, Tatsuya Akutsu, Geoffrey I. Webb, Seiya Imoto, Jiangning Song |
Bioinform. | 1 |
| 2023 | Targeting tumor heterogeneity: multiplex-detection-based multiple instance learning for whole slide image classificationabstractMOTIVATION: Multiple instance learning (MIL) is a powerful technique to classify whole slide images (WSIs) for diagnostic pathology. The key challenge of MIL on WSI classification is to discover the critical instances that trigger the bag label. However, tumor heterogeneity significantly hinders the algorithm's performance. RESULTS: Here, we propose a novel multiplex-detection-based multiple instance learning (MDMIL) which targets tumor heterogeneity by multiplex detection strategy and feature constraints among samples. Specifically, the internal query generated after the probability distribution analysis and the variational query optimized throughout the training process are utilized to detect potential instances in the form of internal and external assistance, respectively. The multiplex detection strategy significantly improves the instance-mining capacity of the deep neural network. Meanwhile, a memory-based contrastive loss is proposed to reach consistency on various phenotypes in the feature space. The novel network and loss function jointly achieve high robustness towards tumor heterogeneity. We conduct experiments on three computational pathology datasets, e.g. CAMELYON16, TCGA-NSCLC, and TCGA-RCC. Benchmarking experiments on the three datasets illustrate that our proposed MDMIL approach achieves superior performance over several existing state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: MDMIL is available for academic purposes at https://github.com/ZacharyWang-007/MDMIL. Zhikang Wang, Yue Bi, Tong Pan, Xiaoyu Wang 0016, Chris Bain, Richard Bassed, Seiya Imoto, Jianhua Yao 0001, Roger J. Daly, Jiangning Song |
Bioinform. | 3 |
| 2022 | Clarion is a multi-label problem transformation method for identifying mRNA subcellular localizationsabstractSubcellular localization of messenger RNAs (mRNAs) plays a key role in the spatial regulation of gene activity. The functions of mRNAs have been shown to be closely linked with their localizations. As such, understanding of the subcellular localizations of mRNAs can help elucidate gene regulatory networks. Despite several computational methods that have been developed to predict mRNA localizations within cells, there is still much room for improvement in predictive performance, especially for the multiple-location prediction. In this study, we proposed a novel multi-label multi-class predictor, termed Clarion, for mRNA subcellular localization prediction. Clarion was developed based on a manually curated benchmark dataset and leveraged the weighted series method for multi-label transformation. Extensive benchmarking tests demonstrated Clarion achieved competitive predictive performance and the weighted series method plays a crucial role in securing superior performance of Clarion. In addition, the independent test results indicate that Clarion outperformed the state-of-the-art methods and can secure accuracy of 81.47, 91.29, 79.77, 92.10, 89.15, 83.74, 80.74, 79.23 and 84.74% for chromatin, cytoplasm, cytosol, exosome, membrane, nucleolus, nucleoplasm, nucleus and ribosome, respectively. The webserver and local stand-alone tool of Clarion is freely available at http://monash.bioweb.cloud.edu.au/Clarion/. Yue Bi, Fuyi Li, Zhikang Wang, Tong Pan, Yuming Guo 0001, Geoffrey I. Webb, Jianhua Yao 0001, Cangzhi Jia, Jiangning Song |
Briefings Bioinform. | 5 |
| 2019 | Cell-like spiking neural P systems with evolution rules
Tong Pan, Suxia Jiang |
Soft Comput. | 1 |
| 2018 | A small universal spiking neural P system with communication on request
Tong Pan |
Neurocomputing | 1 |
| 2007 | A Framework of Agent-Based Collaborative Intelligent Transport SystemabstractIntelligent cooperative decision-making and wireless communication are more and more used in intelligent transport systems. In this paper, According to the thought of agent-oriented collaborative and use wireless communication, we construct an intelligent transport system framework based on the agent, and set up the experimental environment to realize this construction. It achieves the vehicles individual regarding the urgent region in event agile response, simultaneously can cause outside the urgent region the vehicles individual to make the corresponding decision-making. Zhiyi Fang, Guannan Qu, Hongjun Yang, Tong Pan |
CSCWD | 5 |