EDBT 2026 Demo / reviewers in the wild / expert
Yu Li 0006
dblp:34/2997-6
· DBLP profile ↗
56ranked-venue papers
8as first author
37since 2021 · last 2026
0000-0002-3664-6722ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 30 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorSystems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DS-ProGen: A Dual-Structure Deep Language Model for Functional Protein DesignabstractInverse Protein Folding (IPF) is a critical subtask in the field of protein design, aiming to engineer amino acid sequences capable of folding correctly into a specified three-dimensional (3D) conformation. Although substantial progress has been achieved in recent years, existing methods generally rely on either backbone coordinates or molecular surface features alone, which restricts their ability to fully capture the complex chemical and geometric constraints necessary for precise sequence prediction. To address this limitation, we present DS-ProGen, a dual-structure deep language model for functional protein design, which integrates both backbone geometry and surface-level representations. By incorporating backbone coordinates as well as surface chemical and geometric descriptors into a next-amino-acid prediction paradigm, DS-ProGen is able to generate functionally relevant and structurally stable sequences while satisfying both global and local conformational constraints. On the PRIDE dataset, DS-ProGen attains the current state-of-the-art recovery rate of 61.47%, demonstrating the synergistic advantage of multi-modal structural encoding in protein design. Furthermore, DS-ProGen excels in predicting interactions with a variety of biological partners, including ligands, ions, and RNA, confirming its robust functional retention capabilities. Zikang Wang, Jiyue Jiang, Ziqian Lin, Dongchen He, Yuheng Shan, Yanruisheng Shao, Jiuming Wang, Yimin Fan, Yu Li 0006 |
AAAI | 14 |
| 2026 | Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMsabstractYu Li, Xiaoran Shang, Qizhi Pei, Yun Zhu, Xin Gao, Honglin Lin, Zhanping Zhong, Zhuoshi Pan, Zheng Liu, Xiaoyang Wang, Conghui He, Dahua Lin, Feng Zhao, Lijun Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yu Li 0006, Xiaoran Shang, Qizhi Pei, Yun Zhu 0007, Xin Gao 0001, Honglin Lin, Zhanping Zhong, Zhuoshi Pan, Xiaoyang Wang 0007, Conghui He, Dahua Lin, Feng Zhao 0004, Lijun Wu 0003 |
ACL (1) | 1 |
| 2026 | ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from ScratchabstractZheng Liu, Honglin Lin, Xiaoyang Wang, Xin Gao, Yu Li, Mengzhang Cai, Yun Zhu, Zhanping Zhong, Qizhi Pei, Zhuoshi Pan, Xiaoran Shang, Conghui He, Bin Cui, Wentao Zhang, Lijun Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Honglin Lin, Xiaoyang Wang 0007, Xin Gao 0001, Yu Li 0006, Mengzhang Cai, Yun Zhu 0007, Zhanping Zhong, Qizhi Pei, Zhuoshi Pan, Xiaoran Shang, Conghui He, Bin Cui 0001, Wentao Zhang 0001, Lijun Wu 0003 |
ACL (1) | 5 |
| 2026 | REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at OnceabstractZhuoshi Pan, Qizhi Pei, Yu Li, Zinan Tang, QiYao Sun, H. Vicky Zhao, Conghui He, Lijun Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhuoshi Pan, Qizhi Pei, Yu Li 0006, Zinan Tang 0001, Qiyao Sun, H. Vicky Zhao, Conghui He, Lijun Wu 0003 |
ACL (1) | 3 |
| 2026 | InverTune: A Backdoor Defense Method for Multimodal Contrastive Learning via Backdoor-Adversarial Correlation Analysis
Mengyuan Sun 0001, Yu Li 0006, Yunjie Ge, Bo Du 0001, Qian Wang 0002 |
NDSS | 2 |
| 2025 | A Strategic Coordination Framework of Small LMs Matches Large LMs in Data SynthesisabstractXin Gao, Qizhi Pei, Zinan Tang, Yu Li, Honglin Lin, Jiang Wu, Lijun Wu, Conghui He. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xin Gao 0001, Qizhi Pei, Zinan Tang 0001, Yu Li 0006, Honglin Lin, Jiang Wu 0003, Lijun Wu 0003, Conghui He |
ACL (1) | 4 |
| 2025 | MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction FusionabstractQizhi Pei, Lijun Wu, Zhuoshi Pan, Yu Li, Honglin Lin, Chenlin Ming, Xin Gao, Conghui He, Rui Yan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Qizhi Pei, Lijun Wu 0003, Zhuoshi Pan, Yu Li 0006, Honglin Lin, Chenlin Ming, Xin Gao 0001, Conghui He, Rui Yan 0001 |
ACL (1) | 4 |
| 2025 | DAPE V2: Process Attention Score as Feature Map for Length ExtrapolationabstractThe attention mechanism is a fundamental component of the Transformer model, contributing to interactions among distinct tokens. In general, the attention scores are determined simply by the key-query products. However, this work’s occasional trial (combining DAPE and NoPE) of including additional MLPs on attention scores without position encoding indicates that the classical key-query multiplication may limit the performance of Transformers. In this work, we conceptualize attention as a feature map and apply the convolution operator (for neighboring attention scores across different heads) to mimic the processing methods in computer vision. Specifically, the main contribution of this paper is identifying and interpreting the Transformer length extrapolation problem as a result of the limited expressiveness of the naive query and key dot product, and we successfully translate the length extrapolation issue into a well-understood feature map processing problem, which is called Convolutional Data-Adaptive Position Encoding (CDAPE).The novel insight, which can be adapted to various attention-related models, reveals that the current Transformer architecture has the potential for further evolution. Extensive experiments demonstrate that treating attention as a feature map and applying convolution as a processing method significantly enhances Transformer performance. Chuanyang Zheng, Yihang Gao, Jiankai Sun, Jingyao Li 0001, Minbin Huang, Xiaozhe Ren, Michael Kwok-Po Ng, Zhenguo Li, Yu Li 0006 |
ACL (1) | 12 |
| 2025 | RBPtool: A Deep Language Model Framework for Multi-Resolution RBP-RNA Binding Prediction and RNA Molecule DesignabstractJiyue Jiang, Yitao Xu, Zikang Wang, Yihan Ye, Yanruisheng Shao, Yuheng Shan, Jiuming Wang, Xiaodan Fan, Jiao Yuan, Yu Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jiyue Jiang, Zikang Wang, Yihan Ye, Yanruisheng Shao, Yuheng Shan, Jiuming Wang, Xiaodan Fan, Jiao Yuan, Yu Li 0006 |
EMNLP | 10 |
| 2025 | Rethinking Text-based Protein Understanding: Retrieval or LLM?abstractIn recent years, protein-text models have gained significant attention for their potential in protein generation and understanding.Current approaches focus on integrating protein-related knowledge into large language models through continued pretraining and multi-modal alignment, enabling simultaneous comprehension of textual descriptions and protein sequences.Through a thorough analysis of existing model architectures and text-based protein understanding benchmarks, we identify significant data leakage issues present in current benchmarks.Moreover, conventional metrics derived from natural language processing fail to assess the model's performance in this domain accurately.To address these limitations, we reorganize existing datasets and introduce a novel evaluation framework based on biological entities.Motivated by our observation, we propose a retrieval-enhanced method, which significantly outperforms fine-tuned LLMs for protein-totext generation and shows accuracy and efficiency in training-free scenarios. Juntong Wu, Zijing Liu, He Cao, Zishan Shu, Yu Li 0006 |
EMNLP | 9 |
| 2025 | Self-Adjust SoftmaxabstractChuanyang Zheng, Yihang Gao, Guoxuan Chen, Han Shi, Jing Xiong, Xiaozhe Ren, Chao Huang, Zhenguo Li, Yu Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chuanyang Zheng, Yihang Gao, Guoxuan Chen, Xiaozhe Ren, Zhenguo Li, Yu Li 0006 |
EMNLP | 9 |
| 2025 | Visual-Semantic Dual Calibration Network for Zero-Shot Learning
Qingyang Hao, Lei Li 0051, Yu Li 0006, Chun Yuan 0003 |
ICIC (9) | 3 |
| 2025 | Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model ReasoningabstractReasoning capability is pivotal for Large Language Models (LLMs) to solve complex tasks, yet achieving reliable and scalable reasoning remains challenging. While Chain-of-Thought (CoT) prompting has become a mainstream approach, existing methods often suffer from uncontrolled generation, insufficient quality, and limited diversity in reasoning paths.
Recent efforts leverage code to enhance CoT by grounding reasoning in executable steps, but such methods are typically constrained to predefined mathematical problems, hindering scalability and generalizability.
In this work, we propose \texttt{Caco} (Code-Assisted Chain-of-ThOught), a novel framework that automates the synthesis of high-quality, verifiable, and diverse instruction-CoT reasoning data through code-driven augmentation. Unlike prior work, \texttt{Caco} first fine-tunes a code-based CoT generator on existing math and programming solutions in a unified code format, then scales the data generation to a large amount of diverse reasoning traces. Crucially, we introduce automated validation via code execution and rule-based filtering to ensure logical correctness and structural diversity, followed by reverse-engineering filtered outputs into natural language instructions and language CoTs to enrich task adaptability. This closed-loop process enables fully automated, scalable synthesis of reasoning data with guaranteed executability.
Experiments on our created \texttt{Caco}-1.3M dataset demonstrate that \texttt{Caco}-trained models achieve strong competitive performance on mathematical reasoning benchmarks, outperforming existing strong baselines. Further analysis reveals that \texttt{Caco}’s code-anchored verification and instruction diversity contribute to superior generalization across unseen tasks. Our work establishes a paradigm for building self-sustaining, trustworthy reasoning systems without human intervention. Honglin Lin, Qizhi Pei, Zhuoshi Pan, Yu Li 0006, Xin Gao 0001, Juntao Li 0005, Conghui He, Lijun Wu 0003 |
NeurIPS | 4 |
| 2025 | cfDecon: Accurate and Interpretable Methylation-Based Cell Type Deconvolution for Cell-Free DNA
Yimin Fan, Irwin King, Yumei Li 0019, Yu Li 0006 |
RECOMB | 10 |
| 2025 | Artificial intelligence in bioinformatics: a surveyabstractThe widespread adoption of high-throughput sequencing technologies and multi-omics approaches has led to rapid accumulation of genomic, transcriptomic, proteomic, and even single-cell multimodal datasets, resulting in an exponential growth of biological data. The massive scale and inherent complexity of these datasets pose significant challenges for data management, analysis, and interpretation in the field of bioinformatics. Concurrently, artificial intelligence (AI) techniques, particularly deep learning and reinforcement learning, have achieved groundbreaking advances in medical diagnostics, drug discovery, and genomic analyses, providing novel theoretical tools and analytical paradigms for bioinformatics research. AI techniques are now extensively applied to DNA, RNA, and protein sequence prediction and design, 3D structural elucidation, functional annotation, integrative analysis of multi-omics data, and personalized drug design for precision medicine, significantly advancing biological research. This review systematically summarizes recent research progress and representative applications of AI techniques in bioinformatics, specifically discussing suitable scenarios and advantages of traditional machine learning algorithms, deep learning models, and reinforcement learning methods. We highlight AI's transformative impact with quantitative metrics from landmark achievements: accurate near-atomic protein structure prediction (median 0.96 Å on CASP14), robust single-cell modeling (AvgBIO $\approx $ 0.82), high protein design success rates (up to 92%), and sensitive cancer detection (Area Under Curve (AUC) $\approx $ 0.93). Furthermore, the paper provides an in-depth analysis of the latest advancements of AI in specific tasks, including biomedical text mining, multimodal omics integration, and single-cell analyses, while highlighting current challenges such as data noise and sparsity, difficulties in modeling long biological sequences, complexities in multimodal data integration, insufficient model interpretability, and ethical and privacy concerns. Finally, the paper outlines promising future research directions, emphasizing large-scale data mining, cross-domain model generalization, innovations in drug design and personalized medicine, and advocates for establishing an open and collaborative research ecosystem. Jiyue Jiang, Yunke Li, Shiwei Cao, Yuheng Shan, Yuexing Liu, Tianyi Fei, Yule Yu, Yu Li 0006, Jiao Yuan |
Briefings Bioinform. | 9 |
| 2025 | GS-DTI: a graph-structure-aware framework leveraging large language models for drug-target interaction predictionabstractMOTIVATION: Accurate and generalizable prediction of drug-target interactions (DTIs) remains a critical challenge for drug discovery, particularly when addressing underexplored targets and compounds. Recent advances in graph neural networks and large-scale pre-trained models offer new opportunities to capture rich structural and functional features essential for DTI prediction while enhancing the generalization ability. RESULTS: We present GS-DTI, a graph structure-based DTI prediction framework that integrates molecular graph transformers, protein language models, and protein tertiary structure. Our method achieved robust and interpretable DTI predictions. GS-DTI extracts drug features from SMILES-derived molecular graphs using a knowledge-guided pre-trained transformer, while protein features are derived from both sequence and predicted 3D structure for comprehensive representation. A multi-task loss function equipped with contrastive learning is adopted to enhance generalization and functional interpretability. Extensive experiments on the benchmarks and challenging cross-domain settings demonstrate that GS-DTI achieves state-of-the-art performance. Notably, our model improves the MCC by over 10% compared to previous methods in the drug-target pair cold start test. The model can pinpoint the binding pockets of the targets, offering robust interpretability, and case studies show GS-DTI's promising potential in virtual screening for new candidate drugs of BACE1. AVAILABILITY AND IMPLEMENTATION: The GS-DTI source code and processed datasets are available at https://github.com/purvavideha/GSDTI. All experimental data are derived from public sources. Qinze Yu, Jiyue Jiang, Yu Li 0006 |
Bioinform. | 5 |
| 2024 | Task Groupings Regularization: Data-Free Meta-Learning with Heterogeneous Pre-trained ModelsabstractData-Free Meta-Learning (DFML) aims to derive knowledge from a collection of pre-trained models without accessing their original data, enabling the rapid adaptation to new unseen tasks. Current methods often overlook the heterogeneity among pre-trained models, which leads to performance degradation due to task conflicts. In this paper, we empirically and theoretically identify and analyze the model heterogeneity in DFML. We find that model heterogeneity introduces a heterogeneity-homogeneity trade-off, where homogeneous models reduce task conflicts but also increase the overfitting risk. Balancing this trade-off is crucial for learning shared representations across tasks. Based on our findings, we propose Task Groupings Regularization, a novel approach that benefits from model heterogeneity by grouping and aligning conflicting tasks. Specifically, we embed pre-trained models into a task space to compute dissimilarity, and group heterogeneous models together based on this measure. Then, we introduce implicit gradient regularization within each group to mitigate potential conflicts. By encouraging a gradient direction suitable for all tasks, the meta-model captures shared representations that generalize across tasks. Comprehensive experiments showcase the superiority of our approach in multiple benchmarks, effectively tackling the model heterogeneity in challenging multi-domain and multi-architecture scenarios. Yongxian Wei, Li Shen 0008, Zhenyi Wang 0001, Yu Li 0006, Chun Yuan 0003, Dacheng Tao |
ICML | 5 |
| 2024 | HORSE: Hierarchical Representation for Large-Scale Neural Subset SelectionabstractSubset selection tasks, such as anomaly detection and compound selection in AI-assisted drug discovery, are crucial for a wide range of applications. Learning subset-valued functions with neural networks has achieved great success by incorporating permutation invariance symmetry into the architecture. However, existing neural set architectures often struggle to either capture comprehensive information from the superset or address complex interactions within the input. Additionally, they often fail to perform in scenarios where superset sizes surpass available memory capacity. To address these challenges, we introduce the novel concept of the Identity Property, which requires models to integrate information from the originating set, resulting in the development of neural networks that excel at performing effective subset selection from large supersets. Moreover, we present the Hierarchical Representation of Neural Subset Selection (HORSE), an attention-based method that learns complex interactions and retains information from both the input set and the optimal subset supervision signal. Specifically, HORSE enables the partitioning of the input ground set into manageable chunks that can be processed independently and then aggregated, ensuring consistent outcomes across different partitions. Through extensive experimentation, we demonstrate that HORSE significantly enhances neural subset selection performance by capturing more complex information and surpasses state-of-the-art methods in handling large-scale inputs by a margin of up to 20%. Binghui Xie, Yongqiang Chen 0002, Kaiwen Zhou 0001, Yu Li 0006, Wei Meng 0001, James Cheng |
NeurIPS | 5 |
| 2024 | MSA Generation with Seqs2Seqs Pretraining: Advancing Protein Structure PredictionsabstractDeep learning models like AlphaFold2 have revolutionized protein structure prediction, achieving unprecedented accuracy. However, the dependence on robust multiple sequence alignments (MSAs) continues to pose a challenge, especially for proteins that lack a wealth of homologous sequences. To overcome this limitation, we introduce MSA-Generator, a self-supervised generative protein language model. Trained on a sequence-to-sequence task using an automatically constructed dataset, MSA-Generator employs protein-specific attention mechanisms to harness large-scale protein databases, generating virtual MSAs that enrich existing ones and boost prediction accuracy. Our experiments on CASP14 and CASP15 benchmarks reveal significant improvements in LDDT scores, particularly for complex and challenging sequences, enhancing the performance of both AlphaFold2 and RoseTTAFold. The code is released at \url{https://github.com/lezhang7/MSAGen}. Jiayang Chen, Yu Li 0006 |
NeurIPS | 4 |
| 2024 | DAPE: Data-Adaptive Positional Encoding for Length ExtrapolationabstractPositional encoding plays a crucial role in transformers, significantly impact- ing model performance and length generalization. Prior research has introduced absolute positional encoding (APE) and relative positional encoding (RPE) to distinguish token positions in given sequences. However, both APE and RPE remain fixed after model training regardless of input data, limiting their adaptability and flexibility. Hence, we expect that the desired positional encoding should be data-adaptive and can be dynamically adjusted with the given attention. In this paper, we propose a Data-Adaptive Positional Encoding (DAPE) method, which dynamically and semantically adjusts based on input context and learned fixed priors. Experimental validation on real-world datasets (Arxiv, Books3, and CHE) demonstrates that DAPE enhances model performances in terms of trained length and length generalization, where the improvements are statistically significant. The model visualization suggests that our model can keep both local and anti-local information. Finally, we successfully train the model on sequence length 128 and achieve better performance at evaluation sequence length 8192, compared with other static positional encoding methods, revealing the benefit of the adaptive positional encoding method. Chuanyang Zheng, Yihang Gao, Minbin Huang, Jingyao Li 0001, Xiaozhe Ren, Michael Kwok-Po Ng, Zhenguo Li, Yu Li 0006 |
NeurIPS | 11 |
| 2024 | GFETM: Genome Foundation-Based Embedded Topic Model for scATAC-seq Modeling
Yimin Fan, Yu Li 0006, Yue Li 0017 |
RECOMB | 2 |
| 2024 | Progress and opportunities of foundation models in bioinformaticsabstractBioinformatics has undergone a paradigm shift in artificial intelligence (AI), particularly through foundation models (FMs), which address longstanding challenges in bioinformatics such as limited annotated data and data noise. These AI techniques have demonstrated remarkable efficacy across various downstream validation tasks, effectively representing diverse biological entities and heralding a new era in computational biology. The primary goal of this survey is to conduct a general investigation and summary of FMs in bioinformatics, tracing their evolutionary trajectory, current research landscape, and methodological frameworks. Our primary focus is on elucidating the application of FMs to specific biological problems, offering insights to guide the research community in choosing appropriate FMs for tasks like sequence analysis, structure prediction, and function annotation. Each section delves into the intricacies of the targeted challenges, contrasting the architectures and advancements of FMs with conventional methods and showcasing their utility across different biological domains. Further, this review scrutinizes the hurdles and constraints encountered by FMs in biology, including issues of data noise, model interpretability, and potential biases. This analysis provides a theoretical groundwork for understanding the circumstances under which certain FMs may exhibit suboptimal performance. Lastly, we outline prospective pathways and methodologies for the future development of FMs in biological research, facilitating ongoing innovation in the field. This comprehensive examination not only serves as an academic reference but also as a roadmap for forthcoming explorations and applications of FMs in biology. Qing Li 0075, Zhihang Hu, Lei Li 0051, Yimin Fan, Irwin King, Gengjie Jia, Yu Li 0006 |
Briefings Bioinform. | 10 |
| 2024 | ifDEEPre: large protein language-based deep learning enables interpretable and fast predictions of enzyme commission numbersabstractAccurate understanding of the biological functions of enzymes is vital for various tasks in both pathologies and industrial biotechnology. However, the existing methods are usually not fast enough and lack explanations on the prediction results, which severely limits their real-world applications. Following our previous work, DEEPre, we propose a new interpretable and fast version (ifDEEPre) by designing novel self-guided attention and incorporating biological knowledge learned via large protein language models to accurately predict the commission numbers of enzymes and confirm their functions. Novel self-guided attention is designed to optimize the unique contributions of representations, automatically detecting key protein motifs to provide meaningful interpretations. Representations learned from raw protein sequences are strictly screened to improve the running speed of the framework, 50 times faster than DEEPre while requiring 12.89 times smaller storage space. Large language modules are incorporated to learn physical properties from hundreds of millions of proteins, extending biological knowledge of the whole network. Extensive experiments indicate that ifDEEPre outperforms all the current methods, achieving more than 14.22% larger F1-score on the NEW dataset. Furthermore, the trained ifDEEPre models accurately capture multi-level protein biological patterns and infer evolutionary trends of enzymes by taking only raw sequences without label information. Meanwhile, ifDEEPre predicts the evolutionary relationships between different yeast sub-species, which are highly consistent with the ground truth. Case studies indicate that ifDEEPre can detect key amino acid motifs, which have important implications for designing novel enzymes. A web server running ifDEEPre is available at https://proj.cse.cuhk.edu.hk/aihlab/ifdeepre/ to provide convenient services to the public. Meanwhile, ifDEEPre is freely available on GitHub at https://github.com/ml4bio/ifDEEPre/. Qingxiong Tan, Jin Xiao 0002, Jiayang Chen, Zeliang Zhang 0001, Yu Li 0006 |
Briefings Bioinform. | 7 |
| 2024 | scNovel: a scalable deep learning-based network for novel rare cell discovery in single-cell transcriptomicsabstractSingle-cell RNA sequencing has achieved massive success in biological research fields. Discovering novel cell types from single-cell transcriptomics has been demonstrated to be essential in the field of biomedicine, yet is time-consuming and needs prior knowledge. With the unprecedented boom in cell atlases, auto-annotation tools have become more prevalent due to their speed, accuracy and user-friendly features. However, existing tools have mostly focused on general cell-type annotation and have not adequately addressed the challenge of discovering novel rare cell types. In this work, we introduce scNovel, a powerful deep learning-based neural network that specifically focuses on novel rare cell discovery. By testing our model on diverse datasets with different scales, protocols and degrees of imbalance, we demonstrate that scNovel significantly outperforms previous state-of-the-art novel cell detection models, reaching the most AUROC performance(the only one method whose averaged AUROC results are above 94%, up to 16.26% more comparing to the second-best method). We validate scNovel's performance on a million-scale dataset to illustrate the scalability of scNovel further. Applying scNovel on a clinical COVID-19 dataset, three potential novel subtypes of Macrophages are identified, where the COVID-related differential genes are also detected to have consistent expression patterns through deeper analysis. We believe that our proposed pipeline will be an important tool for high-throughput clinical data in a wide range of applications. Chuanyang Zheng, Hongxin Wei, Irwin King, Yu Li 0006 |
Briefings Bioinform. | 7 |
| 2024 | Learning meaningful representation of single-neuron morphology via large-scale pre-trainingabstractSUMMARY: Single-neuron morphology, the study of the structure, form, and shape of a group of specialized cells in the nervous system, is of vital importance to define the type of neurons, assess changes in neuronal development and aging and determine the effects of brain disorders and treatments. Despite the recent surge in the amount of available neuron morphology reconstructions due to advancements in microscopy imaging, existing computational and deep learning methods for modeling neuron morphology have been limited in both scale and accuracy. In this paper, we propose MorphRep, a model for learning meaningful representation of neuron morphology pre-trained with over 250 000 existing neuron morphology data. By encoding the neuron morphology into graph-structured data, using graph transformers for feature encoding and enforcing the consistency between multiple augmented views of neuron morphology, MorphRep achieves the state of the art performance on widely used benchmarking datasets. Meanwhile, MorphRep can accurately characterize the neuron morphology space across neuron morphometrics, fine-grained cell types, brain regions and ages. Furthermore, MorphRep can be applied to distinguish neurons under a wide range of conditions, including genetic perturbation, drug injection, environment change and disease. In summary, MorphRep provides an effective strategy to embed and represent neuron morphology and can be a valuable tool in integrating cell morphology into single-cell multiomics analysis. AVAILABILITY AND IMPLEMENTATION: The codebase has been deposited in https://github.com/YaxuanLi-cn/MorphRep. Yimin Fan, Yaxuan Li 0002, Yunhua Zhong, Lei Li 0051, Yu Li 0006 |
Bioinform. | 6 |
| 2024 | RiboDiffusion: tertiary structure-based RNA inverse folding with generative diffusion modelsabstractMOTIVATION: RNA design shows growing applications in synthetic biology and therapeutics, driven by the crucial role of RNA in various biological processes. A fundamental challenge is to find functional RNA sequences that satisfy given structural constraints, known as the inverse folding problem. Computational approaches have emerged to address this problem based on secondary structures. However, designing RNA sequences directly from 3D structures is still challenging, due to the scarcity of data, the nonunique structure-sequence mapping, and the flexibility of RNA conformation. RESULTS: In this study, we propose RiboDiffusion, a generative diffusion model for RNA inverse folding that can learn the conditional distribution of RNA sequences given 3D backbone structures. Our model consists of a graph neural network-based structure module and a Transformer-based sequence module, which iteratively transforms random sequences into desired sequences. By tuning the sampling weight, our model allows for a trade-off between sequence recovery and diversity to explore more candidates. We split test sets based on RNA clustering with different cut-offs for sequence or structure similarity. Our model outperforms baselines in sequence recovery, with an average relative improvement of 11% for sequence similarity splits and 16% for structure similarity splits. Moreover, RiboDiffusion performs consistently well across various RNA length categories and RNA types. We also apply in silico folding to validate whether the generated sequences can fold into the given 3D RNA backbones. Our method could be a powerful tool for RNA design that explores the vast sequence space and finds novel solutions to 3D structural constraints. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/ml4bio/RiboDiffusion. Ziqian Lin, Dongchen He, Yu Li 0006 |
Bioinform. | 5 |
| 2024 | Meta-Learning Without Data via Unconditional Diffusion ModelsabstractAlthough few-shot learning aims to address data scarcity, it still requires large, annotated datasets for training, which are often unavailable due to cost and privacy concerns. Previous studies have utilized pre-trained diffusion models, either to synthesize auxiliary data besides limited labeled samples, or to employ diffusion models as zero-shot classifiers. However, they are limited to conditional diffusion models needing class prior information (e.g., carefully crafted text prompts) about unseen tasks. To overcome this, we leverage unconditional diffusion models without needs for class information to train a meta-model capable of generalizing to unseen tasks. The framework contains(1)a meta-learning without data approach that uses synthetic data during training; and(2)a diffusion model-based data augmentation to calibrate the distribution shift during testing. During meta-training, we implement aself-taughtclass-learner to gradually capture class concepts, guiding unconditional diffusion models to generate alabeledpseudo dataset. This pseudo dataset is then used to jointly train the class-learner and the meta-model, allowing for iterative refinement and clear differentiation between classes. During meta-testing, we introduce a data augmentation that employs the diffusion models used in meta-training, to narrow the gap between meta-training and meta-testing task distribution. This enables the meta-model trained onsyntheticimages to effectively classifyrealimages in unseen tasks. Comprehensive experiments showcase the superiority and adaptability of our approach in four real-world scenarios. Code available athttps://github.com/WalkerWorldPeace/MLWDUDM. Yongxian Wei, Li Shen 0008, Zhenyi Wang 0001, Lei Li 0051, Yu Li 0006, Chun Yuan 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | AcrNET: predicting anti-CRISPR with deep learningabstractMOTIVATION: As an important group of proteins discovered in phages, anti-CRISPR inhibits the activity of the immune system of bacteria (i.e. CRISPR-Cas), offering promise for gene editing and phage therapy. However, the prediction and discovery of anti-CRISPR are challenging due to their high variability and fast evolution. Existing biological studies rely on known CRISPR and anti-CRISPR pairs, which may not be practical considering the huge number. Computational methods struggle with prediction performance. To address these issues, we propose a novel deep neural network for anti-CRISPR analysis (AcrNET), which achieves significant performance. RESULTS: On both the cross-fold and cross-dataset validation, our method outperforms the state-of-the-art methods. Notably, AcrNET improves the prediction performance by at least 15% regarding the F1 score for the cross-dataset test problem comparing with state-of-art Deep Learning method. Moreover, AcrNET is the first computational method to predict the detailed anti-CRISPR classes, which may help illustrate the anti-CRISPR mechanism. Taking advantage of a Transformer protein language model ESM-1b, which was pre-trained on 250 million protein sequences, AcrNET overcomes the data scarcity problem. Extensive experiments and analysis suggest that the Transformer model feature, evolutionary feature, and local structure feature complement each other, which indicates the critical properties of anti-CRISPR proteins. AlphaFold prediction, further motif analysis, and docking experiments further demonstrate that AcrNET can capture the evolutionarily conserved pattern and the interaction between anti-CRISPR and the target implicitly. AVAILABILITY AND IMPLEMENTATION: Web server: https://proj.cse.cuhk.edu.hk/aihlab/AcrNET/. Training code and pre-trained model are available at. Yumeng Wei, Qingxiong Tan, Licheng Zong, Jiuming Wang, Jiayang Chen, Yu Li 0006 |
Bioinform. | 10 |
| 2023 | Con-AAE: contrastive cycle adversarial autoencoders for single-cell multi-omics alignment and integrationabstractMOTIVATION: We have entered the multi-omics era and can measure cells from different aspects. Hence, we can get a more comprehensive view by integrating or matching data from different spaces corresponding to the same object. However, it is particularly challenging in the single-cell multi-omics scenario because such data are very sparse with extremely high dimensions. Though some techniques can be used to measure scATAC-seq and scRNA-seq simultaneously, the data are usually highly noisy due to the limitations of the experimental environment. RESULTS: To promote single-cell multi-omics research, we overcome the above challenges, proposing a novel framework, contrastive cycle adversarial autoencoders, which can align and integrate single-cell RNA-seq data and single-cell ATAC-seq data. Con-AAE can efficiently map the above data with high sparsity and noise from different spaces to a coordinated subspace, where alignment and integration tasks can be easier. We demonstrate its advantages on several datasets. AVAILABILITY AND IMPLEMENTATION: Zenodo link: https://zenodo.org/badge/latestdoi/368779433. github: https://github.com/kakarotcq/Con-AAE. Zhihang Hu, Tingyang Yu, Yumeng Wei, Juan Shu, Jianzhu Ma, Yu Li 0006 |
Bioinform. | 9 |
| 2022 | Contact-Distil: Boosting Low Homologous Protein Contact Map Prediction by Self-Supervised DistillationabstractAccurate protein contact map prediction (PCMP) is essential for precise protein structure estimation and further biological studies. Recent works achieve significant performance on this task with high quality multiple sequence alignment (MSA). However, the PCMP accuracy drops dramatically while only poor MSA (e.g., absolute MSA count less than 10) is available. Therefore, in this paper, we propose the Contact-Distil to improve the low homologous PCMP accuracy through knowledge distillation on a self-supervised model. Particularly, two pre-trained transformers are exploited to learn the high quality and low quality MSA representation in parallel for the teacher and student model correspondingly. Besides, the co-evolution information is further extracted from pure sequence through a pretrained ESM-1b model, which provides auxiliary knowledge to improve student performance. Extensive experiments show Contact-Distil outperforms previous state-of-the-arts by large margins on CAMEO-L dataset for low homologous PCMP, i.e., around 13.3% and 9.5% improvements against Alphafold2 and MSA Transformer respectively when MSA count less than 10. Qin Wang 0011, Jiayang Chen, Yu Li 0006, Liangzhen Zheng, Sheng Wang 0001, Zhen Li 0026, Shuguang Cui |
AAAI | 4 |
| 2022 | CLMB: Deep Contrastive Learning for Robust Metagenomic Binning
Zhengyuan Jiang, Yu Li 0006 |
RECOMB | 4 |
| 2022 | Self-supervised contrastive learning for integrative single cell RNA-seq data analysisabstractWe present a novel self-supervised Contrastive LEArning framework for single-cell ribonucleic acid (RNA)-sequencing (CLEAR) data representation and the downstream analysis. Compared with current methods, CLEAR overcomes the heterogeneity of the experimental data with a specifically designed representation learning task and thus can handle batch effects and dropout events simultaneously. It achieves superior performance on a broad range of fundamental tasks, including clustering, visualization, dropout correction, batch effect removal, and pseudo-time inference. The proposed method successfully identifies and illustrates inflammatory-related mechanisms in a COVID-19 disease study with 43 695 single cells from peripheral blood mononuclear cells. Wenkai Han, Jiayang Chen, Huawen Zhong, Zhihang Hu, Licheng Zong, Ting-Fung Chan, Irwin King, Xin Gao 0001, Yu Li 0006 |
Briefings Bioinform. | 12 |
| 2022 | Protein-RNA interaction prediction with deep learning: structure mattersabstractProtein-RNA interactions are of vital importance to a variety of cellular activities. Both experimental and computational techniques have been developed to study the interactions. Because of the limitation of the previous database, especially the lack of protein structure data, most of the existing computational methods rely heavily on the sequence data, with only a small portion of the methods utilizing the structural information. Recently, AlphaFold has revolutionized the entire protein and biology field. Foreseeably, the protein-RNA interaction prediction will also be promoted significantly in the upcoming years. In this work, we give a thorough review of this field, surveying both the binding site and binding preference prediction problems and covering the commonly used datasets, features and models. We also point out the potential challenges and opportunities in this field. This survey summarizes the development of the RNA-binding protein-RNA interaction field in the past and foresees its future development in the post-AlphaFold era. Junkang Wei, Licheng Zong, Xin Gao 0001, Yu Li 0006 |
Briefings Bioinform. | 5 |
| 2022 | Deep learning identifies and quantifies recombination hotspot determinantsabstractMOTIVATION: Recombination is one of the essential genetic processes for sexually reproducing organisms, which can happen more frequently in some regions, called recombination hotspots. Although several factors, such as PRDM9 binding motifs, are known to be related to the hotspots, their contributions to the recombination hotspots have not been quantified, and other determinants are yet to be elucidated. Here, we propose a computational method, RHSNet, based on deep learning and signal processing, to identify and quantify the hotspot determinants in a purely data-driven manner, utilizing datasets from various studies, populations, sexes and species. RESULTS: RHSNet can significantly outperform other sequence-based methods on multiple datasets across different species, sexes and studies. In addition to being able to identify hotspot regions and the well-known determinants accurately, more importantly, RHSNet can quantify the determinants that contribute significantly to the recombination hotspot formation in the relation between PRDM9 binding motif, histone modification and GC content. Further cross-sex, cross-population and cross-species studies suggest that the proposed method has the generalization power and potential to identify and quantify the evolutionary determinant motifs. AVAILABILITY AND IMPLEMENTATION: https://github.com/frankchen121212/RHSNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yu Li 0006, Trisevgeni Rapakoulia, Hiroyuki Kuwahara, Kevin Y. Yip, Xin Gao 0001 |
Bioinform. | 1 |
| 2021 | Disease gene prediction with privileged information and heteroscedastic dropoutabstractMOTIVATION: Recently, machine learning models have achieved tremendous success in prioritizing candidate genes for genetic diseases. These models are able to accurately quantify the similarity among disease and genes based on the intuition that similar genes are more likely to be associated with similar diseases. However, the genetic features these methods rely on are often hard to collect due to high experimental cost and various other technical limitations. Existing solutions of this problem significantly increase the risk of overfitting and decrease the generalizability of the models. RESULTS: In this work, we propose a graph neural network (GNN) version of the Learning under Privileged Information paradigm to predict new disease gene associations. Unlike previous gene prioritization approaches, our model does not require the genetic features to be the same at training and test stages. If a genetic feature is hard to measure and therefore missing at the test stage, our model could still efficiently incorporate its information during the training process. To implement this, we develop a Heteroscedastic Gaussian Dropout algorithm, where the dropout probability of the GNN model is determined by another GNN model with a mirrored GNN architecture. To evaluate our method, we compared our method with four state-of-the-art methods on the Online Mendelian Inheritance in Man dataset to prioritize candidate disease genes. Extensive evaluations show that our model could improve the prediction accuracy when all the features are available compared to other methods. More importantly, our model could make very accurate predictions when >90% of the features are missing at the test stage. AVAILABILITY AND IMPLEMENTATION: Our method is realized with Python 3.7 and Pytorch 1.5.0 and method and data are freely available at: https://github.com/juanshu30/Disease-Gene-Prioritization-with-Privileged-Information-and-Heteroscedastic-Dropout. Juan Shu, Yu Li 0006, Sheng Wang 0012, Bowei Xi, Jianzhu Ma |
Bioinform. | 2 |
| 2021 | DeepCellState: An autoencoder-based framework for predicting cell type specific transcriptional states induced by drug treatmentabstractDrug treatment induces cell type specific transcriptional programs, and as the number of combinations of drugs and cell types grows, the cost for exhaustive screens measuring the transcriptional drug response becomes intractable. We developed DeepCellState, a deep learning autoencoder-based framework, for predicting the induced transcriptional state in a cell type after drug treatment, based on the drug response in another cell type. Training the method on a large collection of transcriptional drug perturbation profiles, prediction accuracy improves significantly over baseline and alternative deep learning approaches when applying the method to two cell types, with improved accuracy when generalizing the framework to additional cell types. Treatments with drugs or whole drug families not seen during training are predicted with similar accuracy, and the same framework can be used for predicting the results from other interventions, such as gene knock-downs. Finally, analysis of the trained model shows that the internal representation is able to learn regulatory relationships between genes in a fully data-driven manner. Ramzan Umarov, Yu Li 0006, Erik Arner |
PLoS Comput. Biol. | 2 |
| 2021 | ReFeaFi: Genome-wide prediction of regulatory elements driving transcription initiationabstractRegulatory elements control gene expression through transcription initiation (promoters) and by enhancing transcription at distant regions (enhancers). Accurate identification of regulatory elements is fundamental for annotating genomes and understanding gene expression patterns. While there are many attempts to develop computational promoter and enhancer identification methods, reliable tools to analyze long genomic sequences are still lacking. Prediction methods often perform poorly on the genome-wide scale because the number of negatives is much higher than that in the training sets. To address this issue, we propose a dynamic negative set updating scheme with a two-model approach, using one model for scanning the genome and the other one for testing candidate positions. The developed method achieves good genome-level performance and maintains robust performance when applied to other vertebrate species, without re-training. Moreover, the unannotated predicted regulatory regions made on the human genome are enriched for disease-associated variants, suggesting them to be potentially true regulatory elements rather than false positives. We validated high scoring "false positive" predictions using reporter assay and all tested candidates were successfully validated, demonstrating the ability of our method to discover novel human regulatory regions. Ramzan Umarov, Yu Li 0006, Takahiro Arakawa, Satoshi Takizawa, Xin Gao 0001, Erik Arner |
PLoS Comput. Biol. | 2 |
| 2020 | RNA Secondary Structure Prediction By Learning Unrolled Algorithms
Xinshi Chen, Yu Li 0006, Ramzan Umarov, Xin Gao 0001 |
ICLR | 2 |
| 2020 | Learning To Stop While Learning To PredictabstractThere is a recent surge of interest in designing deep architectures based on the update steps in traditional algorithms, or learning neural networks to improve and replace traditional algorithms. While traditional algorithms have certain stopping criteria for outputting results at different iterations, many algorithm-inspired deep models are restricted to a “fixed-depth” for all inputs. Similar to algorithms, the optimal depth of a deep architecture may be different for different input instances, either to avoid “over-thinking”, or because we want to compute less for operations converged already. In this paper, we tackle this varying depth problem using a steerable architecture, where a feed-forward deep model and a variational stopping policy are learned together to sequentially determine the optimal number of layers for each input instance. Training such architecture is very challenging. We provide a variational Bayes perspective and design a novel and effective training procedure which decomposes the task into an oracle model learning stage and an imitation stage. Experimentally, we show that the learned deep model along with the stopping policy improves the performances on a diverse set of tasks, including learning sparse recovery, few-shot meta learning, and computer vision tasks. Xinshi Chen, Hanjun Dai, Yu Li 0006, Xin Gao 0001 |
ICML | 3 |
| 2020 | DeepSimulator1.5: a more powerful, quicker and lighter simulator for Nanopore sequencingabstractMOTIVATION: Nanopore sequencing is one of the leading third-generation sequencing technologies. A number of computational tools have been developed to facilitate the processing and analysis of the Nanopore data. Previously, we have developed DeepSimulator1.0 (DS1.0), which is the first simulator for Nanopore sequencing to produce both the raw electrical signals and the reads. However, although DS1.0 can produce high-quality reads, for some sequences, the divergence between the simulated raw signals and the real signals can be large. Furthermore, the Nanopore sequencing technology has evolved greatly since DS1.0 was released. It is thus necessary to update DS1.0 to accommodate those changes. RESULTS: We propose DeepSimulator1.5 (DS1.5), all three modules of which have been updated substantially from DS1.0. As for the sequence generator, we updated the sample read length distribution to reflect the newest real reads' features. In terms of the signal generator, which is the core of DeepSimulator, we added one more pore model, the context-independent pore model, which is much faster than the previous context-dependent one. Furthermore, to make the generated signals more similar to the real ones, we added a low-pass filter to post-process the pore model signals. Regarding the basecaller, we added the support for the newest official basecaller, Guppy, which can support both GPU and CPU. In addition, multiple optimizations, related to multiprocessing control, memory and storage management, have been implemented to make DS1.5 a much more amenable and lighter simulator than DS1.0. AVAILABILITY AND IMPLEMENTATION: The main program and the data are available at https://github.com/lykaust15/DeepSimulator. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yu Li 0006, Sheng Wang 0001, Chongwei Bi, Zhaowen Qiu, Mo Li 0005, Xin Gao 0001 |
Bioinform. | 1 |
| 2019 | Linear Kernel Tests via Empirical Likelihood for High-Dimensional DataabstractWe propose a framework for analyzing and comparing distributions without imposing any parametric assumptions via empirical likelihood methods. Our framework is used to study two fundamental statistical test problems: the two-sample test and the goodness-of-fit test. For the two-sample test, we need to determine whether two groups of samples are from different distributions; for the goodness-of-fit test, we examine how likely it is that a set of samples is generated from a known target distribution. Specifically, we propose empirical likelihood ratio (ELR) statistics for the two-sample test and the goodness-of-fit test, both of which are of linear time complexity and show higher power (i.e., the probability of correctly rejecting the null hypothesis) than the existing linear statistics for high-dimensional data. We prove the nonparametric Wilks’ theorems for the ELR statistics, which illustrate that the limiting distributions of the proposed ELR statistics are chi-square distributions. With these limiting distributions, we can avoid bootstraps or simulations to determine the threshold for rejecting the null hypothesis, which makes the ELR statistics more efficient than the recently proposed linear statistic, finite set Stein discrepancy (FSSD). We also prove the consistency of the ELR statistics, which guarantees that the test power goes to 1 as the number of samples goes to infinity. In addition, we experimentally demonstrate and theoretically analyze that FSSD has poor performance or even fails to test for high-dimensional data. Finally, we conduct a series of experiments to evaluate the performance of our ELR statistics as compared to state-of-the-art linear statistics. Lizhong Ding 0001, Yu Li 0006, Shizhong Liao, Yong Liu 0018, Peng Yang 0010, Ling Shao 0001, Xin Gao 0001 |
AAAI | 3 |
| 2019 | Approximate Kernel Selection with Strong Approximate ConsistencyabstractKernel selection is fundamental to the generalization performance of kernel-based learning algorithms. Approximate kernel selection is an efficient kernel selection approach that exploits the convergence property of the kernel selection criteria and the computational virtue of kernel matrix approximation. The convergence property is measured by the notion of approximate consistency. For the existing Nyström approximations, whose sampling distributions are independent of the specific learning task at hand, it is difficult to establish the strong approximate consistency. They mainly focus on the quality of the low-rank matrix approximation, rather than the performance of the kernel selection criterion used in conjunction with the approximate matrix. In this paper, we propose a novel Nyström approximate kernel selection algorithm by customizing a criterion-driven adaptive sampling distribution for the Nyström approximation, which adaptively reduces the error between the approximate and accurate criteria. We theoretically derive the strong approximate consistency of the proposed Nyström approximate kernel selection algorithm. Finally, we empirically evaluate the approximate consistency of our algorithm as compared to state-of-the-art methods. Lizhong Ding 0001, Yong Liu 0018, Shizhong Liao, Yu Li 0006, Peng Yang 0010, Yijie Pan, Ling Shao 0001, Xin Gao 0001 |
AAAI | 4 |
| 2019 | Two Generator Game: Learning to Sample via Linear Goodness-of-Fit TestabstractLearning the probability distribution of high-dimensional data is a challenging problem. To solve this problem, we formulate a deep energy adversarial network (DEAN), which casts the energy model learned from real data into an optimization of a goodness-of-fit (GOF) test statistic. DEAN can be interpreted as a GOF game between two generative networks, where one explicit generative network learns an energy-based distribution that fits the real data, and the other implicit generative network is trained by minimizing a GOF test statistic between the energy-based distribution and the generated data, such that the underlying distribution of the generated data is close to the energy-based distribution. We design a two-level alternative optimization procedure to train the explicit and implicit generative networks, such that the hyper-parameters can also be automatically learned. Experimental results show that DEAN achieves high quality generations compared to the state-of-the-art approaches. Lizhong Ding 0001, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Yong Liu 0018, Yu Li 0006, Ling Shao 0001 |
NeurIPS | 6 |
| 2019 | AuTom-dualx: a toolkit for fully automatic fiducial marker-based alignment of dual-axis tilt series with simultaneous reconstructionabstractMotivation: Dual-axis electron tomography is an important 3 D macro-molecular structure reconstruction technology, which can reduce artifacts and suppress the effect of missing wedge. However, the fully automatic data process for dual-axis electron tomography still remains a challenge due to three difficulties: (i) how to track the mass of fiducial markers automatically; (ii) how to integrate the information from the two different tilt series; and (iii) how to cope with the inconsistency between the two different tilt series. Results: Here we develop a toolkit for fully automatic alignment of dual-axis electron tomography, with a simultaneous reconstruction procedure. The proposed toolkit and its workflow carries out the following solutions: (i) fully automatic detection and tracking of fiducial markers under large-field datasets; (ii) automatic combination of two different tilt series and global calibration of projection parameters; and (iii) inconsistency correction based on distortion correction parameters and the consequently simultaneous reconstruction. With all of these features, the presented toolkit can achieve accurate alignment and reconstruction simultaneously and conveniently under a single global coordinate system. Availability and implementation: The toolkit AuTom-dualx (alignment module dualxmauto and reconstruction module volrec_mltm) are accessible for general application at http://ear.ict.ac.cn, and the key source code is freely available under request. Supplementary information: Supplementary data are available at Bioinformatics online. Renmin Han, Albert F. Lawrence, Peng Yang 0010, Yu Li 0006, Sheng Wang 0001, Zhiyong Liu 0002, Xin Gao 0001, Fa Zhang 0001 |
Bioinform. | 6 |
| 2019 | Promoter analysis and prediction in the human genome using sequence-based deep learning modelsabstractMOTIVATION: Computational identification of promoters is notoriously difficult as human genes often have unique promoter sequences that provide regulation of transcription and interaction with transcription initiation complex. While there are many attempts to develop computational promoter identification methods, we have no reliable tool to analyze long genomic sequences. RESULTS: In this work, we further develop our deep learning approach that was relatively successful to discriminate short promoter and non-promoter sequences. Instead of focusing on the classification accuracy, in this work we predict the exact positions of the transcription start site inside the genomic sequences testing every possible location. We studied human promoters to find effective regions for discrimination and built corresponding deep learning models. These models use adaptively constructed negative set, which iteratively improves the model's discriminative ability. Our method significantly outperforms the previously developed promoter prediction programs by considerably reducing the number of false-positive predictions. We have achieved error-per-1000-bp rate of 0.02 and have 0.31 errors per correct prediction, which is significantly better than the results of other human promoter predictors. AVAILABILITY AND IMPLEMENTATION: The developed method is available as a web server at http://www.cbrc.kaust.edu.sa/PromID/. Ramzan Umarov, Hiroyuki Kuwahara, Yu Li 0006, Xin Gao 0001, Victor V. Solovyev |
Bioinform. | 3 |
| 2019 | PredMP: a web server for de novo prediction and visualization of membrane proteinsabstractMOTIVATION: PredMP is the first web service, to our knowledge, that aims at de novo prediction of the membrane protein (MP) 3D structure followed by the embedding of the MP into the lipid bilayer for visualization. Our approach is based on a high-throughput Deep Transfer Learning (DTL) method that first predicts MP contacts by learning from non-MPs and then predicts the 3D model of the MP using the predicted contacts as distance restraints. This algorithm is derived from our previous Deep Learning (DL) method originally developed for soluble protein contact prediction, which has been officially ranked No. 1 in CASP12. The DTL framework in our approach overcomes the challenge that there are only a limited number of solved MP structures for training the deep learning model. There are three modules in the PredMP server: (i) The DTL framework followed by the contact-assisted folding protocol has already been implemented in RaptorX-Contact, which serves as the key module for 3D model generation; (ii) The 1D annotation module, implemented in RaptorX-Property, is used to predict the secondary structure and disordered regions; and (iii) the visualization module to display the predicted MPs embedded in the lipid bilayer guided by the predicted transmembrane topology. RESULTS: Tested on 510 non-redundant MPs, our server predicts correct folds for ∼290 MPs, which significantly outperforms existing methods. Tested on a blind and live benchmark CAMEO from September 2016 to January 2018, PredMP can successfully model all 10 MPs belonging to the hard category. AVAILABILITY AND IMPLEMENTATION: PredMP is freely accessed on the web at http://www.predmp.com. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sheng Wang 0001, Shiyang Fei, Zongan Wang, Yu Li 0006, Jinbo Xu, Feng Zhao 0004, Xin Gao 0001 |
Bioinform. | 4 |
| 2019 | DeeReCT-PolyA: a robust and generic deep learning method for PAS identificationabstractMOTIVATION: Polyadenylation is a critical step for gene expression regulation during the maturation of mRNA. An accurate and robust method for poly(A) signals (PASs) identification is not only desired for the purpose of better transcripts' end annotation, but can also help us gain a deeper insight of the underlying regulatory mechanism. Although many methods have been proposed for PAS recognition, most of them are PAS motif- and human-specific, which leads to high risks of overfitting, low generalization power, and inability to reveal the connections between the underlying mechanisms of different mammals. RESULTS: In this work, we propose a robust, PAS motif agnostic, and highly interpretable and transferrable deep learning model for accurate PAS recognition, which requires no prior knowledge or human-designed features. We show that our single model trained over all human PAS motifs not only outperforms the state-of-the-art methods trained on specific motifs, but can also be generalized well to two mouse datasets. Moreover, we further increase the prediction accuracy by transferring the deep learning model trained on the data of one species to the data of a different species. Several novel underlying poly(A) patterns are revealed through the visualization of important oligomers and positions in our trained models. Finally, we interpret the deep learning models by converting the convolutional filters into sequence logos and quantitatively compare the sequence logos between human and mouse datasets. AVAILABILITY AND IMPLEMENTATION: https://github.com/likesum/DeeReCT-PolyA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhihao Xia, Yu Li 0006, Bin Zhang 0042, Zhongxiao Li, Yuhui Hu, Wei Chen 0029, Xin Gao 0001 |
Bioinform. | 2 |
| 2018 | An accurate and rapid continuous wavelet dynamic time warping algorithm for end-to-end mapping in ultra-long nanopore sequencingabstractMotivation: Long-reads, point-of-care and polymerase chain reaction-free are the promises brought by nanopore sequencing. Among various steps in nanopore data analysis, the end-to-end mapping between the raw electrical current signal sequence and the reference expected signal sequence serves as the key building block to signal labeling, and the following signal visualization, variant identification and methylation detection. One of the classic algorithms to solve the signal mapping problem is the dynamic time warping (DTW). However, the ultra-long nanopore sequencing and an order of magnitude difference in the sampling speed complexify the scenario and make the classical DTW infeasible to solve the problem. Results: Here, we propose a novel multi-level DTW algorithm, continuous wavelet DTW (cwDTW), based on continuous wavelet transforms with different scales of the two signal sequences. Our algorithm starts from low-resolution wavelet transforms of the two sequences, such that the transformed sequences are short and have similar sampling rates. Then the peaks and nadirs of the transformed sequences are extracted to form feature sequences with similar lengths, which can be easily mapped by the original DTW. Our algorithm then recursively projects the warping path from a lower-resolution level to a higher-resolution one by building a context-dependent boundary and enabling a constrained search for the warping path in the latter. Comprehensive experiments on two real nanopore datasets on human and on Pandoraea pnomenusa demonstrate the efficiency and effectiveness of the proposed algorithm. In particular, cwDTW can gain remarkable acceleration with tiny loss of the alignment accuracy. On the real nanopore datasets, cwDTW can finish an alignment task in few seconds, which is about 3000 times faster than the original DTW. By successfully applying cwDTW on the tasks of signal labeling and ultra-long sequence comparison, we further demonstrate the power and applicability of cwDTW. Availability and implementation: Our program is available at https://github.com/realbigws/cwDTW. Supplementary information: Supplementary data are available at Bioinformatics online. Renmin Han, Yu Li 0006, Xin Gao 0001, Sheng Wang 0001 |
Bioinform. | 2 |
| 2018 | DeepSimulator: a deep simulator for Nanopore sequencingabstractMotivation: Oxford Nanopore sequencing is a rapidly developed sequencing technology in recent years. To keep pace with the explosion of the downstream data analytical tools, a versatile Nanopore sequencing simulator is needed to complement the experimental data as well as to benchmark those newly developed tools. However, all the currently available simulators are based on simple statistics of the produced reads, which have difficulty in capturing the complex nature of the Nanopore sequencing procedure, the main task of which is the generation of raw electrical current signals. Results: Here we propose a deep learning based simulator, DeepSimulator, to mimic the entire pipeline of Nanopore sequencing. Starting from a given reference genome or assembled contigs, we simulate the electrical current signals by a context-dependent deep learning model, followed by a base-calling procedure to yield simulated reads. This workflow mimics the sequencing procedure more naturally. The thorough experiments performed across four species show that the signals generated by our context-dependent model are more similar to the experimentally obtained signals than the ones generated by the official context-independent pore model. In terms of the simulated reads, we provide a parameter interface to users so that they can obtain the reads with different accuracies ranging from 83 to 97%. The reads generated by the default parameter have almost the same properties as the real data. Two case studies demonstrate the application of DeepSimulator to benefit the development of tools in de novo assembly and in low coverage SNP detection. Availability and implementation: The software can be accessed freely at: https://github.com/lykaust15/DeepSimulator. Supplementary information: Supplementary data are available at Bioinformatics online. Yu Li 0006, Renmin Han, Chongwei Bi, Mo Li 0005, Sheng Wang 0001, Xin Gao 0001 |
Bioinform. | 1 |
| 2018 | DEEPre: sequence-based enzyme EC number prediction by deep learningabstractMotivation: Annotation of enzyme function has a broad range of applications, such as metagenomics, industrial biotechnology, and diagnosis of enzyme deficiency-caused diseases. However, the time and resource required make it prohibitively expensive to experimentally determine the function of every enzyme. Therefore, computational enzyme function prediction has become increasingly important. In this paper, we develop such an approach, determining the enzyme function by predicting the Enzyme Commission number. Results: We propose an end-to-end feature selection and classification model training approach, as well as an automatic and robust feature dimensionality uniformization method, DEEPre, in the field of enzyme function prediction. Instead of extracting manually crafted features from enzyme sequences, our model takes the raw sequence encoding as inputs, extracting convolutional and sequential features from the raw encoding based on the classification result to directly improve the prediction performance. The thorough cross-fold validation experiments conducted on two large-scale datasets show that DEEPre improves the prediction performance over the previous state-of-the-art methods. In addition, our server outperforms five other servers in determining the main class of enzymes on a separate low-homology dataset. Two case studies demonstrate DEEPre's ability to capture the functional difference of enzyme isoforms. Availability and implementation: The server could be accessed freely at http://www.cbrc.kaust.edu.sa/DEEPre. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Yu Li 0006, Sheng Wang 0001, Ramzan Umarov, Bingqing Xie, Ming Fan 0003, Lihua Li 0002, Xin Gao 0001 |
Bioinform. | 1 |
| 2018 | DLBI: deep learning guided Bayesian inference for structure reconstruction of super-resolution fluorescence microscopyabstractMotivation: Super-resolution fluorescence microscopy with a resolution beyond the diffraction limit of light, has become an indispensable tool to directly visualize biological structures in living cells at a nanometer-scale resolution. Despite advances in high-density super-resolution fluorescent techniques, existing methods still have bottlenecks, including extremely long execution time, artificial thinning and thickening of structures, and lack of ability to capture latent structures. Results: Here, we propose a novel deep learning guided Bayesian inference (DLBI) approach, for the time-series analysis of high-density fluorescent images. Our method combines the strength of deep learning and statistical inference, where deep learning captures the underlying distribution of the fluorophores that are consistent with the observed time-series fluorescent images by exploring local features and correlation along time-axis, and statistical inference further refines the ultrastructure extracted by deep learning and endues physical meaning to the final image. In particular, our method contains three main components. The first one is a simulator that takes a high-resolution image as the input, and simulates time-series low-resolution fluorescent images based on experimentally calibrated parameters, which provides supervised training data to the deep learning model. The second one is a multi-scale deep learning module to capture both spatial information in each input low-resolution image as well as temporal information among the time-series images. And the third one is a Bayesian inference module that takes the image from the deep learning module as the initial localization of fluorophores and removes artifacts by statistical inference. Comprehensive experimental results on both real and simulated datasets demonstrate that our method provides more accurate and realistic local patch and large-field reconstruction than the state-of-the-art method, the 3B analysis, while our method is more than two orders of magnitude faster. Availability and implementation: The main program is available at https://github.com/lykaust15/DLBI. Supplementary information: Supplementary data are available at Bioinformatics online. Yu Li 0006, Fa Zhang 0001, Pingyong Xu, Mingshu Zhang, Ming Fan 0003, Lihua Li 0002, Xin Gao 0001, Renmin Han |
Bioinform. | 1 |
| 2017 | Sequence2Vec: a novel embedding approach for modeling transcription factor binding affinity landscapeabstractMOTIVATION: An accurate characterization of transcription factor (TF)-DNA affinity landscape is crucial to a quantitative understanding of the molecular mechanisms underpinning endogenous gene regulation. While recent advances in biotechnology have brought the opportunity for building binding affinity prediction methods, the accurate characterization of TF-DNA binding affinity landscape still remains a challenging problem. RESULTS: Here we propose a novel sequence embedding approach for modeling the transcription factor binding affinity landscape. Our method represents DNA binding sequences as a hidden Markov model which captures both position specific information and long-range dependency in the sequence. A cornerstone of our method is a novel message passing-like embedding algorithm, called Sequence2Vec, which maps these hidden Markov models into a common nonlinear feature space and uses these embedded features to build a predictive model. Our method is a novel combination of the strength of probabilistic graphical models, feature space embedding and deep learning. We conducted comprehensive experiments on over 90 large-scale TF-DNA datasets which were measured by different high-throughput experimental technologies. Sequence2Vec outperforms alternative machine learning methods as well as the state-of-the-art binding affinity prediction methods. AVAILABILITY AND IMPLEMENTATION: Our program is freely available at https://github.com/ramzan1990/sequence2vec. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hanjun Dai, Ramzan Umarov, Hiroyuki Kuwahara, Yu Li 0006, Xin Gao 0001 |
Bioinform. | 4 |
| 2010 | Dynamic memory paravirtualization transparent to guest OS
Xiaolin Wang 0001, Yifeng Sun, Yingwei Luo, Zhenlin Wang 0003, Yu Li 0006, Haogang Chen 0002, Xiaoming Li 0001 |
Sci. China Inf. Sci. | 5 |
| 2008 | Semantic knowledge facilities for a web-based recipe database system supporting personalizationabstractAbstract The recent explosive proliferation of interesting and useful data over the Web such as various recipes, while providing people with readily available information, brings out a challenging issue on how to manage such non‐conventional data effectively. To respond to the challenge, we have been developing a Web‐based recipe database system calledDish_Masterto manage recipes in a novel way, which not only covers the static recipe attributes but also elucidates the dynamic cooking behaviors. In this paper, we present several semantic knowledge facilities devised inDish_Master, including a set of semantic modeling and knowledge constructs to effectively represent recipe data, rules and constraints, and user profile aspects. With such a rich set of semantic knowledge facilities,Dish_Masterlays down a solid foundation of providing users with personalized services such as adaptation and recommendation. Users can benefit from the system's real‐time consultation and automatic summarization of cuisine knowledge. The usefulness and elegance ofDish_Masterare demonstrated through an experimental prototype system. Copyright © 2007 John Wiley & Sons, Ltd. Liping Wang 0002, Qing Li 0001, Guozhu Dong, Yu Li 0006 |
Concurr. Comput. Pract. Exp. | 4 |
| 2006 | RecipeCrawler: Collecting Recipe Data from WWW Incrementally
Yu Li 0006, Xiaofeng Meng 0001, Liping Wang 0002, Qing Li 0001 |
WAIM | 1 |
| 2006 | Hybrid Method for Automated News Content Extraction from the Web
Yu Li 0006, Xiaofeng Meng 0001, Qing Li 0001, Liping Wang 0002 |
WISE | 1 |