EDBT 2026 Demo / reviewers in the wild / expert
Mingze Yin
dblp:371/3018
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0009-6595-9849ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
3 papers |
Generative modeling · 68% Language models and text generation · 32% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein function prediction |
1.7 | 2 | 2025 | Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs · KDD (2) 2025 ProtCLIP: Function-Informed Protein Multi-Modal Learning · AAAI 2025 |
Bioinformatics and computational biology
protein design |
1.6 | 2 | 2025 | Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer · AAAI 2025 Bridge-IF: Learning Inverse Protein Folding with Markov Bridges · NeurIPS 2024 |
Machine learning › Generative modeling
generative flow networks |
0.9 | 1 | 2025 | Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer · AAAI 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs · KDD (2) 2025 |
Bioinformatics and computational biology › protein design
antibody design |
0.9 | 1 | 2025 | Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer · AAAI 2025 |
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein representation learning |
0.9 | 1 | 2025 | ProtCLIP: Function-Informed Protein Multi-Modal Learning · AAAI 2025 |
Machine learning › Generative modeling › diffusion model
diffusion bridge |
0.8 | 1 | 2024 | Bridge-IF: Learning Inverse Protein Folding with Markov Bridges · NeurIPS 2024 |
Machine learning › Generative modeling › generative model › continuous-time generative model
markov bridge |
0.8 | 1 | 2024 | Bridge-IF: Learning Inverse Protein Folding with Markov Bridges · NeurIPS 2024 |
Bioinformatics and computational biology › protein design
inverse protein folding |
0.8 | 1 | 2024 | Bridge-IF: Learning Inverse Protein Folding with Markov Bridges · NeurIPS 2024 |
Natural language and speech › Language models and text generation › neural language model
protein language model |
0.3 | 1 | 2025 | Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs · KDD (2) 2025 |
Methods — techniques the papers use, named apart from their topics
protein language model · 3.3contrastive learning · 2.6structure denoising · 1.7products of experts · 1.7potts model · 1.7mixture of experts · 1.7contrastive divergence · 1.7structure encoder · 1.5multimodal pretraining · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody DesignerabstractAntibodies defend our health by binding to antigens with high specificity and potentiality, primarily relying on the Complementarity-Determining Region (CDR). Yet, current experimental methods of discovering new antibody CDRs are heavily time-consuming. Computational design could alleviate this burden; especially, protein language models have proven quite beneficial in many recent studies. However, most existing models solely focus on antibody potentiality and struggle to encapsulate the diverse range of plausible CDR candidates, limiting their effectiveness in real-world scenarios as binding is only one factor in the multitude of drug-forming criteria. In this paper, we introduce PG-AbD, a framework uniting Generative Flow Networks (GFlowNets) and pretrained Protein Language Models (PLMs) to successfully generate highly potent, diverse and novel antibody candidates. We innovatively construct a Products of Experts (PoE) composed by the global-distribution-modeling PLM and the local-distribution-modeling Potts Model to serve as the reward function of GFlowNet. The joint training paradigm is introduced, where PoE is trained by contrastive divergence with the negative samples generated by GFlowNet, and then guides GFlowNet to sample diverse antibody candidates. We evaluate PG-AbD on extensive antibody design benchmarks. It significantly outperforms existing methods in diversity (13.5% on RabDab, 31.1% on SabDab) while maintaining optimal potential and novelty. Generated antibodies are also found to form stable, regular 3D structures with their corresponding antigens, demonstrating the great potential of PG-AbD to accelerate real-world antibody discovery. Mingze Yin, Hanjing Zhou, Yiheng Zhu 0002, Jialu Wu, Wei Wu 0045, Kun Fu 0002, Zheng Wang 0027, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001 |
AAAI | 1 |
| 2025 | ProtCLIP: Function-Informed Protein Multi-Modal LearningabstractMulti-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model. Hanjing Zhou, Mingze Yin, Wei Wu 0045, Kun Fu 0002, Jintai Chen, Jian Wu 0001, Zheng Wang 0027 |
AAAI | 2 |
| 2025 | Group-On: Boosting One-Shot Segmentation with Supportive QueryabstractOne-shot semantic segmentation aims to segment query images given only ONE annotated support image of the same class. This task is challenging because target objects in the support and query images can be largely different in appearance and pose (i.e., intra-class variation). Prior works suggested that incorporating more annotated support images in few-shot settings boosts performances but increases costs due to additional manual labeling. In this paper, we propose a novel and effective approach for ONE-shot semantic segmentation, called Group-On, which packs multiple query images in batches for the benefit of mutual knowledge support within the same category. Specifically, after coarse segmentation masks of the batch of queries are predicted, query-mask pairs act as pseudo support data to enhance mask predictions mutually. To effectively steer such process, we construct an innovative MoME module, where a flexible number of mask experts are guided by a scene-driven router and work together to make comprehensive decisions, fully promoting mutual benefits of queries. Comprehensive experiments on three standard benchmarks show that, in the ONE-shot setting, Group-On significantly outperforms previous works by considerable margins. With only one annotated support image, Group-On can be even competitive with the counterparts using 5 annotated images. Hanjing Zhou, Mingze Yin, Danny Ziyi Chen, Jian Wu 0001, Jintai Chen |
ICME | 2 |
| 2025 | Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMsabstractProteins, as essential biomolecules, play a central role in biological processes, including metabolic reactions and DNA replication. Accurate prediction of their properties and functions is crucial in biological applications. Recent development of protein language models (pLMs) with supervised fine tuning provides a promising solution to this problem. However, the fine-tuned model is tailored for particular downstream prediction task, and achieving general-purpose protein understanding remains a challenge. In this paper, we introduce Structure-Enhanced Protein Instruction Tuning (SEPIT) framework to bridge this gap. Our approach incorporates a novel structure-aware module into pLMs to enrich their structural knowledge, and subsequently integrates these enhanced pLMs with large language models (LLMs) to advance protein understanding. In this framework, we propose a novel instruction tuning pipeline. First, we warm up the enhanced pLMs using contrastive learning and structure denoising. Then, caption-based instructions are used to establish a basic understanding of proteins. Finally, we refine this understanding by employing a mixture of experts (MoEs) to capture more complex properties and functional information with the same number of activated parameters. Moreover, we construct the largest and most comprehensive protein instruction dataset to date, which allows us to train and evaluate the general-purpose protein understanding model. Extensive experiments on both open-ended generation and closed-set answer tasks demonstrate the superior performance of SEPIT over both closed-source general LLMs and open-source LLMs trained with protein knowledge. Wei Wu 0045, Chao Wang 0086, Liyi Chen 0001, Mingze Yin, Yiheng Zhu 0002, Kun Fu 0002, Jieping Ye, Hui Xiong 0001, Zheng Wang 0027 |
KDD (2) | 4 |
| 2024 | Bridge-IF: Learning Inverse Protein Folding with Markov BridgesabstractInverse protein folding is a fundamental task in computational protein design, which aims to design protein sequences that fold into the desired backbone structures. While the development of machine learning algorithms for this task has seen significant success, the prevailing approaches, which predominantly employ a discriminative formulation, frequently encounter the error accumulation issue and often fail to capture the extensive variety of plausible sequences. To fill these gaps, we propose Bridge-IF, a generative diffusion bridge model for inverse folding, which is designed to learn the probabilistic dependency between the distributions of backbone structures and protein sequences. Specifically, we harness an expressive structure encoder to propose a discrete, informative prior derived from structures, and establish a Markov bridge to connect this prior with native sequences. During the inference stage, Bridge-IF progressively refines the prior sequence, culminating in a more plausible design. Moreover, we introduce a reparameterization perspective on Markov bridge models, from which we derive a simplified loss function that facilitates more effective training. We also modulate protein language models (PLMs) with structural conditions to precisely approximate the Markov bridge process, thereby significantly enhancing generation performance while maintaining parameter-efficient training. Extensive experiments on well-established benchmarks demonstrate that Bridge-IF predominantly surpasses existing baselines in sequence recovery and excels in the design of plausible proteins with high foldability. The code is available at https://github.com/violet-sto/Bridge-IF. Jialu Wu, Qiuyi Li, Jiahuan Yan, Mingze Yin, Jieping Ye |
NeurIPS | 5 |