Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yiheng Zhu 0002

dblp:232/1993-2 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-8020-9979ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 36% Language models and text generation · 17% Segmentation and scene understanding · 17%
Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 90% Medical and health informatics · 10%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
generative flow networks
1.522025
Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer · AAAI 2025
Sample-efficient Multi-objective Molecular Optimization with GFlowNets · NeurIPS 2023
Natural language and speech › Language models and text generation
instruction tuning
0.912025
Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs · KDD (2) 2025
Bioinformatics and computational biology › protein design
antibody design
0.912025
Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer · AAAI 2025
Bioinformatics and computational biology
protein design
0.912025
Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer · AAAI 2025
Bioinformatics and computational biology
protein function prediction
0.912025
Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs · KDD (2) 2025
Machine learning › Kernel, tree and ensemble methods
tabular prediction
0.812024
Making Pre-trained Language Models Great on Tabular Prediction · ICLR 2024
Computer vision › Segmentation and scene understanding
annotation-efficient segmentation
0.712023
GCL: Gradient-Guided Contrastive Learning for Medical Image Segmentation with Multi-Perspective Meta Labels · ACM Multimedia 2023
Machine learning › Graph learning
graph generation
0.712023
MolHF: A Hierarchical Normalizing Flow for Molecular Graph Generation · IJCAI 2023
Computer vision › Segmentation and scene understanding
medical image segmentation
0.712023
GCL: Gradient-Guided Contrastive Learning for Medical Image Segmentation with Multi-Perspective Meta Labels · ACM Multimedia 2023
Machine learning › Generative modeling › molecular generation
molecular graph generation
0.712023
MolHF: A Hierarchical Normalizing Flow for Molecular Graph Generation · IJCAI 2023
Machine learning › Generative modeling
normalizing flow
0.712023
MolHF: A Hierarchical Normalizing Flow for Molecular Graph Generation · IJCAI 2023
Bioinformatics and computational biology › molecular informatics
molecular design
0.712023
Sample-efficient Multi-objective Molecular Optimization with GFlowNets · NeurIPS 2023
Bioinformatics and computational biology › molecular informatics › cheminformatics › molecule generation
multi-objective molecular optimization
0.712023
Sample-efficient Multi-objective Molecular Optimization with GFlowNets · NeurIPS 2023
Bioinformatics and computational biology › drug discovery › drug response prediction
cancer drug response prediction
0.612022
TGSA: protein-protein association-based twin graph neural networks for drug response prediction with similarity augmentation · Bioinform. 2022
Bioinformatics and computational biology › drug discovery
drug response prediction
0.612022
TGSA: protein-protein association-based twin graph neural networks for drug response prediction with similarity augmentation · Bioinform. 2022
Medical and health informatics
precision medicine
0.612022
TGSA: protein-protein association-based twin graph neural networks for drug response prediction with similarity augmentation · Bioinform. 2022
Natural language and speech › Language models and text generation › neural language model
protein language model
0.312025
Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs · KDD (2) 2025
Natural language and speech › Language models and text generation
pre-trained language model
0.212024
Making Pre-trained Language Models Great on Tabular Prediction · ICLR 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.212023
GCL: Gradient-Guided Contrastive Learning for Medical Image Segmentation with Multi-Perspective Meta Labels · ACM Multimedia 2023
Bioinformatics and computational biology › protein analysis › protein-protein interaction
protein-protein interaction network
0.212022
TGSA: protein-protein association-based twin graph neural networks for drug response prediction with similarity augmentation · Bioinform. 2022

Methods — techniques the papers use, named apart from their topics

contrastive learning · 2.4structure denoising · 1.7protein language model · 1.7products of experts · 1.7potts model · 1.7mixture of experts · 1.7contrastive divergence · 1.7relative magnitude tokenization · 0.8intra-feature attention · 0.8multi-objective bayesian optimization · 0.7hypernetwork · 0.7hierarchical representation learning · 0.7GFlowNets · 0.7
YearPublicationVenuePosition
2025 Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer
abstract
Antibodies defend our health by binding to antigens with high specificity and potentiality, primarily relying on the Complementarity-Determining Region (CDR). Yet, current experimental methods of discovering new antibody CDRs are heavily time-consuming. Computational design could alleviate this burden; especially, protein language models have proven quite beneficial in many recent studies. However, most existing models solely focus on antibody potentiality and struggle to encapsulate the diverse range of plausible CDR candidates, limiting their effectiveness in real-world scenarios as binding is only one factor in the multitude of drug-forming criteria. In this paper, we introduce PG-AbD, a framework uniting Generative Flow Networks (GFlowNets) and pretrained Protein Language Models (PLMs) to successfully generate highly potent, diverse and novel antibody candidates. We innovatively construct a Products of Experts (PoE) composed by the global-distribution-modeling PLM and the local-distribution-modeling Potts Model to serve as the reward function of GFlowNet. The joint training paradigm is introduced, where PoE is trained by contrastive divergence with the negative samples generated by GFlowNet, and then guides GFlowNet to sample diverse antibody candidates. We evaluate PG-AbD on extensive antibody design benchmarks. It significantly outperforms existing methods in diversity (13.5% on RabDab, 31.1% on SabDab) while maintaining optimal potential and novelty. Generated antibodies are also found to form stable, regular 3D structures with their corresponding antigens, demonstrating the great potential of PG-AbD to accelerate real-world antibody discovery.
Mingze Yin, Hanjing Zhou, Yiheng Zhu 0002, Jialu Wu, Wei Wu 0045, Kun Fu 0002, Zheng Wang 0027, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001
AAAI3
2025 Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs
abstract
Proteins, as essential biomolecules, play a central role in biological processes, including metabolic reactions and DNA replication. Accurate prediction of their properties and functions is crucial in biological applications. Recent development of protein language models (pLMs) with supervised fine tuning provides a promising solution to this problem. However, the fine-tuned model is tailored for particular downstream prediction task, and achieving general-purpose protein understanding remains a challenge. In this paper, we introduce Structure-Enhanced Protein Instruction Tuning (SEPIT) framework to bridge this gap. Our approach incorporates a novel structure-aware module into pLMs to enrich their structural knowledge, and subsequently integrates these enhanced pLMs with large language models (LLMs) to advance protein understanding. In this framework, we propose a novel instruction tuning pipeline. First, we warm up the enhanced pLMs using contrastive learning and structure denoising. Then, caption-based instructions are used to establish a basic understanding of proteins. Finally, we refine this understanding by employing a mixture of experts (MoEs) to capture more complex properties and functional information with the same number of activated parameters. Moreover, we construct the largest and most comprehensive protein instruction dataset to date, which allows us to train and evaluate the general-purpose protein understanding model. Extensive experiments on both open-ended generation and closed-set answer tasks demonstrate the superior performance of SEPIT over both closed-source general LLMs and open-source LLMs trained with protein knowledge.
Wei Wu 0045, Chao Wang 0086, Liyi Chen 0001, Mingze Yin, Yiheng Zhu 0002, Kun Fu 0002, Jieping Ye, Hui Xiong 0001, Zheng Wang 0027
KDD (2)5
2024 Making Pre-trained Language Models Great on Tabular Prediction
abstract
The transferability of deep neural networks (DNNs) has made significant progress in image and language processing. However, due to the heterogeneity among tables, such DNN bonus is still far from being well exploited on tabular data prediction (e.g., regression or classification tasks). Condensing knowledge from diverse domains, language models (LMs) possess the capability to comprehend feature names from various tables, potentially serving as versatile learners in transferring knowledge across distinct tables and diverse prediction tasks, but their discrete text representation space is inherently incompatible with numerical feature values in tables. In this paper, we present TP-BERTa, a specifically pre-trained LM for tabular data prediction. Concretely, a novel relative magnitude tokenization converts scalar numerical feature values to finely discrete, high-dimensional tokens, and an intra-feature attention approach integrates feature values with the corresponding feature names. Comprehensive experiments demonstrate that our pre-trained TP-BERTa leads the performance among tabular DNNs and is competitive with Gradient Boosted Decision Tree models in typical tabular data regime.
Jiahuan Yan, Bo Zheng 0011, Yiheng Zhu 0002, Danny Ziyi Chen, Jimeng Sun 0001, Jian Wu 0001, Jintai Chen
ICLR4
2023 MolHF: A Hierarchical Normalizing Flow for Molecular Graph Generation
abstract
Molecular de novo design is a critical yet challenging task in scientific fields, aiming to design novel molecular structures with desired property profiles. Significant progress has been made by resorting to generative models for graphs. However, limited attention is paid to hierarchical generative models, which can exploit the inherent hierarchical structure (with rich semantic information) of the molecular graphs and generate complex molecules of larger size that we shall demonstrate to be difficult for most existing models. The primary challenge to hierarchical generation is the non-differentiable issue caused by the generation of intermediate discrete coarsened graph structures. To sidestep this issue, we cast the tricky hierarchical generation problem over discrete spaces as the reverse process of hierarchical representation learning and propose MolHF, a new hierarchical flow-based model that generates molecular graphs in a coarse-to-fine manner. Specifically, MolHF first generates bonds through a multi-scale architecture, then generates atoms based on the coarsened graph structure at each scale. We demonstrate that MolHF achieves state-of-the-art performance in random generation and property optimization, implying its high capacity to model data distribution. Furthermore, MolHF is the first flow-based model that can be applied to model larger molecules (polymer) with more than 100 heavy atoms. The code and models are available at https://github.com/violet-sto/MolHF.
Yiheng Zhu 0002, Zhenqiu Ouyang, Ben Liao, Jialu Wu, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001
IJCAI1
2023 GCL: Gradient-Guided Contrastive Learning for Medical Image Segmentation with Multi-Perspective Meta Labels
abstract
Since annotating medical images for segmentation tasks commonly incurs expensive costs, it is highly desirable to design an annotation-efficient method to alleviate the annotation burden. Recently, contrastive learning has exhibited a great potential in learning robust representations to boost downstream tasks with limited labels. In medical imaging scenarios, ready-made meta labels (i.e., specific attribute information of medical images) inherently reveal semantic relationships among images, which have been used to define positive pairs in previous work. However, the multi-perspective semantics revealed by various meta labels are usually incompatible and can incur intractable "semantic contradiction" when combining different meta labels. In this paper, we tackle the issue of "semantic contradiction" in a gradient-guided manner using our proposed Gradient Mitigator method, which systematically unifies multi-perspective meta labels to enable a pre-trained model to attain a better high-level semantic recognition ability. Moreover, we emphasize that the fine-grained discrimination ability is vital for segmentation-oriented pre-training, and develop a novel method called Gradient Filter to dynamically screen pixel pairs with the most discriminating power based on the magnitude of gradients. Comprehensive experiments on four medical image segmentation datasets verify that our new method GCL: (1) learns informative image representations and considerably boosts segmentation performance with limited labels, and (2) shows promising generalizability on out-of-distribution datasets.
Jintai Chen, Jiahuan Yan, Yiheng Zhu 0002, Danny Ziyi Chen, Jian Wu 0001
ACM Multimedia4
2023 Sample-efficient Multi-objective Molecular Optimization with GFlowNets
abstract
Many crucial scientific problems involve designing novel molecules with desired properties, which can be formulated as a black-box optimization problem over the *discrete* chemical space. In practice, multiple conflicting objectives and costly evaluations (e.g., wet-lab experiments) make the *diversity* of candidates paramount. Computational methods have achieved initial success but still struggle with considering diversity in both objective and search space. To fill this gap, we propose a multi-objective Bayesian optimization (MOBO) algorithm leveraging the hypernetwork-based GFlowNets (HN-GFN) as an acquisition function optimizer, with the purpose of sampling a diverse batch of candidate molecular graphs from an approximate Pareto front. Using a single preference-conditioned hypernetwork, HN-GFN learns to explore various trade-offs between objectives. We further propose a hindsight-like off-policy strategy to share high-performing molecules among different preferences in order to speed up learning for HN-GFN. We empirically illustrate that HN-GFN has adequate capacity to generalize over preferences. Moreover, experiments in various real-world MOBO settings demonstrate that our framework predominantly outperforms existing methods in terms of candidate quality and sample efficiency. The code is available at https://github.com/violet-sto/HN-GFN.
Yiheng Zhu 0002, Jialu Wu, Chaowen Hu, Jiahuan Yan, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001
NeurIPS1
2022 TGSA: protein-protein association-based twin graph neural networks for drug response prediction with similarity augmentation
abstract
MOTIVATION: Drug response prediction (DRP) plays an important role in precision medicine (e.g. for cancer analysis and treatment). Recent advances in deep learning algorithms make it possible to predict drug responses accurately based on genetic profiles. However, existing methods ignore the potential relationships among genes. In addition, similarity among cell lines/drugs was rarely considered explicitly. RESULTS: We propose a novel DRP framework, called TGSA, to make better use of prior domain knowledge. TGSA consists of Twin Graph neural networks for Drug Response Prediction (TGDRP) and a Similarity Augmentation (SA) module to fuse fine-grained and coarse-grained information. Specifically, TGDRP abstracts cell lines as graphs based on STRING protein-protein association networks and uses Graph Neural Networks (GNNs) for representation learning. SA views DRP as an edge regression problem on a heterogeneous graph and utilizes GNNs to smooth the representations of similar cell lines/drugs. Besides, we introduce an auxiliary pre-training strategy to remedy the identified limitations of scarce data and poor out-of-distribution generalization. Extensive experiments on the GDSC2 dataset demonstrate that our TGSA consistently outperforms all the state-of-the-art baselines under various experimental settings. We further evaluate the effectiveness and contributions of each component of TGSA via ablation experiments. The promising performance of TGSA shows enormous potential for clinical applications in precision medicine. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/violet-sto/TGSA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yiheng Zhu 0002, Zhenqiu Ouyang, Ruiwei Feng, Danny Ziyi Chen, Jian Wu 0001
Bioinform.1