Hehuan Ma

dblp:263/2003 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-5971-0053ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Learning from Guidelines: Structured Prompt Optimization for Expert Annotation Tasks
abstract
Deep learning has significantly advanced numerous fields by training on extensive annotated datasets. However, this data-driven paradigm faces limitations such as limited adaptability and high annotation costs, particularly when precise adherence to detailed, domain-specific guidelines is required in annotation. This challenge raises a critical question: Can models effectively shift from data-driven learning to autonomously leveraging guidelines with minimal annotated examples? To address this, we propose the Guideline-Driven Prompt (GDP) optimization framework, which shifts the learning paradigm from data-driven training to guideline-driven reasoning. GDP leverages Retrieval Augmented Generation (RAG) to retrieve essential fragments from complex guidelines and synthesize them into structured, executable prompts. A tree-based optimization algorithm systematically constructs and refines these prompts, explicitly capturing the intricate logic embedded in professional guidelines through a latent pipeline structure. Empirical evaluations on four datasets ranging from diverse domains and different tasks demonstrate that GDP effectively transitions the learning process from data-intensive methods to a guideline-driven approach in tasks requiring detailed and complex guideline adherence, reducing dependence on extensive annotated datasets.
Leon Wenliang Zhong, Thao M. Dang, Feng Jiang 0012, Hehuan Ma, Yuzhi Guo, Jean Gao, Junzhou Huang
AAAI5
2026 Guidelines as Environments: A World Model Approach to Rule Following
abstract
Haiqing Li, Wenliang Zhong, Yinhao Wu, Hehuan Ma, Yuzhi Guo, Thao M. Dang, Junzhou Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Leon Wenliang Zhong, Yinhao Wu, Hehuan Ma, Yuzhi Guo, Thao M. Dang, Junzhou Huang
ACL (1)4
2025 GoBERT: Gene Ontology Graph Informed BERT for Universal Gene Function Prediction
abstract
Exploring the functions of genes and gene products is crucial to a wide range of fields, including medical research, evolutionary biology, and environmental science. However, discovering new functions largely relies on expensive and exhaustive wet lab experiments. Existing methods of automatic function annotation or prediction mainly focus on protein function prediction with sequence, 3D-structures or protein family information. In this study, we propose to tackle the gene function prediction problem by exploring Gene Ontology graph and annotation with BERT (GoBERT) to decipher the underlying relationships among gene functions. Our proposed novel function prediction task utilizes existing functions as inputs and generalizes the function prediction to gene and gene products. Specifically, two pre-train tasks are designed to jointly train GoBERT to capture both explicit and implicit relations of functions. Neighborhood prediction is a self-supervised multi-label classification task that captures the explicit function relations. Specified masking and recovering task helps GoBERT in finding implicit patterns among functions. The pre-trained GoBERT possess the ability to predict novel functions for various gene and gene products based on known functional annotations. Extensive experiments, biological case studies, and ablation studies are conducted to demonstrate the superiority of our proposed GoBERT.
Yuwei Miao, Yuzhi Guo, Hehuan Ma, Jingquan Yan, Feng Jiang 0012, Rui Liao, Junzhou Huang
AAAI3
2025 Zero-Shot Composed Image Retrieval via Dual-Stream Instruction-Aware Distillation
Leon Wenliang Zhong, Robert A. Barton, Weizhi An, Feng Jiang 0012, Hehuan Ma, Yuzhi Guo, Abhishek Dan, Shioulin Sam, Karim Bouyarmane, Junzhou Huang
ICCV5
2025 HAGE: Hierarchical Alignment Gene-Enhanced Pathology Representation Learning with Spatial Transcriptomics
Thao M. Dang, Yuzhi Guo, Hehuan Ma, Feng Jiang 0012, Yuwei Miao, Qifeng Zhou, Jean Gao, Junzhou Huang
MICCAI (1)4
2025 Text-Guided Multi-instance Learning for Scoliosis Screening via Gait Video Analysis
Yuzhi Guo, Feng Jiang 0012, Thao M. Dang, Hehuan Ma, Qifeng Zhou, Jean Gao, Junzhou Huang
MICCAI (6)5
2025 TRIDENT: Tri-Modal Molecular Representation Learning with Taxonomic Annotations and Local Correspondence
abstract
Molecular property prediction aims to learn representations that map chemical structures to functional properties. While multimodal learning has emerged as a powerful paradigm to learn molecular representations, prior works have largely overlooked textual and taxonomic information of molecules for representation learning. We introduce TRIDENT, a novel framework that integrates molecular SMILES, textual descriptions, and taxonomic functional annotations to learn rich molecular representations. To achieve this, we curate a comprehensive dataset of molecule-text pairs with structured, multi-level functional annotations. Instead of relying on conventional contrastive loss, TRIDENT employs a volume-based alignment objective to jointly align tri-modal features at the global level, enabling soft, geometry-aware alignment across modalities. Additionally, TRIDENT introduces a novel local alignment objective that captures detailed relationships between molecular substructures and their corresponding sub-textual descriptions. A momentum-based mechanism dynamically balances global and local alignment, enabling the model to learn both broad functional semantics and fine-grained structure-function mappings. TRIDENT achieves state-of-the-art performance on 18 downstream tasks, demonstrating the value of combining SMILES, textual, and taxonomic functional annotations for molecular property prediction. Our code and data are available at https://github.com/uta-smile/TRIDENT.
Feng Jiang 0012, Mangal Prakash, Hehuan Ma, Jianyuan Deng, Yuzhi Guo, Amina Mollaysa, Tommaso Mansi, Rui Liao, Junzhou Huang
NeurIPS3
2025 Segment Any Cell: A SAM-Based Auto-Prompting Fine-Tuning Framework for Nuclei Segmentation
abstract
In the rapidly evolving field of AI research, foundational models like BERT and GPT have significantly advanced language and vision tasks. The advent of pretrain-prompting models, such as ChatGPT and segment anything model (SAM), has further revolutionized image segmentation. However, their applications in specialized areas, particularly in nuclei segmentation within medical imaging, reveal a key challenge: the generation of high-quality, informative prompts is as crucial as applying state-of-the-art (SOTA) fine-tuning techniques on foundation models. To address this, we introduce segment any cell (SAC), an innovative framework that enhances SAM specifically for nuclei segmentation. SAC integrates a low-rank adaptation (LoRA) within the attention layer of the Transformer to improve the fine-tuning process, outperforming existing SOTA methods. It also introduces an innovative auto-prompt generator that produces effective prompts to guide segmentation, a critical factor in handling the complexities of nuclei segmentation in biomedical imaging. Our extensive experiments demonstrate the superiority of SAC in nuclei segmentation tasks, proving its effectiveness as a tool for pathologists and researchers. Our contributions include a novel prompt generation strategy, automated adaptability for diverse segmentation tasks, the innovative application of low-rank attention adaptation in SAM, and a versatile framework for semantic and instance segmentation challenges.
Saiyang Na, Yuzhi Guo, Feng Jiang 0012, Hehuan Ma, Jean Gao, Junzhou Huang
IEEE Trans. Neural Networks Learn. Syst.4
2024 Causal Subgraphs and Information Bottlenecks: Redefining OOD Robustness in Graph Neural Networks
Weizhi An, Leon Wenliang Zhong, Feng Jiang 0012, Hehuan Ma, Junzhou Huang
ECCV (88)4
2024 PathM3: A Multimodal Multi-task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning
Qifeng Zhou, Leon Wenliang Zhong, Yuzhi Guo, Michael Xiao, Hehuan Ma, Junzhou Huang
MICCAI (4)5
2024 GTE: a graph learning framework for prediction of T-cell receptors and epitopes binding specificity
abstract
The interaction between T-cell receptors (TCRs) and peptides (epitopes) presented by major histocompatibility complex molecules (MHC) is fundamental to the immune response. Accurate prediction of TCR-epitope interactions is crucial for advancing the understanding of various diseases and their prevention and treatment. Existing methods primarily rely on sequence-based approaches, overlooking the inherent topology structure of TCR-epitope interaction networks. In this study, we present $GTE$, a novel heterogeneous Graph neural network model based on inductive learning to capture the topological structure between TCRs and Epitopes. Furthermore, we address the challenge of constructing negative samples within the graph by proposing a dynamic edge update strategy, enhancing model learning with the nonbinding TCR-epitope pairs. Additionally, to overcome data imbalance, we adapt the Deep AUC Maximization strategy to the graph domain. Extensive experiments are conducted on four public datasets to demonstrate the superiority of exploring underlying topological structures in predicting TCR-epitope interactions, illustrating the benefits of delving into complex molecular networks. The implementation code and data are available at https://github.com/uta-smile/GTE.
Feng Jiang 0012, Yuzhi Guo, Hehuan Ma, Saiyang Na, Leon Wenliang Zhong, Tao Wang 0161, Junzhou Huang
Briefings Bioinform.3
2022 Self-Supervised Pre-training for Protein Embeddings Using Tertiary Structures
abstract
The protein tertiary structure largely determines its interaction with other molecules. Despite its importance in various structure-related tasks, fully-supervised data are often time-consuming and costly to obtain. Existing pre-training models mostly focus on amino-acid sequences or multiple sequence alignments, while the structural information is not yet exploited. In this paper, we propose a self-supervised pre-training model for learning structure embeddings from protein tertiary structures. Native protein structures are perturbed with random noise, and the pre-training model aims at estimating gradients over perturbed 3D structures. Specifically, we adopt SE(3)-invariant features as model inputs and reconstruct gradients over 3D coordinates with SE(3)-equivariance preserved. Such paradigm avoids the usage of sophisticated SE(3)-equivariant models, and dramatically improves the computational efficiency of pre-training models. We demonstrate the effectiveness of our pre-training model on two downstream tasks, protein structure quality assessment (QA) and protein-protein interaction (PPI) site prediction. Hierarchical structure embeddings are extracted to enhance corresponding prediction models. Extensive experiments indicate that such structure embeddings consistently improve the prediction accuracy for both downstream tasks.
Yuzhi Guo, Jiaxiang Wu 0001, Hehuan Ma, Junzhou Huang
AAAI3
2022 Cross-dependent graph neural networks for molecular property prediction
abstract
MOTIVATION: The crux of molecular property prediction is to generate meaningful representations of the molecules. One promising route is to exploit the molecular graph structure through graph neural networks (GNNs). Both atoms and bonds significantly affect the chemical properties of a molecule, so an expressive model ought to exploit both node (atom) and edge (bond) information simultaneously. Inspired by this observation, we explore the multi-view modeling with GNN (MVGNN) to form a novel paralleled framework, which considers both atoms and bonds equally important when learning molecular representations. In specific, one view is atom-central and the other view is bond-central, then the two views are circulated via specifically designed components to enable more accurate predictions. To further enhance the expressive power of MVGNN, we propose a cross-dependent message-passing scheme to enhance information communication of different views. The overall framework is termed as CD-MVGNN. RESULTS: We theoretically justify the expressiveness of the proposed model in terms of distinguishing non-isomorphism graphs. Extensive experiments demonstrate that CD-MVGNN achieves remarkably superior performance over the state-of-the-art models on various challenging benchmarks. Meanwhile, visualization results of the node importance are consistent with prior knowledge, which confirms the interpretability power of CD-MVGNN. AVAILABILITY AND IMPLEMENTATION: The code and data underlying this work are available in GitHub at https://github.com/uta-smile/CD-MVGNN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hehuan Ma, Yatao Bian, Yu Rong 0001, Wenbing Huang 0001, Tingyang Xu, Weiyang Xie, Geyan Ye, Junzhou Huang
Bioinform.1
2021 Hierarchical Graph Capsule Network
abstract
Graph Neural Networks (GNNs) draw their strength from explicitly modeling the topological information of structured data. However, existing GNNs suffer from limited capability in capturing the hierarchical graph representation which plays an important role in graph classification. In this paper, we innovatively propose hierarchical graph capsule network (HGCN) that can jointly learn node embeddings and extract graph hierarchies. Specifically, disentangled graph capsules are established by identifying heterogeneous factors underlying each node, such that their instantiation parameters represent different properties of the same entity. To learn the hierarchical representation, HGCN characterizes the part-whole relationship between lower-level capsules (part) and higher-level capsules (whole) by explicitly considering the structure information among the parts. Experimental studies demonstrate the effectiveness of HGCN and the contribution of each component. Code: https://github.com/uta-smile/HGCN
Peilin Zhao, Yu Rong 0001, Chaochao Yan, Chunyuan Li, Hehuan Ma, Junzhou Huang
AAAI6
2021 Gradient-Norm Based Attentive Loss for Molecular Property Prediction
abstract
Molecular property prediction is one fundamental yet challenging task for drug discovery. Many studies have addressed this problem by designing deep learning algorithms, e.g., sequence-based models and graph-based models. However, the underlying data distribution is rarely explored. We discover that there exist easy samples and hard samples in the molecule datasets, and the overall distribution is usually imbalanced. Current research mainly treats them equally during the model training, while we believe that they shall not share the same weights since neural networks training is dominated by the majority class. Therefore, we propose to utilize a self-attention mechanism to generate a learnable weight for each data sample according to the associated gradient norm. The learned attention value is then embedded into the prediction models to construct an attentive loss for the network updating and back-propagation. It is empirically demonstrated that our proposed method can consistently boost the prediction performance for both classification and regression tasks.
Hehuan Ma, Yu Rong 0001, Yuzhi Guo, Chaochao Yan, Junzhou Huang
BIBM1
2021 Exploring Robustness of Unsupervised Domain Adaptation in Semantic Segmentation
abstract
Recent studies imply that deep neural networks are vulnerable to adversarial examples, i.e., inputs with a slight but intentional perturbation are incorrectly classified by the network. Such vulnerability makes it risky for some security-related applications (e.g., semantic segmentation in autonomous cars) and triggers tremendous concerns on the model reliability. For the first time, we comprehensively evaluate the robustness of existing UDA methods and propose a robust UDA approach. It is rooted in two observations: i) the robustness of UDA methods in semantic segmentation remains unexplored, which poses a security concern in this field; and ii) although commonly used self-supervision (e.g., rotation and jigsaw) benefits model robustness in classification and recognition tasks, they fail to provide the critical supervision signals that are essential in semantic segmentation. These observations motivate us to propose adversarial self-supervision UDA (or ASSUDA) that maximizes the agreement between clean images and their adversarial examples by a contrastive loss in the output space. Extensive empirical studies on commonly used benchmarks demonstrate that ASSUDA is resistant to adversarial attacks.
Chunyuan Li, Weizhi An, Hehuan Ma, Yuzhi Guo, Yu Rong 0001, Peilin Zhao, Junzhou Huang
ICCV4
2020 Protein Ensemble Learning with Atrous Spatial Pyramid Networks for Secondary Structure Prediction
abstract
The secondary structure of proteins is significant for studying the three-dimensional structure and functions of proteins. Several models from image understanding and natural language modeling have been successfully adapted in the protein sequence study area, such as Long Short-term Memory (LSTM) network and Convolutional Neural Network (CNN). Recently, Gated Convolutional Neural Network (GCNN) has been proposed for natural language processing and reduces latency while achieving high levels of sentence scoring. Conditionally Parameterized Convolution (CondConv), which use extra sample-dependant modules to conditionally adjust the convolutional network, have achieved great success in the image processing area. In this paper, we propose a novel Conditionally Parameterized Convolutional network (CondGCNN) which utilize the power of both CondConv and GCNN, and we leverage an ensemble encoder to combine the capabilities of both LSTM and CondGCNN to encode protein sequences to obtain better sequential features from proteins. In addition, due to the similarity between the image segmentation problem and the secondary structure prediction problem, we propose an ASP network (Atrous Spatial Pyramid Pooling (ASPP) based network) as the secondary structure generator in our proposed framework. We have conducted extensive ablation studies over each component in the proposed model to verify its effectiveness. Extensive experiments show that the proposed method can achieve higher performance on protein secondary structure prediction task than existing methods on CB513, Caspll and CASP12 datasets. Our method is expected to be useful for protein structure and further protein functions prediction.
Yuzhi Guo, Jiaxiang Wu 0001, Hehuan Ma, Sheng Wang 0001, Junzhou Huang
BIBM3
2020 WeightAln: Weighted Homologous Alignment for Protein Structure Property Prediction
abstract
Accurately predicting protein structure properties is essential in analyzing the structure and function of a protein, such as secondary structure, solvent accessibility, and dihedral angles. Multiple Sequence Alignment (MSA), which is a sequence alignment of multiple homologous protein sequences for the target protein, is widely used in the protein structure property prediction. The most popular strategy to exploit MSA is converting it into a position-specific scoring matrice (PSSM), then inputs the PSSM to the relevant prediction networks. PSSM is obtained by simply counting the frequency of amino acids presented at each position in the corresponding MSA, which means, each sequence in the MSA has the same weight to the target protein. However, simply setting the weights of homologous protein sequences of a protein as same cannot sufficiently model the complex relationships between them. Moreover, some sequences within the MSA are redundant, which raises a tantalizing question: can we generate a different weight for each sequence in the MSA and use the weighted PSSM to improve the performance of protein structure property prediction? To help answer this question, we present WeightAln framework, which to our knowledge, is the first attempt to generate learnable MSA weights for protein prediction tasks. We prove the effectiveness of our method by conducting extensive experiments on three protein structure property prediction tasks.
Yuzhi Guo, Jiaxiang Wu 0001, Hehuan Ma, Xinliang Zhu, Junzhou Huang
BIBM3
2020 Improving Molecular Property Prediction on Limited Data with Deep Multi-Label Learning
abstract
Acquiring labeled data has been widely recognized as a major challenge in molecular property prediction. Since it generally requires a series of specialized biochemical experiments which are time-consuming, costly, as well as labor-intensive. The deficiency of labeled property data makes it difficult to learn a good prediction model. Here, we propose an RNN-based multi-label molecular property prediction method to alleviate the data scarcity issue in two stages: 1) utilize the abundant unlabeled SMILES data to pre-train a seq2seq model whose encoder learns to generate molecular fingerprint based on the given SMILES; and 2) finetune the pre-trained model on the labeled molecular property data. Since labeled data is limited, we train those properties with limited sample size jointly with other properties which contain relatively sufficient samples. This approach brings in the idea of multi-label training, which is able to pre-train and fine-tune the encoder network, as well as train the prediction network with a data augmentation strategy. Extensive experiments on molecular property prediction demonstrate that our proposed method has achieved superior performance compared with the state-of-the-art approaches on properties with limited sample size.
Hehuan Ma, Chaochao Yan, Yuzhi Guo, Sheng Wang 0001, Hongmao Sun, Junzhou Huang
BIBM1
2020 Bagging MSA Learning: Enhancing Low-Quality PSSM with Deep Learning for Accurate Protein Structure Property Prediction
Yuzhi Guo, Jiaxiang Wu 0001, Hehuan Ma, Sheng Wang 0001, Junzhou Huang
RECOMB3