VLDB 2026 Research / reviewers in the wild / expert
Hongbin Ye
dblp:274/3132
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0001-5727-5599ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Synergizing Multigrid Algorithms with Vision Transformer: A Novel Approach to Enhance the Seismic Foundation ModelabstractDue to the rapid advancement and homogenization of Artificial Intelligence (AI) technology development, transformer-based foundation models have revolutionized scientific applications, such as drug discovery, materials research, and astronomy. However, seismic data presents unique characteristics that require specialized processing techniques for pretraining foundation models in seismic contexts with high- and low-frequency features playing crucial roles. Existing Vision Transformer (ViT) with sequential image tokenization fails to efficiently and effectively capture both high- and low-frequency seismic information because they ignore the intrinsic structural patterns of seismograms. This work introduces ADATG, a novel adaptive two-grid training strategy with Hilbert encoding, explicitly tailored for seismogram data and leveraging the hierarchical structures inherent in seismic data. Specifically, our approach employs spectrum decomposition to separate high- and low-frequency components, and hierarchical Hilbert encoding to represent the data effectively. Moreover, inspired by the frequency principle, we propose an adaptive training strategy that initially emphasizes coarse-level information and then progressively refines the model's focus on fine-level features. Extensive experiments demonstrate the effectiveness and efficiency of our method. This research highlights the importance of data encoding and training strategies informed by the distinct characteristics of high- and low-frequency features in seismic images, ultimately enhancing the pretraining of visual seismic foundation models. Huiwen Wu, Hongbin Ye |
AAAI | 4 |
| 2025 | EEMS: Edge-Prompt Enhanced Medical Image Segmentation Based on Learnable Gating MechanismabstractMedical image segmentation is vital for diagnosis, treatment planning, and disease monitoring but is challenged by complex factors like ambiguous edges and background noise. We introduce EEMS, a new model for segmentation, combining an Edge-Aware Enhancement Unit (EAEU) and a Multi-scale Prompt Generation Unit (MSPGU). EAEU enhances edge perception via multi-frequency feature extraction, accurately defining boundaries. MSPGU integrates high-level semantic and low-level spatial features using a prompt-guided approach, ensuring precise target localization. The Dual-Source Adaptive Gated Fusion Unit (DAGFU) merges edge features from EAEU with semantic features from MSPGU, enhancing segmentation accuracy and robustness. Tests on datasets like ISIC2018 confirm EEMS's superior performance and reliability as a clinical tool. Quanjun Li, Zimeng Li 0001, Hongbin Ye, Yupeng Liu 0003, Haolun Li 0001, Xuhang Chen 0002 |
BIBM | 5 |
| 2024 | ProTeM: Unifying Protein Function Prediction via Text Matching
Ming Qin, Yuhao Wang 0006, Hongbin Ye, Zongbing Wang, Weihao Gao, Shangsong Liang, Qiang Zhang 0026, Keyan Ding |
ICANN (8) | 5 |
| 2024 | InstructIE: A Bilingual Instruction-based Information Extraction Dataset
Honghao Gui, Shuofei Qiao, Jintian Zhang, Hongbin Ye, Mengshu Sun, Lei Liang 0002, Jeff Z. Pan, Huajun Chen, Ningyu Zhang 0001 |
ISWC (3) | 4 |
| 2023 | Active Finetuning Protein Language Model: A Budget-Friendly Method for Directed EvolutionabstractDirected evolution is a widely-used strategy of protein engineering to improve protein function via mimicking natural mutation and selection. Machine learning-assisted directed evolution (MLDE) approaches aim to learn a fitness predictor, thereby efficiently searching for optimal mutants within the vast combinatorial mutation space. Since annotating mutants is both costly and labor-intensive, how to efficiently sample and utilize informative protein mutants to train the predictor is a critical problem in MLDE. Previous MLDE works just simply utilized pre-trained protein language models (PPLMs) for sampling without tailoring to the specific target protein of interest, which has not fully exploited the potential of PPLMs. In this work, we propose a novel method, the Actively-Finetuned Protein language model for Directed Evolution(AFP-DE), which leverages PPLMs to actively sample and fine-tune themselves, continuously improving the model’s sampling and overall performance through iterations, to achieve efficient directed protein evolution. Extensive experiments have shown the effectiveness of our method in generating optimal mutants with minimal annotation effort, outperforming previous works even with fewer annotated mutants, making it budget-friendly for biological experiments. Ming Qin, Keyan Ding, Bin Wu 0025, Haihong Yang, Hongbin Ye, Huajun Chen, Qiang Zhang 0026 |
ECAI | 7 |
| 2023 | LOGEN: Few-Shot Logical Knowledge-Conditioned Text Generation With Self-TrainingabstractNatural language generation from structured data mainly focuses on surface-level descriptions, suffering from uncontrollable content selection and low fidelity. Previous works leverage logical forms to facilitate logical knowledge-conditioned text generation. Though achieving remarkable progress, they are data-hungry, which makes the adoption for real-world applications challenging with limited data. To this end, this paper proposes a unified framework for logical knowledge-conditioned text generation in the few-shot setting. With only a few seeds logical forms (e.g., 20/100 shot), our approach leverages self-training and samples pseudo logical forms based on content and structure consistency. Experimental results demonstrate that our approach can obtain better few-shot performance than baselines. Shumin Deng, Hongbin Ye, Chuanqi Tan, Mosha Chen, Songfang Huang, Fei Huang 0002, Huajun Chen, Ningyu Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Learning to Ask for Data-Efficient Event Argument Extraction (Student Abstract)abstractEvent argument extraction (EAE) is an important task for information extraction to discover specific argument roles. In this study, we cast EAE as a question-based cloze task and empirically analyze fixed discrete token template performance. As generating human-annotated question templates is often time-consuming and labor-intensive, we further propose a novel approach called “Learning to Ask,” which can learn optimized question templates for EAE without human annotations. Experiments using the ACE-2005 dataset demonstrate that our method based on optimized questions achieves state-of-the-art performance in both the few-shot and supervised settings. Hongbin Ye, Ningyu Zhang 0001, Zhen Bi, Shumin Deng, Chuanqi Tan, Hui Chen 0018, Fei Huang 0002, Huajun Chen |
AAAI | 1 |
| 2022 | Generative Knowledge Graph Construction: A ReviewabstractGenerative Knowledge Graph Construction (KGC) refers to those methods that leverage the sequence-to-sequence framework for building knowledge graphs, which is flexible and can be adapted to widespread tasks.In this study, we summarize the recent compelling progress in generative knowledge graph construction.We present the advantages and weaknesses of each paradigm in terms of different generation targets and provide theoretical insight and empirical analysis.Based on the review, we suggest promising research directions for the future.Our contributions are threefold: (1) We present a detailed, complete taxonomy for the generative KGC methods; (2) We provide a theoretical and empirical analysis of the generative KGC methods; (3) We propose several research directions that can be developed in the future. Hongbin Ye, Ningyu Zhang 0001, Hui Chen 0018, Huajun Chen |
EMNLP | 1 |
| 2022 | Ontology-enhanced Prompt-tuning for Few-shot LearningabstractFew-shot Learning (FSL) is aimed to make predictions based on a limited number of samples. Structured data such as knowledge graphs and ontology libraries has been leveraged to benefit the few-shot setting in various tasks. However, the priors adopted by the existing methods suffer from challenging knowledge missing, knowledge noise, and knowledge heterogeneity, which hinder the performance for few-shot learning. In this study, we explore knowledge injection for FSL with pre-trained language models and propose ontology-enhanced prompt-tuning (OntoPrompt). Specifically, we develop the ontology transformation based on the external knowledge graph to address the knowledge missing issue, which fulfills and converts structure knowledge to text. We further introduce span-sensitive knowledge injection via a visible matrix to select informative knowledge to handle the knowledge noise issue. To bridge the gap between knowledge and text, we propose a collective training algorithm to optimize representations jointly. We evaluate our proposed OntoPrompt in three tasks, including relation extraction, event extraction, and knowledge graph completion, with eight datasets. Experimental results demonstrate that our approach can obtain better few-shot performance than baselines. Hongbin Ye, Ningyu Zhang 0001, Shumin Deng, Xiang Chen 0016, Hui Chen 0018, Feiyu Xiong, Xi Chen 0003, Huajun Chen |
WWW | 1 |
| 2022 | Robust triple extraction with cascade bidirectional capsule network
Ningyu Zhang 0001, Shumin Deng, Hongbin Ye, Wei Zhang 0127, Huajun Chen |
Expert Syst. Appl. | 3 |
| 2021 | Contrastive Triple Extraction with Generative TransformerabstractTriple extraction is an essential task in information extraction for natural language processing and knowledge graph construction. In this paper, we revisit the end-to-end triple extraction task for sequence generation. Since generative triple extraction may struggle to capture long-term dependencies and generate unfaithful triples, we introduce a novel model, contrastive triple extraction with a generative transformer. Specifically, we introduce a single shared transformer module for encoder-decoder-based generation. To generate faithful results, we propose a novel triplet contrastive training object. Moreover, we introduce two mechanisms to further improve model performance (i.e., batch-wise dynamic attention-masking and triple-wise calibration). Experimental results on three datasets (i.e., NYT, WebNLG, and MIE) show that our approach achieves better performance than that of baselines. Hongbin Ye, Ningyu Zhang 0001, Shumin Deng, Mosha Chen, Chuanqi Tan, Fei Huang 0002, Huajun Chen |
AAAI | 1 |
| 2021 | AliCG: Fine-grained and Evolvable Conceptual Graph Construction for Semantic Search at AlibabaabstractConceptual graphs, which is a particular type of Knowledge Graphs, play an essential role in semantic search. Prior conceptual graph construction approaches typically extract high-frequent, coarse-grained, and time-invariant concepts from formal texts such as Wikipedia. In real applications, however, it is necessary to extract less-frequent, fine-grained, and time-varying conceptual knowledge and build taxonomy in an evolving manner. In this paper, we introduce an approach to implementing and deploying the conceptual graph at Alibaba. Specifically, We propose a framework called AliCG which is capable of a) extracting fine-grained concepts by a novel bootstrapping with alignment consensus approach, b) mining long-tail concepts with a novel low-resource phrase mining approach, c) updating the graph dynamically via a concept distribution estimation method based on implicit and explicit user behaviors. We have deployed the conceptual graph at Alibaba UC Browser. Extensive offline evaluation as well as online A/B testing demonstrate the efficacy of our approach. Ningyu Zhang 0001, Qianghuai Jia, Shumin Deng, Xiang Chen 0016, Hongbin Ye, Hui Chen 0018, Huaixiao Tou, Gang Huang 0004, Nengwei Hua, Huajun Chen |
KDD | 5 |
| 2021 | Contrastive Information Extraction With Generative TransformerabstractInformation extraction tasks such as triple extraction and event extraction are of great importance for natural language processing and knowledge graph construction. In this paper, we revisit the end-to-end information extraction task for sequence generation. Since generative information extraction may struggle to capture long-term dependencies and generate unfaithful triples, we introduce a novel model, contrastive information extraction with a generative transformer. Specifically, we introduce a single shared transformer module for an encoder-decoder-based generation. To generate faithful results, we propose a novel triplet contrastive training object. Moreover, we introduce two mechanisms to further improve model performance (i.e., batch-wise dynamic attention-masking and triple-wise calibration). Experimental results on five datasets (i.e., NYT, WebNLG, MIE, ACE-2005, and MUC-4) show that our approach achieves better performance than baselines. Ningyu Zhang 0001, Hongbin Ye, Shumin Deng, Chuanqi Tan, Mosha Chen, Songfang Huang, Fei Huang 0002, Huajun Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Bridging Text and Knowledge with Multi-Prototype Embedding for Few-Shot Relational Triple ExtractionabstractCurrent supervised relational triple extraction approaches require huge amounts of labeled data and thus suffer from poor performance in few-shot settings.However, people can grasp new knowledge by learning a few instances.To this end, we take the first step to study the few-shot relational triple extraction, which has not been well understood.Unlike previous single-task few-shot problems, relational triple extraction is more challenging as the entities and relations have implicit correlations.In this paper, We propose a novel multi-prototype embedding network model to jointly extract the composition of relational triples, namely, entity pairs and corresponding relations.To be specific, we design a hybrid prototypical learning mechanism that bridges text and knowledge concerning both entities and relations.Thus, implicit correlations between entities and relations are injected.Additionally, we propose a prototype-aware regularization to learn more representative prototypes.Experimental results demonstrate that the proposed method can improve the performance of the few-shot triple extraction. Haiyang Yu 0003, Ningyu Zhang 0001, Shumin Deng, Hongbin Ye, Wei Zhang 0127, Huajun Chen |
COLING | 4 |