EDBT 2026 Demo / reviewers in the wild / expert
Haixing Dai
dblp:210/2444
· DBLP profile ↗
12ranked-venue papers
1as first author
12since 2021 · last 2025
0000-0003-0409-6129ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning lifespan brain anatomical correspondence via cortical developmental continuity transfer
Lu Zhang 0050, Zhengwang Wu, Xiaowei Yu 0001, Yanjun Lyu, Zihao Wu 0001, Haixing Dai, Lin Zhao 0004, Li Wang 0026, Gang Li 0001, Xianqiao Wang, Tianming Liu 0001, Dajiang Zhu |
Medical Image Anal. | 6 |
| 2025 | AugGPT: Leveraging ChatGPT for Text Data AugmentationabstractText data augmentation is an effective strategy for overcoming the challenge of limited sample sizes in many natural language processing (NLP) tasks. This challenge is especially prominent in the few-shot learning (FSL) scenario, where the data in the target domain is generally much scarcer and of lowered quality. A natural and widely used strategy to mitigate such challenges is to perform data augmentation to better capture data invariance and increase the sample size. However, current text data augmentation methods either can’t ensure the correct labeling of the generated data (lacking faithfulness), or can’t ensure sufficient diversity in the generated data (lacking compactness), or both. Inspired by the recent success of large language models (LLM), especially the development of ChatGPT, we propose a text data augmentation approach based on ChatGPT (named ”AugGPT”). AugGPT rephrases each sentence in the training samples into multiple conceptually similar but semantically different samples. The augmented samples can then be used in downstream model training. Experiment results on multiple few-shot learning text classification tasks show the superior performance of the proposed AugGPT approach over state-of-the-art text data augmentation methods in terms of testing accuracy and distribution of the augmented samples. Haixing Dai, Zhengliang Liu, Wenxiong Liao, Zihao Wu 0001, Lin Zhao 0004, Shaochen Xu, Fang Zeng, Wei Liu 0146, Ninghao Liu 0001, Sheng Li 0001, Dajiang Zhu, Hongmin Cai, Lichao Sun 0001, Quanzheng Li, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001 |
IEEE Trans. Big Data | 1 |
| 2025 | Exploring New Frontiers in Agricultural NLP: Investigating the Potential of Large Language Models for Food ApplicationsabstractThis paper explores new frontiers in agricultural natural language processing (NLP) by investigating the effectiveness of food-related text corpora for pretraining transformer-based language models. Specifically, we focus on semantic matching, establishing mappings between food descriptions and nutrition data through fine-tuning AgriBERT with the FoodOn ontology. Our work introduces an expanded comparison with state-of-the-art language models such as GPT-4, Mistral-large, Claude 3 Sonnet, and Gemini 1.0 Ultra. This exploratory investigation, rather than a direct comparison, aims to understand how AgriBERT, a domain-specific, fine-tuned, open-source model, complements the broad knowledge and generative abilities of these advanced LLMs in addressing the unique challenges of the agricultural sector. We also experiment with other applications, such as cuisine prediction from ingredients, expanding our research to include various NLP tasks beyond semantic matching. Overall, this paper underscores the potential of integrating domain-specific models like AgriBERT with advanced LLMs to enhance the performance and applicability of agricultural NLP applications. Saed Rezayi, Zhengliang Liu, Zihao Wu 0001, Chandra Dhakal, Bao Ge, Haixing Dai, Gengchen Mai, Ninghao Liu 0001, Chen Zhen, Tianming Liu 0001, Sheng Li 0001 |
IEEE Trans. Big Data | 6 |
| 2025 | Exploring the Trade-Offs: Unified Large Language Models vs Local Fine-Tuned Models for Highly-Specific Radiology NLI TaskabstractRecently, ChatGPT and GPT-4 have emerged and gained immense global attention due to their unparalleled performance in language processing. Despite demonstrating impressive capability in various open-domain tasks, their adequacy in highly specific fields like radiology remains untested. Radiology presents unique linguistic phenomena distinct from open-domain data due to its specificity and complexity. Assessing the performance of large language models (LLMs) in such specific domains is crucial not only for a thorough evaluation of their overall performance but also for providing valuable insights into future model design directions: whether model design should be generic or domain-specific. To this end, in this study, we evaluate the performance of ChatGPT/GPT-4 on a radiology natural language inference (NLI) task and compare it to other models fine-tuned specifically on task-related data samples. We also conduct a comprehensive investigation on ChatGPT/GPT-4’s reasoning ability by introducing varying levels of inference difficulty. Our results show that 1) ChatGPT and GPT-4 outperform other LLMs in the radiology NLI task and 2) other specifically fine-tuned Bert-based models require significant amounts of data samples to achieve comparable performance to ChatGPT/GPT-4. These findings not only demonstrate the feasibility and promise of constructing a generic model capable of addressing various tasks across different domains, but also highlight several key factors crucial for developing a unified model, particularly in a medical context, paving the way for future artificial general intelligence (AGI) systems. We release our code and data to the research community. Zihao Wu 0001, Lu Zhang 0050, Xiaowei Yu 0001, Zhengliang Liu, Lin Zhao 0004, Yiwei Li 0002, Haixing Dai, Chong Ma 0004, Gang Li 0001, Wei Liu 0146, Quanzheng Li, Dinggang Shen, Xiang Li 0001, Dajiang Zhu, Tianming Liu 0001 |
IEEE Trans. Big Data | 8 |
| 2025 | Mask-Guided Vision Transformer for Few-Shot LearningabstractLearning with little data is challenging but often inevitable in various application scenarios where the labeled data are limited and costly. Recently, few-shot learning (FSL) gained increasing attention because of its generalizability of prior knowledge to new tasks that contain only a few samples. However, for data-intensive models such as vision transformer (ViT), current fine-tuning-based FSL approaches are inefficient in knowledge generalization and, thus, degenerate the downstream task performances. In this article, we propose a novel mask-guided ViT (MG-ViT) to achieve an effective and efficient FSL on the ViT model. The key idea is to apply a mask on image patches to screen out the task-irrelevant ones and to guide the ViT focusing on task-relevant and discriminative patches during FSL. Particularly, MG-ViT only introduces an additional mask operation and a residual connection, enabling the inheritance of parameters from pretrained ViT without any other cost. To optimally select representative few-shot samples, we also include an active learning-based sample selection method to further improve the generalizability of MG-ViT-based FSL. We evaluate the proposed MG-ViT on classification, object detection, and segmentation tasks using gradient-weighted class activation mapping (Grad-CAM) to generate masks. The experimental results show that the MG-ViT model significantly improves the performance and efficiency compared with general fine-tuning-based ViT and ResNet models, providing novel insights and a concrete approach toward generalizing data-intensive and large-scale deep learning models for FSL. Yuzhong Chen 0002, Zhenxiang Xiao, Yi Pan 0001, Lin Zhao 0004, Haixing Dai, Zihao Wu 0001, Changhe Li, Changying Li, Dajiang Zhu, Tianming Liu 0001, Xi Jiang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Mask-guided BERT for few-shot text classification
Wenxiong Liao, Zhengliang Liu, Haixing Dai, Zihao Wu 0001, Yiyang Zhang 0003, Yuzhong Chen 0002, Xi Jiang 0001, Dajiang Zhu, Sheng Li 0001, Wei Liu 0146, Tianming Liu 0001, Quanzheng Li, Hongmin Cai, Xiang Li 0001 |
Neurocomputing | 3 |
| 2024 | BI-AVAN: A Brain-Inspired Adversarial Visual Attention Network for Characterizing Human Visual Attention From Neural ActivityabstractVisual attention is a fundamental mechanism in the human brain, and it inspires the design of attention mechanisms in deep neural networks. However, most of the visual attention studies adopted eye-tracking data rather than the direct measurement of brain activity to characterize human visual attention. In addition, the adversarial relationship between the attention-related objects and attention-neglected background in the human visual system was not fully exploited. To bridge these gaps, we propose a novel brain-inspired adversarial visual attention network (BI-AVAN) to characterize human visual attention directly from functional brain activity. Our BI-AVAN model imitates the biased competition process between attention-related/neglected objects to identify and locate the visual objects in a movie frame the human brain focuses on in an unsupervised manner. We use independent eye-tracking data as ground truth for validation and experimental results show that our model achieves robust and promising results when inferring meaningful human visual attention and mapping the relationship between brain activities and visual stimuli. Our BI-AVAN model contributes to the emerging field of leveraging the brain's functional architecture to inspire and guide the model design in artificial intelligence (AI), e.g., deep neural networks. Heng Huang 0003, Lin Zhao 0004, Haixing Dai, Lu Zhang 0050, Xintao Hu, Dajiang Zhu, Tianming Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Individual Functional Network Abnormalities Mapping via Graph Representation-Based Neural Architecture Search
Qing Li 0027, Haixing Dai, Jinglei Lv, Lin Zhao 0004, Zhengliang Liu, Zihao Wu 0001, Xia Wu 0001, Claire Coles, Xiaoping Hu 0001, Tianming Liu 0001, Dajiang Zhu |
ADMA (3) | 2 |
| 2023 | Coarse-to-fine Knowledge Graph Domain Adaptation based on Distantly-supervised Iterative TrainingabstractThe knowledge graph (KG) is a highly needed basis to support the high-fidelity and high-interpretability modeling of various tasks in healthcare artificial intelligence. In this work, we focus on constructing an oncology knowledge graph that will be used in downstream cancer research and solution development. Modern supervised learning for knowledge graph construction requires a large amount of manually labeled data, which makes the process time-consuming and labor-intensive. Although there exists multiple research on named entity recognition and relation extraction based on distantly supervised learning, constructing a domain-specific knowledge graph from large collections of textual data without manual annotations is still an urgent problem to be solved. In response, we propose an integrated framework for adapting and re-learning knowledge graphs from a general domain (biomedical in our case) to a fine-defined domain (oncology). In this framework, we apply distant-supervision on cross-domain knowledge graph adaptation. Consequently, no manual data annotation is required to train the model. We introduce a novel iterative training strategy to facilitate the discovery of domain-specific named entities and triplets. Experimental results indicate that the proposed framework can perform domain adaptation and construction of knowledge graphs efficiently. Wenxiong Liao, Zhengliang Liu, Yiyang Zhang 0003, Fei Qi 0007, Siqi Ding, Hui Ren 0001, Zihao Wu 0001, Haixing Dai, Sheng Li 0001, Lingfei Wu 0001, Ninghao Liu 0001, Quanzheng Li, Tianming Liu 0001, Xiang Li 0001, Hongmin Cai |
BIBM | 9 |
| 2023 | A generic framework for embedding human brain function with temporally correlated autoencoder
Lin Zhao 0004, Zihao Wu 0001, Haixing Dai, Zhengliang Liu, Xintao Hu, Dajiang Zhu, Tianming Liu 0001 |
Medical Image Anal. | 3 |
| 2022 | Embedding Human Brain Function via Transformer
Lin Zhao 0004, Zihao Wu 0001, Haixing Dai, Zhengliang Liu, Dajiang Zhu, Tianming Liu 0001 |
MICCAI (1) | 3 |
| 2021 | Exploring the Functional Difference of Gyri/Sulci via Hierarchical Interpretable Autoencoder
Lin Zhao 0004, Haixing Dai, Xi Jiang 0001, Dajiang Zhu, Tianming Liu 0001 |
MICCAI (7) | 2 |