Zihao Wu 0001

dblp:185/2651-1 · DBLP profile ↗
← Back
23ranked-venue papers
1as first author
23since 2021 · last 2025
0000-0001-7483-6570ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ECHOPulse: ECG Controlled Echocardio-gram Video Generation
abstract
Echocardiography (ECHO) is essential for cardiac assessments, but its video quality and interpretation heavily relies on manual expertise, leading to inconsistent results from clinical and portable devices. ECHO video generation offers a solution by improving automated monitoring through synthetic data and generating high-quality videos from routine health data. However, existing models often face high computational costs, slow inference, and rely on complex conditional prompts that require experts' annotations. To address these challenges, we propose ECHOPulse, an ECG-conditioned ECHO video generation model. ECHOPulse introduces two key advancements: (1) it accelerates ECHO video generation by leveraging VQ-VAE tokenization and masked visual token modeling for fast decoding, and (2) it conditions on readily accessible ECG signals, which are highly coherent with ECHO videos, bypassing complex conditional prompts. To the best of our knowledge, this is the first work to use time-series prompts like ECG signals for ECHO video generation. ECHOPulse not only enables controllable synthetic ECHO data generation but also provides updated cardiac function information for disease monitoring and prediction beyond ECG alone. Evaluations on three public and private datasets demonstrate state-of-the-art performance in ECHO video generation across both qualitative and quantitative measures. Additionally, ECHOPulse can be easily generalized to other modality generation tasks, such as cardiac MRI, fMRI, and 3D CT generation. We will make the synthetic ECHO dataset, along with the code and model, publicly available upon acceptance.
Yiwei Li 0002, Sekeun Kim, Zihao Wu 0001, Hanqi Jiang, Yi Pan 0001, Pengfei Jin, Sifan Song, Xiaowei Yu 0001, Tianze Yang, Tianming Liu 0001, Quanzheng Li, Xiang Li 0001
ICLR3
2025 HARP: Human-Assisted Regrouping With Permutation Invariant Critic for Multi-Agent Reinforcement Learning
abstract
Human-in-the-loop reinforcement learning integrates human expertise to accelerate agent learning and provide critical guidance and feedback in complex fields. However, many existing approaches focus on single-agent tasks and require continuous human involvement during the training process, significantly increasing the human workload and limiting scalability. In this paper, we propose HARP (HumanAssisted Regrouping with Permutation Invariant Critic), a multi-agent reinforcement learning framework designed for group-oriented tasks. HARP integrates automatic agent regrouping with strategic human assistance during deployment, enabling and allowing non-experts to offer effective guidance with minimal intervention. During training, agents dynamically adjust their groupings to optimize collaborative task completion. When deployed, they actively seek human assistance and utilize the Permutation Invariant Group Critic to evaluate and refine human-proposed groupings, allowing non-expert users to contribute valuable suggestions. In multiple collaboration scenarios, our approach is able to leverage limited guidance from non-experts and enhance performance. The project can be found at https://github.com/huawen-hu/HARP.
Huawen Hu, Enze Shi, Chenxi Yue, Shuocun Yang, Zihao Wu 0001, Yiwei Li 0002, Tianyang Zhong, Tianming Liu 0001, Shu Zhang 0006
ICRA5
2025 Understanding LLMs: A comprehensive overview from training to inference
Tianle Han, Jiaming Tian, Yutong Zhang 0019, Jiaqi Wang 0010, Xiaohui Gao, Tianyang Zhong, Yi Pan 0001, Shaochen Xu, Zihao Wu 0001, Zhengliang Liu, Xin Zhang 0151, Shu Zhang 0001, Xintao Hu, Ning Qiang, Tianming Liu 0001, Bao Ge
Neurocomputing13
2025 Learning lifespan brain anatomical correspondence via cortical developmental continuity transfer
Lu Zhang 0050, Zhengwang Wu, Xiaowei Yu 0001, Yanjun Lyu, Zihao Wu 0001, Haixing Dai, Lin Zhao 0004, Li Wang 0026, Gang Li 0001, Xianqiao Wang, Tianming Liu 0001, Dajiang Zhu
Medical Image Anal.5
2025 AugGPT: Leveraging ChatGPT for Text Data Augmentation
abstract
Text data augmentation is an effective strategy for overcoming the challenge of limited sample sizes in many natural language processing (NLP) tasks. This challenge is especially prominent in the few-shot learning (FSL) scenario, where the data in the target domain is generally much scarcer and of lowered quality. A natural and widely used strategy to mitigate such challenges is to perform data augmentation to better capture data invariance and increase the sample size. However, current text data augmentation methods either can’t ensure the correct labeling of the generated data (lacking faithfulness), or can’t ensure sufficient diversity in the generated data (lacking compactness), or both. Inspired by the recent success of large language models (LLM), especially the development of ChatGPT, we propose a text data augmentation approach based on ChatGPT (named ”AugGPT”). AugGPT rephrases each sentence in the training samples into multiple conceptually similar but semantically different samples. The augmented samples can then be used in downstream model training. Experiment results on multiple few-shot learning text classification tasks show the superior performance of the proposed AugGPT approach over state-of-the-art text data augmentation methods in terms of testing accuracy and distribution of the augmented samples.
Haixing Dai, Zhengliang Liu, Wenxiong Liao, Zihao Wu 0001, Lin Zhao 0004, Shaochen Xu, Fang Zeng, Wei Liu 0146, Ninghao Liu 0001, Sheng Li 0001, Dajiang Zhu, Hongmin Cai, Lichao Sun 0001, Quanzheng Li, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001
IEEE Trans. Big Data6
2025 Exploring New Frontiers in Agricultural NLP: Investigating the Potential of Large Language Models for Food Applications
abstract
This paper explores new frontiers in agricultural natural language processing (NLP) by investigating the effectiveness of food-related text corpora for pretraining transformer-based language models. Specifically, we focus on semantic matching, establishing mappings between food descriptions and nutrition data through fine-tuning AgriBERT with the FoodOn ontology. Our work introduces an expanded comparison with state-of-the-art language models such as GPT-4, Mistral-large, Claude 3 Sonnet, and Gemini 1.0 Ultra. This exploratory investigation, rather than a direct comparison, aims to understand how AgriBERT, a domain-specific, fine-tuned, open-source model, complements the broad knowledge and generative abilities of these advanced LLMs in addressing the unique challenges of the agricultural sector. We also experiment with other applications, such as cuisine prediction from ingredients, expanding our research to include various NLP tasks beyond semantic matching. Overall, this paper underscores the potential of integrating domain-specific models like AgriBERT with advanced LLMs to enhance the performance and applicability of agricultural NLP applications.
Saed Rezayi, Zhengliang Liu, Zihao Wu 0001, Chandra Dhakal, Bao Ge, Haixing Dai, Gengchen Mai, Ninghao Liu 0001, Chen Zhen, Tianming Liu 0001, Sheng Li 0001
IEEE Trans. Big Data3
2025 Exploring the Trade-Offs: Unified Large Language Models vs Local Fine-Tuned Models for Highly-Specific Radiology NLI Task
abstract
Recently, ChatGPT and GPT-4 have emerged and gained immense global attention due to their unparalleled performance in language processing. Despite demonstrating impressive capability in various open-domain tasks, their adequacy in highly specific fields like radiology remains untested. Radiology presents unique linguistic phenomena distinct from open-domain data due to its specificity and complexity. Assessing the performance of large language models (LLMs) in such specific domains is crucial not only for a thorough evaluation of their overall performance but also for providing valuable insights into future model design directions: whether model design should be generic or domain-specific. To this end, in this study, we evaluate the performance of ChatGPT/GPT-4 on a radiology natural language inference (NLI) task and compare it to other models fine-tuned specifically on task-related data samples. We also conduct a comprehensive investigation on ChatGPT/GPT-4’s reasoning ability by introducing varying levels of inference difficulty. Our results show that 1) ChatGPT and GPT-4 outperform other LLMs in the radiology NLI task and 2) other specifically fine-tuned Bert-based models require significant amounts of data samples to achieve comparable performance to ChatGPT/GPT-4. These findings not only demonstrate the feasibility and promise of constructing a generic model capable of addressing various tasks across different domains, but also highlight several key factors crucial for developing a unified model, particularly in a medical context, paving the way for future artificial general intelligence (AGI) systems. We release our code and data to the research community.
Zihao Wu 0001, Lu Zhang 0050, Xiaowei Yu 0001, Zhengliang Liu, Lin Zhao 0004, Yiwei Li 0002, Haixing Dai, Chong Ma 0004, Gang Li 0001, Wei Liu 0146, Quanzheng Li, Dinggang Shen, Xiang Li 0001, Dajiang Zhu, Tianming Liu 0001
IEEE Trans. Big Data1
2025 Core-Periphery Multi-Modality Feature Alignment for Zero-Shot Medical Image Analysis
abstract
Multi-modality learning, exemplified by the language-image pair pre-trained CLIP model, has demonstrated remarkable performance in enhancing zero-shot capabilities and has gained significant attention recently. However, simply applying language-image pre-trained CLIP to medical image analysis encounters substantial domain shifts, resulting in severe performance degradation due to inherent disparities between natural (non-medical) and medical image characteristics. To address this challenge and uphold or even enhance CLIP's zero-shot capability in medical image analysis, we develop a novel approach, Core-Periphery feature alignment for CLIP (CP-CLIP), to model medical images and corresponding clinical text jointly. To achieve this, we design an auxiliary neural network whose structure is organized by the core-periphery (CP) principle. This auxiliary CP network not only aligns medical image and text features into a unified latent space more efficiently but also ensures alignment driven by principles of brain network organization. In this way, our approach effectively mitigates and further enhances CLIP's zero-shot performance in medical image analysis. More importantly, the proposed CP-CLIP exhibits excellent explanatory capability, enabling the automatic identification of critical disease-related regions in clinical analysis. Extensive experiments and evaluation across five public datasets covering different diseases underscore the superiority of our CP-CLIP in zero-shot medical image prediction and critical features detection, showing its promising utility in multimodal feature alignment in current medical applications.
Xiaowei Yu 0001, Lu Zhang 0050, Zihao Wu 0001, Dajiang Zhu
IEEE Trans. Medical Imaging3
2025 A Unified and Biologically Plausible Relational Graph Representation of Vision Transformers
abstract
Vision transformer (ViT) and its variants have achieved remarkable success in various tasks. The key characteristic of these ViT models is to adopt different aggregation strategies of spatial patch information within the artificial neural networks (ANNs). However, there is still a key lack of unified representation of different ViT architectures for systematic understanding and assessment of model representation performance. Moreover, how those well-performing ViT ANNs are similar to real biological neural networks (BNNs) is largely unexplored. To answer these fundamental questions, we, for the first time, propose a unified and biologically plausible relational graph representation of ViT models. Specifically, the proposed relational graph representation consists of two key subgraphs: an aggregation graph and an affine graph. The former considers ViT tokens as nodes and describes their spatial interaction, while the latter regards network channels as nodes and reflects the information communication between channels. Using this unified relational graph representation, we found that: 1) model performance was closely related to graph measures; 2) the proposed relational graph representation of ViT has high similarity with real BNNs; and 3) there was a further improvement in model performance when training with a superior model to constrain the aggregation graph.
Yuzhong Chen 0002, Zhenxiang Xiao, Lin Zhao 0004, Lu Zhang 0050, Zihao Wu 0001, Dajiang Zhu, Dezhong Yao 0001, Xintao Hu, Tianming Liu 0001, Xi Jiang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 Mask-Guided Vision Transformer for Few-Shot Learning
abstract
Learning with little data is challenging but often inevitable in various application scenarios where the labeled data are limited and costly. Recently, few-shot learning (FSL) gained increasing attention because of its generalizability of prior knowledge to new tasks that contain only a few samples. However, for data-intensive models such as vision transformer (ViT), current fine-tuning-based FSL approaches are inefficient in knowledge generalization and, thus, degenerate the downstream task performances. In this article, we propose a novel mask-guided ViT (MG-ViT) to achieve an effective and efficient FSL on the ViT model. The key idea is to apply a mask on image patches to screen out the task-irrelevant ones and to guide the ViT focusing on task-relevant and discriminative patches during FSL. Particularly, MG-ViT only introduces an additional mask operation and a residual connection, enabling the inheritance of parameters from pretrained ViT without any other cost. To optimally select representative few-shot samples, we also include an active learning-based sample selection method to further improve the generalizability of MG-ViT-based FSL. We evaluate the proposed MG-ViT on classification, object detection, and segmentation tasks using gradient-weighted class activation mapping (Grad-CAM) to generate masks. The experimental results show that the MG-ViT model significantly improves the performance and efficiency compared with general fine-tuning-based ViT and ResNet models, providing novel insights and a concrete approach toward generalizing data-intensive and large-scale deep learning models for FSL.
Yuzhong Chen 0002, Zhenxiang Xiao, Yi Pan 0001, Lin Zhao 0004, Haixing Dai, Zihao Wu 0001, Changhe Li, Changying Li, Dajiang Zhu, Tianming Liu 0001, Xi Jiang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
abstract
Advancements in foundation models have made it possible to conduct applications in various downstream tasks. Especially, the new era has witnessed a remarkable capability to extend large language models (LLMs) for tackling tasks of 3-D scene understanding. Current methods rely heavily on 3-D point clouds, but the 3-D point cloud reconstruction of an indoor scene often results in information loss. Some textureless planes or repetitive patterns are prone to omission and manifest as voids within the reconstructed 3-D point clouds. Besides, objects with complex structures tend to introduce distortion of details caused by misalignments between the captured images and the dense reconstructed point clouds. The 2-D multiview images present visual consistency with 3-D point clouds and provide more detailed representations of scene components, which can naturally compensate for these deficiencies. Based on these insights, we propose Argus, a novel 3-D multimodal framework that leverages multiview images for enhanced 3-D scene understanding with LLMs. In general, Argus can be treated as a 3-D large multimodal foundation model (3D-LMM) since it takes various modalities as input (text instructions, 2-D multiview images, and 3-D point clouds) and expands the capability of LLMs to tackle 3-D tasks. Argus involves fusing and integrating multiview images and camera poses into view-as-scene features, which interact with the 3-D features to create comprehensive and detailed 3-D-aware scene embeddings. Our approach compensates for the information loss while reconstructing 3-D point clouds and helps LLMs better understand the 3-D world. Extensive experiments demonstrate that our method outperforms existing 3D-LMMs in various downstream tasks.
Yifan Xu 0028, Hanqi Jiang, Ruifei Ma, Yiwei Li 0002, Zihao Wu 0001, Zeju Li, Xiangde Liu
IEEE Trans. Neural Networks Learn. Syst.7
2024 CP-CLIP: Core-Periphery Feature Alignment CLIP for Zero-Shot Medical Image Analysis
Xiaowei Yu 0001, Zihao Wu 0001, Lu Zhang 0050, Jing Zhang 0010, Yanjun Lyu, Dajiang Zhu
MICCAI (3)2
2024 Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
abstract
In the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly aligned from the data, without considering the explicit relationships in the medical context. This data-reliance may lead to low generalization of the learned alignment relationships. In this work, we propose the Eye-gaze Guided Multi-modal Alignment (EGMA) framework to harness eye-gaze data for better alignment of medical visual and textual features. We explore the natural auxiliary role of radiologists' eye-gaze data in aligning medical images and text, and introduce a novel approach by using eye-gaze data, collected synchronously by radiologists during diagnostic evaluations. We conduct downstream tasks of image classification and image-text retrieval on four medical datasets, where EGMA achieved state-of-the-art performance and stronger generalization across different datasets. Additionally, we explore the impact of varying amounts of eye-gaze data on model performance, highlighting the feasibility and utility of integrating this auxiliary data into multi-modal alignment framework.
Chong Ma 0004, Hanqi Jiang, Wenting Chen, Yiwei Li 0002, Zihao Wu 0001, Xiaowei Yu 0001, Zhengliang Liu, Lei Guo 0002, Dajiang Zhu, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001
NeurIPS5
2024 Real-time Core-Periphery Guided ViT with Smart Data Layout Selection on Mobile Devices
abstract
Mobile devices have become essential enablers for AI applications, particularly in scenarios that require real-time performance. Vision Transformer (ViT) has become a fundamental cornerstone in this regard due to its high accuracy. Recent efforts have been dedicated to developing various transformer architectures that offer im- proved accuracy while reducing the computational requirements. However, existing research primarily focuses on reducing the theoretical computational complexity through methods such as local attention and model pruning, rather than considering realistic performance on mobile hardware. Although these optimizations reduce computational demands, they either introduce additional overheads related to data transformation (e.g., Reshape and Transpose) or irregular computation/data-access patterns. These result in significant overhead on mobile devices due to their limited bandwidth, which even makes the latency worse than vanilla ViT on mobile. In this paper, we present ECP-ViT, a real-time framework that employs the core-periphery principle inspired by the brain functional networks to guide self-attention in ViTs and enable the deployment of ViT models on smartphones. We identify the main bottleneck in transformer structures caused by data transformation and propose a hardware-friendly core-periphery guided self-attention to decrease computation demands. Additionally, we design the system optimizations for intensive data transformation in pruned models. ECP-ViT, with the proposed algorithm-system co-optimizations, achieves a speedup of 4.6× to 26.9× on mobile GPUs across four datasets: STL-10, CIFAR100, TinyImageNet, and ImageNet.
Zhihao Shu, Xiaowei Yu 0001, Zihao Wu 0001, Wenqi Jia 0003, Yinchen Shi, Miao Yin, Tianming Liu 0001, Dajiang Zhu, Wei Niu 0002
NeurIPS3
2024 Mask-guided BERT for few-shot text classification
Wenxiong Liao, Zhengliang Liu, Haixing Dai, Zihao Wu 0001, Yiyang Zhang 0003, Yuzhong Chen 0002, Xi Jiang 0001, Dajiang Zhu, Sheng Li 0001, Wei Liu 0146, Tianming Liu 0001, Quanzheng Li, Hongmin Cai, Xiang Li 0001
Neurocomputing4
2024 CE-GAN: Community Evolutionary Generative Adversarial Network for Alzheimer's Disease Risk Prediction
abstract
In the studies of neurodegenerative diseases such as Alzheimer's Disease (AD), researchers often focus on the associations among multi-omics pathogeny based on imaging genetics data. However, current studies overlook the communities in brain networks, leading to inaccurate models of disease development. This paper explores the developmental patterns of AD from the perspective of community evolution. We first establish a mathematical model to describe functional degeneration in the brain as the community evolution driven by entropy information propagation. Next, we propose an interpretable Community Evolutionary Generative Adversarial Network (CE-GAN) to predict disease risk. In the generator of CE-GAN, community evolutionary convolutions are designed to capture the evolutionary patterns of AD. The experiments are conducted using functional magnetic resonance imaging (fMRI) data and single nucleotide polymorphism (SNP) data. CE-GAN achieves 91.67% accuracy and 91.83% area under curve (AUC) in AD risk prediction tasks, surpassing advanced methods on the same dataset. In addition, we validated the effectiveness of CE-GAN for pathogeny extraction. The source code of this work is available at https://github.com/fmri123456/CE-GAN.
Xia-an Bi, Zicheng Yang, YangJun Huang, Zhao-Xu Xing, Luyun Xu, Zihao Wu 0001, Zhengliang Liu, Xiang Li 0001, Tianming Liu 0001
IEEE Trans. Medical Imaging6
2023 Coupling Artificial Neurons in BERT and Biological Neurons in the Human Brain
abstract
Linking computational natural language processing (NLP) models and neural responses to language in the human brain on the one hand facilitates the effort towards disentangling the neural representations underpinning language perception, on the other hand provides neurolinguistics evidence to evaluate and improve NLP models. Mappings of an NLP model’s representations of and the brain activities evoked by linguistic input are typically deployed to reveal this symbiosis. However, two critical problems limit its advancement: 1) The model’s representations (artificial neurons, ANs) rely on layer-level embeddings and thus lack fine-granularity; 2) The brain activities (biological neurons, BNs) are limited to neural recordings of isolated cortical unit (i.e., voxel/region) and thus lack integrations and interactions among brain functions. To address those problems, in this study, we 1) define ANs with fine-granularity in transformer-based NLP models (BERT in this study) and measure their temporal activations to input text sequences; 2) define BNs as functional brain networks (FBNs) extracted from functional magnetic resonance imaging (fMRI) data to capture functional interactions in the brain; 3) couple ANs and BNs by maximizing the synchronization of their temporal activations. Our experimental results demonstrate 1) The activations of ANs and BNs are significantly synchronized; 2) the ANs carry meaningful linguistic/semantic information and anchor to their BN signatures; 3) the anchored BNs are interpretable in a neurolinguistic context. Overall, our study introduces a novel, general, and effective framework to link transformer-based NLP models and neural activities in response to language and may provide novel insights for future studies such as brain-inspired evaluation and development of NLP models.
Mengyue Zhou, Gaosheng Shi, Lin Zhao 0004, Zihao Wu 0001, Tianming Liu 0001, Xintao Hu
AAAI6
2023 Individual Functional Network Abnormalities Mapping via Graph Representation-Based Neural Architecture Search
Qing Li 0027, Haixing Dai, Jinglei Lv, Lin Zhao 0004, Zhengliang Liu, Zihao Wu 0001, Xia Wu 0001, Claire Coles, Xiaoping Hu 0001, Tianming Liu 0001, Dajiang Zhu
ADMA (3)6
2023 Coarse-to-fine Knowledge Graph Domain Adaptation based on Distantly-supervised Iterative Training
abstract
The knowledge graph (KG) is a highly needed basis to support the high-fidelity and high-interpretability modeling of various tasks in healthcare artificial intelligence. In this work, we focus on constructing an oncology knowledge graph that will be used in downstream cancer research and solution development. Modern supervised learning for knowledge graph construction requires a large amount of manually labeled data, which makes the process time-consuming and labor-intensive. Although there exists multiple research on named entity recognition and relation extraction based on distantly supervised learning, constructing a domain-specific knowledge graph from large collections of textual data without manual annotations is still an urgent problem to be solved. In response, we propose an integrated framework for adapting and re-learning knowledge graphs from a general domain (biomedical in our case) to a fine-defined domain (oncology). In this framework, we apply distant-supervision on cross-domain knowledge graph adaptation. Consequently, no manual data annotation is required to train the model. We introduce a novel iterative training strategy to facilitate the discovery of domain-specific named entities and triplets. Experimental results indicate that the proposed framework can perform domain adaptation and construction of knowledge graphs efficiently.
Wenxiong Liao, Zhengliang Liu, Yiyang Zhang 0003, Fei Qi 0007, Siqi Ding, Hui Ren 0001, Zihao Wu 0001, Haixing Dai, Sheng Li 0001, Lingfei Wu 0001, Ninghao Liu 0001, Quanzheng Li, Tianming Liu 0001, Xiang Li 0001, Hongmin Cai
BIBM8
2023 An automatic classifier for monitoring applied behaviors of cage-free laying hens with deep learning
Xiao Yang 0027, Ramesh Bahadur Bist, Sachin Subedi, Zihao Wu 0001, Tianming Liu 0001, Lilong Chai
Eng. Appl. Artif. Intell.4
2023 A generic framework for embedding human brain function with temporally correlated autoencoder
Lin Zhao 0004, Zihao Wu 0001, Haixing Dai, Zhengliang Liu, Xintao Hu, Dajiang Zhu, Tianming Liu 0001
Medical Image Anal.2
2022 AgriBERT: Knowledge-Infused Agricultural Language Models for Matching Food and Nutrition
abstract
Pretraining domain-specific language models remains an important challenge which limits their applicability in various areas such as agriculture. This paper investigates the effectiveness of leveraging food related text corpora (e.g., food and agricultural literature) in pretraining transformer-based language models. We evaluate our trained language model, called AgriBERT, on the task of semantic matching, i.e., establishing mapping between food descriptions and nutrition data, which is a long-standing challenge in the agricultural domain. In particular, we formulate the task as an answer selection problem, fine-tune the trained language model with the help of an external source of knowledge (e.g., FoodOn ontology), and establish a baseline for this task. The experimental results reveal that our language model substantially outperforms other language models and baselines in the task of matching food description and nutrition.
Saed Rezayi, Zhengliang Liu, Zihao Wu 0001, Chandra Dhakal, Bao Ge, Chen Zhen, Tianming Liu 0001, Sheng Li 0001
IJCAI3
2022 Embedding Human Brain Function via Transformer
Lin Zhao 0004, Zihao Wu 0001, Haixing Dai, Zhengliang Liu, Dajiang Zhu, Tianming Liu 0001
MICCAI (1)2