VLDB 2026 Research / reviewers in the wild / expert
Tian Bai 0002
dblp:05/6070-2
· DBLP profile ↗
43ranked-venue papers
13as first author
35since 2021 · last 2026
0000-0001-8060-4725ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 8 first-author · 13 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aware Distillation for Robust Vision-Language Tracking Under Linguistic SparsityabstractVision-language object tracking overcomes the limitations of relying solely on visual features by leveraging language descriptions of objects to provide cross-modal semantic information, thereby enhancing model robustness in complex scenarios. However, most existing high-performance vision-language trackers are trained jointly on pure visual data and vision-language multimodal data. Due to the relative sparsity of language annotations in the data, the trackers tend to prioritize the localization role of visual features, diminishing the model's attention to language information. To mitigate this issue, we propose a novel vision-language tracker: Aware Distillation for Robust Vision-Language Tracking under Linguistic Sparsity (ADTrack). We introduce a knowledge distillation framework employing a knowledge-rich teacher model and a lightweight student model to establish modality correlations between vision and language, enabling efficient modeling between visual information and language descriptions. Specifically, our lightweight student module simultaneously distills language encoding capabilities from large language models through teacher-guided learning on input language, while performing target-aware perception on template images using language descriptions to generate more effective template features for subsequent visual extraction. Furthermore, to ensure perceptual robustness in linguistically sparse scenarios, we simulate language-deficient conditions during training and employ contrastive learning to enhance model adaptability. Extensive experiments demonstrate that ADTrack reduces parameters by over 50% while achieving state-of-the-art (SOTA) performance and speed on vision-language tracking benchmarks, including LaSOT, LaSOText, TNL2K, OTB-Lang and MGIT. Guangtong Zhang, Bineng Zhong 0001, Shirui Yang, Tian Bai 0002 |
AAAI | 5 |
| 2026 | Selective distillation of language tokens for redundancy suppression in vision-language tracking
Tian Bai 0002, Shirui Yang, Guangtong Zhang |
Expert Syst. Appl. | 1 |
| 2026 | CogECI: Context Grounded Document-level Event Causality Identification via Large Language Models
Zefan Zhang, Xumeng Zhang, Shijie Jiang, Tian Bai 0002 |
Knowl. Based Syst. | 4 |
| 2026 | TDP-DETR: Temporal dynamics perception framework for video moment retrieval and highlight detectionabstractVideo Moment Retrieval (VMR) and Highlight Detection (HD) aim to localize query-relevant temporal segments and evaluate clip-level saliency within untrimmed videos. Accurate temporal boundary perception is essential for VMR and HD. While current models have made significant progress, they still struggle to achieve precise action semantic alignment with temporally dynamic video content and are prone to boundary perception bias when action-related visual semantic cues experience fluctuations in specific frames. In this paper, we propose a Temporal Dynamics Perception DEtection TRansformer (TDP-DETR) that models action temporal dynamics from two complementary perspectives: temporal persistence and temporal progression. For temporal persistence, we introduce a dynamic masking strategy for action duration-aware temporal modeling, enabling the model to infer action persistence from query semantics and incorporate it as a temporal prior for boundary prediction. For temporal progression, we design an action state difference perception module that captures frame-to-frame action state variations, allowing the model to perceive action progression speed and thereby improve anticipation of action boundaries. Extensive experiments on three MR/HD benchmarks demonstrate that our method consistently outperforms existing state-of-the-art approaches. Our code will be public soon. Huilin An, Zefan Zhang, Shijie Jiang, Kehua Zhu, Tian Bai 0002 |
Neural Networks | 5 |
| 2026 | Dual-level dynamic heterogeneous graph network for video question answering
Zefan Zhang, Tian Bai 0002 |
Neural Networks | 4 |
| 2026 | SSGraphDTI: A Drug-Target Interaction Prediction Method Integrated Structural and Dynamic Systemic Biology AttributesabstractDrug-Target Interaction (DTI) is a crucial aspect of pharmaceutical development. However, biochemical experiments are prohibitively expensive to identify these interactions on a large scale, while the computational approach is still on the way to making a highly reliable prediction. For the purpose of promoting prediction accuracy, drug-related molecular networks are gradually introduced to this task to furnish valuable information. We hypothesized that integrating structural and systemic biological attributes could effectively enhance the performance of DTI prediction and proposed a novel DTI prediction model, SSGraphDTI, which integrated two aforementioned attributes. Specifically, the structural attributes of drugs and targets are extracted using independent convolutional neural network based models from the Simplified Molecular Input Line Entry System of drugs and the amino acid sequences of targets, respectively. Meanwhile, the systemic biological attributes of drug-target pairs are obtained through graph representation learning on the dynamically constructed heterogeneous drug-target interaction network. SSGraphDTI was meticulously trained and rigorously tested on the benchmark Dataset_DrugBank, achieving an improvement of approximately 1.0% across five metrics compared to recent comparable methods. These results underscore the potential of combining both structural and systemic information for accurate DTI prediction. Benefiting from the fact that the input consists solely of structural data without requiring interaction information, the model effectively addresses the "cold-start problem" in drug discovery. Furthermore, by extracting systemic attributes directly from the dynamically constructed DTI networks, the model maintains strong predictive performance even when data is limited. Haotian Guan, Tian Bai 0002, Jingtong Zhao, Han Wang 0028 |
IEEE J. Biomed. Health Informatics | 2 |
| 2026 | DAMON: Difference-Aware Medical Visual Question Answering via Multimodal Large Language ModelabstractDifference-aware Medical Visual Question Answering (MVQA) aims to answer questions regarding disease-related content and the visual differences between the paired medical images, which is crucial for assessing disease progression and guiding further treatment planning. Although current medical Multimodal Large Language Models (MLLMs) have shown promising results in MVQA, they still exhibit poor generalization performance in difference-aware MVQA due to two key challenges. Firstly, existing difference-aware MVQA datasets are biased toward temporal variations of individual diseases, limiting their ability to model multi-disease coexistence and overlapping symptoms in real-world clinical scenarios. Secondly, disease-level semantic alignment becomes more challenging with multi-image inputs, as they introduce more redundant and interfering visual features. To address the first challenge, we introduce DAMON-QA, a large-scale difference-aware MVQA dataset designed to support visual difference analysis across multiple diseases. Leveraging this dataset, we train MLLMs and propose a Difference-Aware Medical visual questiON answering (DAMON) model. To tackle the second challenge, we further propose a Disease-driven Prompt Module (DPM) to identify the relevant diseases and guide the disease difference analysis process. Experiments on MIMIC-Diff-VQA show that our DAMON model achieves state-of-the-art (SOTA) performance. Zefan Zhang, Ruihong Zhao, Tian Bai 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | Depth-Assisted Camouflaged Object Segmentation via Frequency-Domain Fusion and High-Order InteractionabstractRecently, some studies have introduced depth cues to solve camouflaged object segmentation (COS) tasks and significantly improve segmentation performance. However, current methods still have two limitations: 1) they are confined to first-order or second-order interaction modeling; 2) they neglect the frequency-domain complementary characteristics between RGB and depth modalities. In this work, we propose a Frequency-domain Fusion and High-order Interaction framework, named F$^{2}$HI, to alleviate the above limitations. Specifically, F$^{2}$HI consists of two key components: Frequency Domain Interaction Fusion (FDIF) and the High-order Interaction Unit (HOIU). The FDIF decouples the frequency information of RGB and depth modalities into high-frequency and low-frequency components, utilizing the dominant frequency components of one modality to enhance the weaker frequency components of the other modality, thereby achieving complementary enhancement in the frequency domain. Moreover, it employs cross-attention to facilitate global multimodal interaction. The HOIU employs cascaded self-attention to parse cross-region features generated by multi-object camouflage scenarios, thereby achieving high-order feature interaction. Compared with 27 state-of-the-art COS methods, F$^{2}$HI achieves competitive performance on four mainstream COS benchmarks. Fuming Sun, Tian Bai 0002 |
IEEE Trans. Multim. | 4 |
| 2026 | RWKV-Inspired Multi-Modal Relation Modeling for Vision-Language TrackingabstractVision-language object tracking can provide more state representations for targets by introducing the language modality, achieving more robust tracking and localization. Therefore, designing multi-modal interactions to achieve feature alignment between vision and language has been one of the research hotspots. However, existing multi-modal interaction methods face two key issues: on the one hand, they lack effective exploration of modeling the relationship between the contextual information of language sequences and visual features; on the other hand, the introduction of modalities leads to increased computational time costs in multi-modal interactions, which severely affects the real-time performance of vision-language tracking algorithms. To address these challenges, we propose a vision-language tracking framework called RWKV-Inspired Multi-modal Relation Modeling for Vision-Language Tracking (RrmTrack). We introduce a novel modality interaction method specific to vision-language object tracking based on RWKV, providing customized interaction for different modalities in vision-language tracking and effectively reducing the computational time cost of cross-modal interaction. Specifically, this method uses a time mixing module to model the relationship between language information and image features, and a channel mixing module to facilitate information interaction between images. By combining parallelized training with a linear attention mechanism and efficient RNN inference, it enables accurate and fast target localization in vision-language tracking. Additionally, we propose a novel feature extraction structure that integrates Siamese and One-stream architectures. An information restoration module is designed to reduce the information interference introduced by the search image to the template image during interaction. RrmTrack achieves state-of-the-art results and speed on multiple vision-language object tracking benchmarks, including TNL2k, LaSOT, OTB-Lang, LaSOText, and MGIT. Guangtong Zhang, Bineng Zhong 0001, Yuhao Mu, Tian Bai 0002 |
IEEE Trans. Multim. | 5 |
| 2025 | MP: Endowing Large Language Models with Lateral ThinkingabstractThe recent studies show that Large Language Models (LLMs) often fall short in tasks demanding creative, lateral thinking due to lacking a clear awareness of their own reasoning processes. To cope with this issue, we propose a novel metacognitive prompting method (titled as MP) by mimicking human metacognition. Through integrating metacognitive principles, MP endows LLMs with lateral thinking ability, thereby enhancing their abilities to strategize, monitor, and reflect on their responses when dealing with creative tasks. The experimental results with five base LLMs across three lateral thinking datasets demonstrate that: All LLMs armed with MP consistently outperform the representative baseline methods. For example, MP demonstrates superior performance over CoT prompting across Sentence Puzzle (+5.00%), Word Puzzle (+10.07%), BiRdQA (+6.48%), and RiddleSense (+2.65%) with GPT-3.5-turbo model. In particular, the deployment of MP with GPT-4 achieves significant performance improvements that even surpass human performance on BRAINTEASER benchmark, demonstrating the transformative potential of MP in enhancing the creative problem-solving abilities of LLMs. Tian Bai 0002, Yongwang Cao, Hai-Tao Yu 0003 |
AAAI | 1 |
| 2025 | Prototype-Guided Multimodal Relation Extraction based on Entity AttributesabstractMultimodal Relation Extraction (MRE) aims to predict relations between head and tail entities based on the context of sentence-image pairs. Most existing MRE methods progressively incorporate textual and visual inputs to dominate the learning process, assuming both contribute significantly to the task. However, the diverse visual appearances and text with ambiguous semantics contain less-informative contexts for the corresponding relation. To tackle these challenges, we highlight the importance of semantically invariant entity attributes that encompass fine-grained categories. Towards this, we propose a novel Prototype-Guided Multimodal Relation Extraction (PG-MRE) framework based on Entity Attributes. Specifically, we first generate detailed entity explanations using Large Language Models (LLMs) to supplement the attribute semantics. Then, the Attribute Prototype Module (APM) refines attribute categories and condenses scattered entity attribute features into cluster-level prototypes. Furthermore, prototype-aligned attribute features guide diverse visual appearance features to produce compact and distinctive multimodal representations in the Relation Prototype Module (RPM). Extensive experiments demonstrate that our method gains superior relation classification capability (especially in scenarios involving various unseen entities), achieving new state-of-the-art performances on MNRE dataset. Zefan Zhang, Tian Bai 0002 |
AAAI | 4 |
| 2025 | HyperDTI-Lite: Hyperbolic Geometry for Enhanced Drug-Target Interaction Prediction via Heterogeneous Feature FusionabstractAccurate prediction of drug-target interactions (DTIs) is crucial for accelerating drug discovery and repurposing efforts. However, existing computational models face significant challenges in effectively integrating heterogeneous features and capturing the complex hierarchical relationships inherent in biological systems. To address these limitations, we propose HyperDTI-Lite, a novel and efficient DTI prediction framework that integrates homologous heterogeneous features, including structural features and physicochemical properties, with semantic embeddings obtained from pretrained language models. By projecting these diverse features into a unified hyperbolic embedding space through a hyperbolic multilayer perceptron, HyperDTILite effectively captures complex dependencies between drug and target features. Extensive experiments on the benchmark DrugBank dataset demonstrate that HyperDTI-Lite consistently outperforms state-of-the-art methods across multiple evaluation metrics, achieving superior predictive accuracy. Comprehensive ablation studies validate the effectiveness of each key component in our framework. This work highlights the advantages of hyperbolic geometry for modeling complex biological data and provides a robust foundation for future drug discovery research. Haotian Guan, Tian Bai 0002, Chuande Yang, Jingtong Zhao, Han Wang 0028 |
BIBM | 2 |
| 2025 | Seeing the Unseen: A Semantic Alignment and Context-Aware Prompt Framework for Open-Vocabulary Camouflaged Object Segmentation
Tian Bai 0002, Jing Sun 0012, Fuming Sun |
ICCV | 2 |
| 2025 | Video-Level Multimodal Relation Extraction with Event-Entity Semantic ConsistencyabstractPrevious research on Multimodal Relation Extraction (MRE) has primarily focused on identifying textual relations enhanced by static visual clues from images, benefiting fields such as multimedia analysis and knowledge graphs. With the rapid rise of video content on social media platforms, Multimodal Relation Extraction (MRE) systems face new challenges. To bridge this gap, we introduce Video-level Multimodal Relation Extraction (VMRE), a novel task aimed at extracting relational facts from videos. To advance this research, we present Vid-MRE, a new dataset containing 32 relation types and 12,402 multimodal relational facts, annotated across 3,970 pairs of textual news titles and corresponding videos. Since this task demands precise event and entity grounding to filter out excessive noise in the video, we propose an Event-Entity Semantic Consistency Network (E2SCN) to capture relational clues in the video effectively. Experimental results demonstrate that incorporating video content into the model significantly improves relation identification performance but also introduces more noise. Our E2SCN method effectively reduces the noise, enhancing fine-grained multimodal event and entity alignments while achieving state-of-the-art (SOTA) performance. Zefan Zhang, Kailong Suo, Tian Bai 0002 |
ACM Multimedia | 5 |
| 2025 | Prompt-guided orthogonal multimodal fusion for cancer survival prediction
Lan Huang 0002, Shuyu Guo, Tian Bai 0002, Ruihong Zhao, Ke Tao |
Inf. Sci. | 3 |
| 2025 | Unsupervised Adversarial Domain Adaptation with Hierarchical Semantic Consistency for Cross-Modal Nuclei Detection
Shuyu Guo, Lan Huang 0002, Yu-Hao Mu, Tian Bai 0002 |
J. Comput. Sci. Technol. | 4 |
| 2025 | Prompting visual dialog with implicit logical knowledge
Zefan Zhang, Tian Bai 0002 |
Knowl. Inf. Syst. | 4 |
| 2025 | ESNet: An Efficient Skeleton-guided Network for camouflaged object detection
Tian Bai 0002, Fuming Sun |
Knowl. Based Syst. | 2 |
| 2025 | Bio-inspired two-stage network for efficient RGB-D salient object detection
Tian Bai 0002, Fuming Sun |
Neural Networks | 2 |
| 2024 | Robust Federated Semi-Supervised Learning for Medical Image Classification via Pseudo-Label FilteringabstractFederated learning (FL) enables collaborative model training across multiple medical institutions to ensure data security. However, due to the variations in medical imaging equipment and regions at different medical institutions, FL methods usually suffer from insufficient data annotations and irrelevant noise within private datasets. To address these issues, a robust federated semi-supervised learning method via pseudo-label filtering (PFRFed) is introduced to utilize unlabeled data while mitigating the impact of noise data. Compared with existing federated semi-supervised learning methods, we propose a pseudo-label filtering mechanism with double dynamic thresholds, which allows the model to adopt more unlabeled data by adjusting the confidence and entropy thresholds at each stage of model training. Moreover, to reduce the degradation caused by noise data in private datasets from different clients, a noise-tolerant loss function and a grouping aggregation method based on the local model similarity are employed. The comparative experiments demonstrate the effectiveness of PFRFed, which has achieved the best classification accuracy of 95.20% and 88.72% on two public medical datasets. Also, PFRFed exhibits heightened resilience to variations in noisy data ratio and labeled data ratio, reaffirming its versatility and robustness. Shuyu Guo, Mingzhu Zhu, Tian Bai 0002 |
BIBM | 4 |
| 2024 | 2D-3D Feature Co-Embedding Network with Sparse Annotation for 3D Medical Image SegmentationabstractSupervised methods on 3D medical image segmentation need large amounts of annotated data, but annotating is time-consuming. Also, existing 3D segmentation methods capture more global structural information but overlook local detailed features, which negatively impacts the segmentation of small tissues. In this paper, we propose a novel weakly-supervised 2D-3D Feature Co-Embedding Network (2D-3D CoENet) that includes 2D and 3D encoding layers, simultaneously extracting 2D local detailed and 3D global structural features. To reduce annotation costs, we use fewer labeled slices as ground truth and pseudo-labels are generated by 2D-3D CoENet for other slices. Additionally, multi-view learning is introduced to capture more 2D local detailed information, and a Multi-view Semantic Consistency loss (MSC loss) is proposed to constrain features from multiple perspectives. To further enhance the local detailed texture features, we propose an Edge Enhancement Module (EEM) in the 3D segmentation network to enhance the edge detail features. Our experimental results on the SKI10 dataset and OAI ZIB dataset demonstrate that our method outperforms the SOTA weakly-supervised segmentation methods. Moreover, our approach achieves results that are comparable to the fully-supervised upper bound results. Mingzhu Zhu, Shuyu Guo, Jianhang Jiao, Tian Bai 0002 |
BIBM | 5 |
| 2024 | MATCC: A Novel Approach for Robust Stock Price Prediction Incorporating Market Trends and Cross-time CorrelationsabstractStock price prediction has been a challenging problem due to non-stationary dynamics and complex market dependencies. Existing work has two limitations: 1. Previous studies have underestimated the importance of market trends, relying solely on stock data to learn patterns and capture market regularities implicitly. However, due to random stock fluctuations and trading noise caused by market sentiment, it is difficult to learn underlying market trends, resulting in poor model performance. 2. Prior research has predominantly concentrated on time-aligned feature correlations, with limited exploration of cross-time stock correlations. To address these issues, we propose a novel framework, MATCC (Market Trend and Cross-time Correlation model). It explicitly extracts market trends as guiding information, decomposes stock data into trend and fluctuation components, and employs a carefully designed structure for mining cross-time correlation. Extensive experiments demonstrate that MATCC significantly outperforms previous works in both ranking and portfolio-based metrics. Additionally, we illustrate the influence of trends and correlations on stock prediction through visualization. We publish our code at https://github.com/caozhiy/MATCC. Jiayu Xu 0005, Chengqi Dong, Tian Bai 0002 |
CIKM | 5 |
| 2024 | Caption-Aware Multimodal Relation Extraction with Mutual Information MaximizationabstractMultimodal Relation Extraction (MRE) has achieved great improvements. However, modern MRE models are easily affected by irrelevant objects during multimodal alignment which are called error sensitivity issues. The main reason is that visual features are not fully aligned with textual features and the reasoning process may suppress redundant and noisy information at the risk of losing critical information. In light of this, we propose a Caption-Aware Multimodal Relation Extraction Network with Mutual Information Maximization (CAMIM). Specifically, we first generate detailed image captions through the Large Language Model (LLM). Then, the Caption-Aware Module (CAM) hierarchically aligns the fine-grained visual entities and textual entities for reasoning. In addition, for preserving crucial information within different modalities, we leverage a Mutual Information Maximization method to regulate the multimodal reasoning module. Experiments show that our model outperforms the state-of-the-art MRE models on the benchmark dataset MNRE. Further ablation studies prove the pluggable and effective performance of our Caption-Aware Module and Mutual Information Maximization method. Our code is available at https://github.com/zefanZhang-cn/CAMIM. Zefan Zhang, Tian Bai 0002 |
ACM Multimedia | 4 |
| 2024 | Integrating grid features and geometric coordinates for enhanced image captioning
Fengzhi Zhao, Zhezhou Yu, Tao Wang 0180, Tian Bai 0002 |
Appl. Intell. | 5 |
| 2023 | A Novel Drug-Drug Interaction Prediction Model Based on Line Subgraph Generation StrategyabstractDrug-Drug Interaction (DDI) prediction task is helpful for better-understanding drugs. In this paper, we propose a novel drug-drug interaction prediction model based on line subgraph generation strategy, named DDI-LSG model. Our DDI-LSG model consists of three main parts which include drug relation graph construction, line subgraph generation strategy, and graph-level classification. To consider more relationships among drugs, we propose a node feature-enhancing method to encode drug features in drug relation graph construction process. To consider drugs and DDI as equivalent factors of our DDI-LSG model, we introduce line graph transformation to integrate DDI with drug feature enhancing vector. Combining Jaccard similarity with cosine similarity, we propose a line subgraph generation strategy to evaluate node relation and extract key structures around target DDI in the line graph. Then, we reformulate the DDI prediction task into a graph-level classification task for the line subgraph of the target DDI. Therefore, in the final part of our DDI-LSG model, we use a graph-level classifier to classify the line subgraphs. Our DDI-LSG model outperforms better experiment results than baselines. Ablation results have validated the node feature enhancing method and line subgraph generation strategy. Tian Bai 0002, Chu Li 0002, Xinyue Peng, Haotian Guan, Zefan Zhang, Guishen Wang |
BIBM | 1 |
| 2023 | A Morphology Focused Cell Detection Model for Histopathology ImagesabstractThe accurate automatic recognition of cell locations is of great significance for downstream tasks in pathology. Due to the various size and distribution of different cell types, previous cell detection methods applied fixed circles as labels to localize cell position by default, which makes less precise and more difficult to adapt various cell morphology. In this paper, we integrate adaptive areas of interests based on morphology to address the limitation in detecting cells with irregular shapes in pathological images. Moreover, a EGSI module is proposed to extract global distribution of cells in slices, which enables model to acquire more rich semantic information. We exhibit favorable F1score of our method on four public datasets. Experimental results demonstrate that the proposed method could detect cells more accurately in those images with different cell morphologies and dense cell distribution. Zhe Wang 0007, Fangyue Wei, Shuyu Guo, Xiaoting Che, Tian Bai 0002 |
BIBM | 5 |
| 2023 | Cross-domain endoscopic image translation and landmark detection based on consistency regularization cycle generative adversarial network
Lan Huang 0002, Yuzhao Wang, Yingfang Zhang, Shuyu Guo, Ke Tao, Tian Bai 0002 |
Expert Syst. Appl. | 6 |
| 2023 | A Hybrid VAE Based Network Embedding Method for Biomedical Relation Mining
Tian Bai 0002, Lan Huang 0002 |
Neural Process. Lett. | 1 |
| 2022 | Multi-site MRI classification using Weighted federated learning based on Mixture of Experts domain adaptationabstractDeep learning often requires large amounts of data from different institutions. Federated learning, as a distributed training framework, enables multiple participants to collaboratively train models without collecting data together and hence protecting data privacy, but the datasets from different institutions usually bring the problem of domain shift, which affects the performance of the model. When addressing domain shift, previous works often use a single global model to share parameters. Therefore, we propose a novel method to train multiple public models with different structures under the federated framework to improve the reliability and robustness of the public models. And each participant keeps its own domain-tuned private model, the private model does not share parameters with other participants. We use Mixture of Experts (MoE) domain adaptation to dynamically combine different public models and private model, which utilizes the similarity between different datasets to update the parameters of the public models. We apply the proposed method to the multi-site Magnetic resonance imaging (MRI) end-to-end classification, and the experiments demonstrate its effectiveness. Tian Bai 0002, Yingfang Zhang, Yuzhao Wang, Yanguo Qin, Fa Zhang 0001 |
BIBM | 1 |
| 2022 | Method for Preoperative Prediction of Microvascular Invasion of Hepatocellular CarcinomaabstractMicrovascular invasion (MVI) in hepatocellular carcinoma (HCC) is of great guiding significance for the formulating treatment strategies and accessing the prognosis before the surgery. However, in traditional medicine, the gold standard for the diagnosis of MVI is obtained by examining pathological images which can only be obtained by sampling and sectioning tumors after surgery. At this time, MVI results have lost the timeliness of guiding tumor resection surgery. In order to solve this problem, existing studies began to use deep learning-based methods for preoperative prediction of MVI using non-invasive imaging. Most of these methods adopt the fusion methods of multi-sequence images to predict MVI, but fail to make full use of the characteristics of multiply sequences as prior knowledge to combine into the model, resulting in no further improvement of prediction performance. So we propose a multi-sequence image difference and correlation deep learning model. The model can extract the difference and correlation information between sequences from different scales and combine them into the model. To validate proposed model, we collected a data set consists of 120 HCC patients, including 50 MVI-positive patients. Compared with existing studies, our method has greatly improved in all evaluation metrics. Tian Bai 0002, Tongjia Chu, Fa Zhang 0001 |
BIBM | 2 |
| 2022 | Context-aware learning for cancer cell nucleus recognition in pathology imagesabstractMOTIVATION: Nucleus identification supports many quantitative analysis studies that rely on nuclei positions or categories. Contextual information in pathology images refers to information near the to-be-recognized cell, which can be very helpful for nucleus subtyping. Current CNN-based methods do not explicitly encode contextual information within the input images and point annotations. RESULTS: In this article, we propose a novel framework with context to locate and classify nuclei in microscopy image data. Specifically, first we use state-of-the-art network architectures to extract multi-scale feature representations from multi-field-of-view, multi-resolution input images and then conduct feature aggregation on-the-fly with stacked convolutional operations. Then, two auxiliary tasks are added to the model to effectively utilize the contextual information. One for predicting the frequencies of nuclei, and the other for extracting the regional distribution information of the same kind of nuclei. The entire framework is trained in an end-to-end, pixel-to-pixel fashion. We evaluate our method on two histopathological image datasets with different tissue and stain preparations, and experimental results demonstrate that our method outperforms other recent state-of-the-art models in nucleus identification. AVAILABILITY AND IMPLEMENTATION: The source code of our method is freely available at https://github.com/qjxjy123/DonRabbit. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tian Bai 0002, Jiayu Xu 0005, Zhenting Zhang, Shuyu Guo |
Bioinform. | 1 |
| 2022 | Traditional Chinese medicine entity relation extraction based on CNN with segment attention
Tian Bai 0002, Haotian Guan, Lan Huang 0002 |
Neural Comput. Appl. | 1 |
| 2021 | Self-Ensembling Semi-Supervised Model for Bone X-ray Images Landmark DetectionabstractBone image landmark detection plays a vital role in orthopedic diseases diagnosis and analysis. As a promising approach, the data-driven convolutional neural network has been widely applied in landmarks detection. However, the model’s performance relies heavily on numerous high-quality labeled data, which are difficult to access. A practical solution for this issue in bone images landmark detection has not been provided, so we propose a framework of semi-supervised bone images landmark detection encouraged by the recent success of the self-ensembling method. The student and teacher networks share the same structure in our framework, where unlabeled data are utilized by enforcing the prediction consistency of two networks under different perturbations to enhance the network’s performance. The two networks use a mutual learning approach to improve the efficiency of using unlabeled data. In addition, we also propose a transformative perturbation scheme to prevent model overfitting during training and enhance the generalization capability of the model on unseen data. Experimental results on an in-house hip dataset collected from the actual process of clinical diagnosis and a publicly available dataset demonstrate that our framework has excellent performance compared with other semi-supervised methods. Tian Bai 0002, Shenyao Liu, Yuzhao Wang |
BIBM | 1 |
| 2021 | A Novel Pseudo-Labeling Approach for Cell Detection Based on Adaptive Threshold
Tian Bai 0002, Zhenting Zhang |
ISBRA | 1 |
| 2021 | Recognizing art work image from natural type: a deep adaptive depiction fusion method
Lan Huang 0002, Yuzhao Wang, Tian Bai 0002 |
Vis. Comput. | 3 |
| 2020 | Multi-field of View Aggregation and Context Encoding for Single-Stage Nucleus Recognition
Tian Bai 0002, Jiayu Xu 0005, Fuyong Xing |
MICCAI (5) | 1 |
| 2020 | A novel deep learning method for extracting unspecific biomedical relationabstractSummary Biomedical relation extraction is an important research subject in Natural language processing (NLP). Deep learning technology has shown greater value in improving accuracy of relation extraction results recently. Existing methods mostly focus on extracting (1) specific relation from short texts (eg, drug‐drug interaction and protein‐protein interaction) and (2) unspecific relation from full text corpora. However, extracting unspecific relation from short text, which is more and more important in practical use, is rarely studied. In this paper, a new model called MAT‐LSTM is proposed to extract unspecific relation from short text in biomedical literatures. Experiments on two Biocreative benchmark datasets and one BioNLP benchmark datasets were made to measure the validity of the proposed model MAT‐LSTM, and better performance is achieved. The MAT‐LSTM model is also applied practically in extracting unspecific relation contained in the PubMed literatures. The results extracted from PubMed by using the proposed model were verified by experts mostly, indicating the practical value of the MAT‐LSTM model. Tian Bai 0002, Lan Huang 0002, Fuyong Xing |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | A novel MEDLINE topic indexing method using image presentation
Lan Huang 0002, Shuyu Guo, Leiguang Gong, Tian Bai 0002 |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | MGOGP: a gene module-based heuristic algorithm for cancer-related gene prioritizationabstractBACKGROUND: Prioritizing genes according to their associations with a cancer allows researchers to explore genes in more informed ways. By far, Gene-centric or network-centric gene prioritization methods are predominated. Genes and their protein products carry out cellular processes in the context of functional modules. Dysfunctional gene modules have been previously reported to have associations with cancer. However, gene module information has seldom been considered in cancer-related gene prioritization. RESULTS: In this study, we propose a novel method, MGOGP (Module and Gene Ontology-based Gene Prioritization), for cancer-related gene prioritization. Different from other methods, MGOGP ranks genes considering information of both individual genes and their affiliated modules, and utilize Gene Ontology (GO) based fuzzy measure value as well as known cancer-related genes as heuristics. The performance of the proposed method is comprehensively validated by using both breast cancer and prostate cancer datasets, and by comparison with other methods. Results show that MGOGP outperforms other methods, and successfully prioritizes more genes with literature confirmed evidence. CONCLUSIONS: This work will aid researchers in the understanding of the genetic architecture of complex diseases, and improve the accuracy of diagnosis and the effectiveness of therapy. Lingtao Su, Guixia Liu, Tian Bai 0002, Qingshan Ma |
BMC Bioinform. | 3 |
| 2017 | An improved fruit fly optimization algorithm for solving traveling salesman problemabstractThe traveling salesman problem (TSP), a typical non-deterministic polynomial (NP) hard problem, has been used in many engineering applications. As a new swarm-intelligence optimization algorithm, the fruit fly optimization algorithm (FOA) is used to solve TSP, since it has the advantages of being easy to understand and having a simple implementation. However, it has problems, including a slow convergence rate for the algorithm, easily falling into the local optimum, and an insufficient optimi-zation precision. To address TSP effectively, three improvements are proposed in this paper to improve FOA. First, the vision search process is reinforced in the foraging behavior of fruit flies to improve the convergence rate of FOA. Second, an elimination mechanism is added to FOA to increase the diversity. Third, a reverse operator and a multiplication operator are proposed. They are performed on the solution sequence in the fruit fly’s smell search and vision search processes, respectively. In the experiment, 10 benchmarks selected from TSPLIB are tested. The results show that the improved FOA outperforms other alternatives in terms of the convergence rate and precision. Lan Huang 0002, Gui-chao Wang, Tian Bai 0002, Zhe Wang 0007 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2016 | A method for exploring implicit concept relatedness in biomedical knowledge networkabstractBACKGROUND: Biomedical information and knowledge, structural and non-structural, stored in different repositories can be semantically connected to form a hybrid knowledge network. How to compute relatedness between concepts and discover valuable but implicit information or knowledge from it effectively and efficiently is of paramount importance for precision medicine, and a major challenge facing the biomedical research community. RESULTS: In this study, a hybrid biomedical knowledge network is constructed by linking concepts across multiple biomedical ontologies as well as non-structural biomedical knowledge sources. To discover implicit relatedness between concepts in ontologies for which potentially valuable relationships (implicit knowledge) may exist, we developed a Multi-Ontology Relatedness Model (MORM) within the knowledge network, for which a relatedness network (RN) is defined and computed across multiple ontologies using a formal inference mechanism of set-theoretic operations. Semantic constraints are designed and implemented to prune the search space of the relatedness network. CONCLUSIONS: Experiments to test examples of several biomedical applications have been carried out, and the evaluation of the results showed an encouraging potential of the proposed approach to biomedical knowledge discovery. Tian Bai 0002, Leiguang Gong, Yan Wang 0028, Casimir A. Kulikowski, Lan Huang 0002 |
BMC Bioinform. | 1 |
| 2015 | Implicit knowledge discovery in biomedical ontologies: Computing interesting relatednessesabstractOntologies, seen as effective representations for sharing and reusing knowledge, have become increasingly important in biomedicine, usually focusing on taxonomic knowledge specific to a subject. Efforts have been made to uncover implicit knowledge within large biomedical ontologies by exploring semantic similarity and relatedness between concepts. However, much less attention has been paid to another potentially helpful approach: discovering implicit knowledge across multiple ontologies of different types, such as disease ontologies, symptom ontologies, and gene ontologies. In this paper, we propose a unified approach to the problem of ontology based implicit knowledge discovery - a Multi-Ontology Relatedness Model (MORM), which includes the formation of multiple related ontologies, a relatedness network and a formal inference mechanism based on set-theoretic operations. Experiments for biomedical applications have been carried out, and preliminary results show the potential value of the proposed approach for biomedical knowledge discovery. Tian Bai 0002, Leiguang Gong, Casimir A. Kulikowski, Lan Huang 0002 |
BIBM | 1 |
| 2013 | An improved k-prototypes clustering algorithm for mixed numeric and categorical data
Jinchao Ji, Tian Bai 0002, Chunguang Zhou, Zhe Wang 0007 |
Neurocomputing | 2 |