Bo Liu 0113

dblp:58/2670-113 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-2165-245XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 PromptEmo: Learning Emotion with Bilateral Textual Prompts in Multi-Domain Open-set Scenarios
abstract
Facial Expression Recognition (FER) is crucial to human-computer interaction. Existing cross-domain FER (CD-FER) methods mainly focus on single-source closed-set scenarios, transferring knowledge from a single source domain to a target domain with identical class sets. However, CD-FER faces two real-world challenges: 1) the need to leverage information from multiple sources, leading to multi-domain shift, and 2) the necessity to recognize unseen target classes, resulting in class shift. These issues give rise to a novel and challenging task, which we define as Multi-domain Open-set FER (MO-FER). In this paper, we propose PromptEmo, a novel CLIP-based framework that leverages bilateral textual prompts to address both shifts in the MO-FER task. Leveraging the generalizability of LLM, PromptEmo constructs trainable positive prompts with LLM-generated emotion descriptions for seen classes, as well as template-derived negative prompts to enhance the reasoning for unseen classes. Then, we introduce a modal-task optimization paradigm organized from two perspectives: textual semantics and visual domains, yielding Intra-modal Space-specific Optimization (ISO) and Cross-modal Emotion-aware Interaction (CEI) strategies. ISO refines the CLIP-based textual space to ensure semantic separation between bilateral prompts and improves the latent visual space by promoting inter-domain alignment. Founded on ISO, CEI facilitates effective vision-language interactions, resulting in four joint loss terms that improve emotion recognition by shaping a domain-invariant, discriminative feature space. PromptEmo surpasses the current SOTA method by 7.7% AUC on unseen classes across four FER datasets, serving as a strong baseline for the MO-FER task.
Xinyi Zeng, Yuxiang Yang 0009, Pinxian Zeng, Wenxia Yin, Bo Liu 0113, Xi Wu 0004, Yan Wang 0015
AAAI5
2026 From Knowledge to Causality: Self-supervised Representation Learning for Granger Causal Discovery in Groups of Time Series
Bo Liu 0113, Di Dai, Hongyan Li 0002, Shenda Hong
DASFAA (4)1
2026 LLM-GC: Advancing Granger Causal Discovery from Time Series with Multimodel Language Modeling
abstract
Recent advances in neural Granger causal methods have shown promise in modeling temporal nonlinear dependencies. However, existing approaches remain confined to raw time-series data, inherently lacking contextual semantics and tending to overfit, which undermines their real-world applicability. To address these challenges, we propose LLM-GC, a novel LLM-empowered multimodal Granger causality discovery framework that enriches unimodal temporal dynamics with semantic priors and world knowledge distilled from large language models (LLMs). LLM-GC leverages dual-modality encoding to capture and align temporal and contextual dynamics by Cross-Modal Dual Retrieval while avoiding causal entanglement across modalities. To extract multimodal causal features, we introduce a causality-aware self-attention mechanism by simply inverting the conventional self-attention structure, enabling a shared causality augmenter to effectively highlight consistent causal patterns across modalities. LLM-GC is the first to bridge LLMs and Granger causality, and experiments on synthetic and real-world benchmark datasets demonstrate that LLM-GC outperforms existing state-of-the-art methods in Granger causal discovery.
Bo Liu 0113, Hongyan Li 0002, Shenda Hong
WSDM1
2026 CD-Former: A Cross-Modal Dual-Interaction Transformer With Whole-Slide Image Pyramids and Genomics for Survival Prediction
abstract
Survival prediction is crucial for cancer patients as it provides essential early prognostic information for treatment planning and decision making. Despite impressive performance , current multi-modal survival prediction methods that integrate pathology and genomic data face two main challenges: (1) Whole-slide images (WSIs) generally exhibit hierarchical structures, but the interactions of phenotypes at different resolutions remain unexplored. More importantly, the potential semantic discrepancy arising from diverse resolutions is often ignored. (2) The absence of effective interactions between the inherent hierarchical structures of WSIs and genomic data. To address these challenges, in this paper, we propose Cross-modal Dual-interaction Transformer (CD-Former), a robust hierarchical framework for multi-modal survival prediction. Our CD-Former involves two key components: (1) an Multimodal Cross-Scale Calibration (MCSC) module for effectively capturing correlations across multiple resolutions and calibrating fine-grained features, thereby bridging the semantic discrepancy caused by different WSI resolutions; and (2) a hierarchical interaction module termed Multi-modal Dual-interaction (M2Di) for fully exploring multi-resolution cross-modal correlations and interactions, which comprises a Patch-level Cross-Attention Block (PCAB) and a Region-level Cross-Attention Block (RCAB) to investigate cross-modal associations between patch- or region-level features of WSI and genomic data. Additionally, we employ a scale-oriented WSI enhancer to capture the interactions among various components of WSIs. The experimental results demonstrate the effectiveness of our proposed framework, which achieves state-of-the-art performance compared to previous studies.
Lifan Long, Xingchen Peng, Bo Liu 0113, Xi Wu 0004, Daoqiang Zhang, Yan Wang 0015
IEEE Trans. Circuits Syst. Video Technol.4
2026 MGTP: Multi-Granularity Textual Prompts for Low-Dose Brain PET Image Denoising via Adversarial Diffusion Model
abstract
Positron emission tomography (PET) is an advanced nuclear imaging technique and has been widely applied in clinic. However, radiation risks associated with standard-dose PET imaging raise health concerns, whereas the quality of low-dose PET images fails to meet clinical requirements. To reduce the tracer dose while maintaining image quality, it is of great interest to estimate high-quality PET images from low-dose images. However, existing low-dose PET image denoising methods primarily focus on image data, overlooking crucial information in non-image textual data such as patients' clinical tabular and textual descriptions of general image quality. This neglect can lead to subpar denoising quality with inaccurate contexts and poor details. To address these problems, in this paper, we propose Multi-Granularity Textual Prompts, namely MGTP, to denoise low-dose PET images via an adversarial diffusion model. Different from prior methods that rely solely on image conditioning, our MGTP innovatively introduces textual prompts spanning diverse granularities to capture both high-level semantic-related contexts and low-level degradation-related details. To harmonize multi-granularity textual prompts with low-dose PET images, we design a Cross-Modality Selective Conditioning (CMSC) module, which prioritizes semantic- and detail-relevant information while eliminating irrelevant components. The resulting features are fed into diffusion model as conditions, enforcing a more controlled diffusion process. In addition, we develop a Masked Prompt Reconstruction Network (MPR-Net) to enhance the preservation of semantics and details in denoised images, mitigating distortions brought by the random noise in the diffusion process. Experiments on clinical PET data show that our method achieves the state-of-the-art performance.
Xinyi Zeng, Pinxian Zeng, Bo Liu 0113, Xi Wu 0004, Deng Xiong, Jiliu Zhou, Yan Wang 0015, Dinggang Shen
IEEE J. Biomed. Health Informatics4
2025 DiffuGC: Diffusion Model Can Help Discover Granger Causality from Interventional Time Series
abstract
Discovering Granger causality from time series data is fundamental to understanding dynamic systems, yet most existing methods struggle with unknown intervention targets or causal structures in real-world scenarios. In this paper, we propose DiffuGC, a novel diffusion-based framework that unifies observational and interventional causal discovery through a generative denoising process. By introducing diffusive interventions, which apply progressive interventions without any prior knowledge, DiffuGC amplifies causal signals while preserving structural information. Furthermore, we introduce a denoising NoiFormer with adaptive attention to both short- and long-term causal dependencies, which disentangles trend and seasonal components to enable accurate reconstruction of causal structures from interventional data. To the best of our knowledge, we are the first to integrate diffusion models with interventional Granger causal discovery. Extensive experiments on synthetic, quasi-real, and real-world benchmarks demonstrate that DiffuGC consistently outperforms state-of-the-art baselines in both observational and interventional data. Moreover, we introduce an intriguing notion, Causality Acceleration, characterized by the early emergence of informative causal patterns within the diffusion path, which may open up promising directions for future research on efficient and adaptive causal discovery.
Bo Liu 0113, Hongyan Li 0002, Shenda Hong
ICDM1
2025 HiLa: Hierarchical Vision-Language Collaboration for Cancer Survival Prediction
Lu Wen, Yuchen Fei, Bo Liu 0113, Luping Zhou, Dinggang Shen, Yan Wang 0015
MICCAI (5)4
2025 GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning
abstract
Medical visual question answering aims to support clinical decision-making by enabling models to answer natural language questions based on medical images. While recent advances in multi-modal learning have significantly improved performance, current methods still suffer from limited answer reliability and poor interpretability, impairing the ability of clinicians and patients to understand and trust model outputs. To address these limitations, this work first proposes a Region-Aware Multimodal Chain-of-Thought (RMCoT) dataset, in which the process of producing an answer is preceded by a sequence of intermediate reasoning steps that explicitly ground relevant visual regions of the medical image, thereby providing fine-grained explainability. Furthermore, we introduce a novel verifiable reward mechanism for reinforcement learning to guide post-training, improving the alignment between the model's reasoning process and its final answer. Remarkably, our method achieves comparable performance using only one-eighth of the training data, demonstrating the efficiency and effectiveness of the proposal. The dataset is available at https://www.med-vqa.com/GEMeX/.
Bo Liu 0113, Along He, Huazhu Fu, Xiao-Ming Wu 0003
ACM Multimedia1
2025 Uncertainty-Aware Medical Diagnostic Phrase Identification and Grounding
abstract
Medical phrase grounding is crucial for identifying relevant regions in medical images based on phrase queries, facilitating accurate image analysis and diagnosis. However, current methods rely on manual extraction of key phrases from medical reports, reducing efficiency and increasing the workload for clinicians. Additionally, the lack of model confidence estimation limits clinical trust and usability. In this paper, we introduce a novel task-Medical Report Grounding (MRG)-which aims to directly identify diagnostic phrases and their corresponding grounding boxes from medical reports in an end-to-end manner. To address this challenge, we propose uMedGround, a a robust and reliable framework that leverages a multimodal large language model to predict diagnostic phrases by embedding a unique token, < $\mathtt {BOX}$BOX >, into the vocabulary to enhance detection capabilities. A vision encoder-decoder processes the embedded token and input image to generate grounding boxes. Critically, uMedGround incorporates an uncertainty-aware prediction model, significantly improving the robustness and reliability of grounding predictions. Experimental results demonstrate that uMedGround outperforms state-of-the-art medical phrase grounding methods and fine-tuned large visual-language models, validating its effectiveness and reliability. This study represents a pioneering exploration of the MRG task, marking the first-ever endeavor in this domain. Additionally, we demonstrate the applicability of uMedGround in medical visual question answering and class-based localization tasks, where it highlights visual evidence aligned with key diagnostic phrases, supporting clinicians in interpreting various types of textual inputs, including free-text reports, visual question answering queries, and class labels.
Ke Zou, Yang Bai 0011, Bo Liu 0113, Zhihao Chen 0004, Yang Zhou 0017, Xuedong Yuan, Meng Wang 0038, Xiaojing Shen, Xiaochun Cao, Huazhu Fu
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 MCAD: Multi-modal Conditioned Adversarial Diffusion Model for High-Quality PET Image Reconstruction
Xinyi Zeng, Pinxian Zeng, Bo Liu 0113, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015
MICCAI (7)4
2024 ABP: Asymmetric Bilateral Prompting for Text-Guided Medical Image Segmentation
Xinyi Zeng, Pinxian Zeng, Aibing Li, Bo Liu 0113, Chengdi Wang, Yan Wang 0015
MICCAI (9)5
2024 Towards Medical Vision-Language Contrastive Pre-training via Study-Oriented Semantic Exploration
abstract
Contrastive vision-language pre-training has shown great promise in representation transfer learning and cross-modality learning in the medical field. However, without fully exploiting the intrinsic properties and correlations of multimodal medical data within patient studies, current research fails to explore all the potential of available data, leading to suboptimal performance on representation learning. In this paper, we propose a novel pre-training framework for learning better medical vision-language embedding, oriented on patients' study-level data. Based on the order-agnostic property of radiology report, we adopt a two-stage feature extraction method for more representative textual characterization. Then, by leveraging momentum encoders and memory queues, study-level semantics are explored with three contrastive objectives to provide comprehensive supervision from three perspectives, i.e., cross-modal, multi-modal, and uni-modal, such that the potential information neglected by previous research can be fully exploited. The superiority of the proposed framework is demonstrated by the impressive improvements on four typical downstream tasks, including zero-shot/data-efficient image classification, image segmentation, and cross-modal retrieval.
Bo Liu 0113, Yan Wang 0015
ACM Multimedia1
2023 Improving Medical Vision-Language Contrastive Pretraining With Semantics-Aware Triage
abstract
Medical contrastive vision-language pretraining has shown great promise in many downstream tasks, such as data-efficient/zero-shot recognition. Current studies pretrain the network with contrastive loss by treating the paired image-reports as positive samples and the unpaired ones as negative samples. However, unlike natural datasets, many medical images or reports from different cases could have large similarity especially for the normal cases, and treating all the unpaired ones as negative samples could undermine the learned semantic structure and impose an adverse effect on the representations. Therefore, we design a simple yet effective approach for better contrastive learning in medical vision-language field. Specifically, by simplifying the computation of similarity between medical image-report pairs into the calculation of the inter-report similarity, the image-report tuples are divided into positive, negative, and additional neutral groups. With this better categorization of samples, more suitable contrastive loss is constructed. For evaluation, we perform extensive experiments by applying the proposed model-agnostic strategy to two state-of-the-art pretraining frameworks. The consistent improvements on four common downstream tasks, including cross-modal retrieval, zero-shot/data-efficient image classification, and image segmentation, demonstrate the effectiveness of the proposed strategy in medical field.
Bo Liu 0113, Donghuan Lu, Dong Wei 0004, Xian Wu 0001, Yan Wang 0015, Yu Zhang 0185, Yefeng Zheng 0001
IEEE Trans. Medical Imaging1