Guanyu Zhou

dblp:81/11348 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-6196-9451ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 35% Language models and text generation · 20% Trustworthy machine learning · 17%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
hallucination mitigation
1.922026
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination · AAAI 2026
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality · ICLR 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
1.922026
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination · AAAI 2026
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality · ICLR 2025
Machine learning › Trustworthy machine learning
robustness
1.922026
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination · AAAI 2026
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality · ICLR 2025
Computer vision › Image recognition and object detection
medical image analysis
1.012026
DGSAN: Dual-Graph Spatiotemporal Attention Network for Pulmonary Nodule Malignancy Prediction · AAAI 2026
Computer vision › Vision and language
multimodal grounding
1.012026
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination · AAAI 2026
Medical and health informatics › oncology
cancer diagnosis
1.012026
DGSAN: Dual-Graph Spatiotemporal Attention Network for Pulmonary Nodule Malignancy Prediction · AAAI 2026
Medical and health informatics
clinical decision support
1.012026
DGSAN: Dual-Graph Spatiotemporal Attention Network for Pulmonary Nodule Malignancy Prediction · AAAI 2026
Machine learning › Deep learning architectures and training › attention mechanism
attention alignment
0.912025
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality · ICLR 2025
Computer vision › Vision and language
multimodal hallucination
0.912025
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality · ICLR 2025
Machine learning › Graph learning
graph neural network
0.312026
DGSAN: Dual-Graph Spatiotemporal Attention Network for Pulmonary Nodule Malignancy Prediction · AAAI 2026
Natural language and speech › Language models and text generation
preference optimization
0.312026
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination · AAAI 2026
Machine learning › Graph learning › graph neural network
spatio-temporal graph
0.312026
DGSAN: Dual-Graph Spatiotemporal Attention Network for Pulmonary Nodule Malignancy Prediction · AAAI 2026
Machine learning › Probabilistic and Bayesian machine learning
boltzmann machine
0.212013
Learning and Selecting Features Jointly with Point-wise Gated Boltzmann Machines · ICML (2) 2013
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection
0.212013
Learning and Selecting Features Jointly with Point-wise Gated Boltzmann Machines · ICML (2) 2013
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.212013
Learning and Selecting Features Jointly with Point-wise Gated Boltzmann Machines · ICML (2) 2013

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 2.0graph attention network · 2.0cross-modal attention · 2.0preference sample construction · 1.0preference optimization · 1.0DPO · 1.0structural causal modeling · 0.9counterfactual reasoning · 0.9causal inference · 0.9backdoor adjustment · 0.9
YearPublicationVenuePosition
2026 OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
abstract
Recently, Omni-modal large language models (OLLMs) have sparked a new wave of research, achieving impressive results in tasks such as audio-video understanding and real-time environment perception. However, hallucination issues still persist. Similar to the bimodal setting, the priors from the text modality tend to dominate, leading OLLMs to rely more heavily on textual cues while neglecting visual and audio information. In addition, fully multimodal scenarios introduce new challenges. Most existing models align visual or auditory modalities with text independently during training, while ignoring the intrinsic correlations between video and its corresponding audio. This oversight results in hallucinations when reasoning requires interpreting hidden audio cues embedded in video content. To address these challenges, we propose OmniDPO, a preference-alignment framework designed to mitigate hallucinations in OLLMs. Specifically, OmniDPO incorporates two strategies: (1) constructing text-preference sample pairs to enhance the model’s understanding of audio-video interactions; and (2) constructing multimodal-preference sample pairs to strengthen the model’s attention to visual and auditory information. By tackling both challenges, OmniDPO effectively improves multimodal grounding and reduces hallucination. Experiments conducted on two OLLMs demonstrate that OmniDPO not only effectively mitigates multimodal hallucinations but also significantly enhances the models' reasoning capabilities across modalities.
Junzhe Chen 0001, Tianshu Zhang 0002, Shiyu Huang 0001, Yuwei Niu, Rongzhou Zhang, Guanyu Zhou, Lijie Wen 0001
AAAI7
2026 DGSAN: Dual-Graph Spatiotemporal Attention Network for Pulmonary Nodule Malignancy Prediction
abstract
Lung cancer continues to be the leading cause of cancer-related deaths globally. Early detection and diagnosis of pulmonary nodules are essential for improving patient survival rates. Although previous research has integrated multimodal and multi-temporal information, outperforming single modality and single time point, the fusion methods are limited to inefficient vector concatenation and simple mutual attention, highlighting the need for more effective multimodal information fusion. To address these challenges, we introduce a Dual-Graph Spatiotemporal Attention Network, which leverages temporal variations and multimodal data to enhance the accuracy of predictions. Our methodology involves developing a Global-Local Feature Encoder to better capture the local, global, and fused characteristics of pulmonary nodules. Additionally, a Dual-Graph Construction method organizes multimodal features into inter-modal and intra-modal graphs. Furthermore, a Hierarchical Cross-Modal Graph Fusion Module is introduced to refine feature integration. We also compiled a novel multimodal dataset named the NLST-cmst dataset as a comprehensive source of support for related research. Our extensive experiments, conducted on both the NLST-cmst and curated CSTL-derived datasets, demonstrate that our DGSAN significantly outperforms state-of-the-art methods in classifying pulmonary nodules with exceptional computational efficiency.
Zhaojie Fang, Guanyu Zhou, Yin Shen, Huoling Luo, Ahmed El-Azab, Ruiquan Ge, Changmiao Wang
AAAI3
2026 AsynFormer: Transformer capturing asynchronous cross-variate dependencies for efficient multivariate time series forecasting
Yanglei Gan, Run Lin, Guanyu Zhou, Yao Liu 0019, Qiao Liu 0003
Knowl. Based Syst.5
2026 No modality left behind: Adapting to missing modalities via knowledge distillation for brain tumor segmentation
Shenghao Zhu, Yifei Chen 0019, Guanyu Zhou, Yuanhan Wang, Fei-wei Qin, Changmiao Wang, Qiyuan Tian
Medical Image Anal.5
2026 MUIT-TTA: Annotation-free intracranial hemorrhage segmentation via pseudo-anomaly synthesis and test-time adaptation
Jinying Zong, Yifei Chen 0019, Mingxuan Liu 0001, Changwei Wu, Beining Wu, Guanyu Zhou, Fei-wei Qin
Pattern Recognit.7
2025 WARPNet: Scale-Wise Autoregressive Cross-Modal Synthesis for Accurate and Detail-Preserving MRI-to-PET Generation
abstract
Due to the inherent limitation of MRI in directly capturing early metabolic abnormalities associated with neurological disorders, and considering the high cost and radiation risks associated with PET scans, cross-modal MRI-to-PET image synthesis has emerged as a critical pathway for early and precise diagnosis. However, current methods generally suffer from structural distortion, blurred details, and computational inefficiencies, significantly restricting their clinical applicability. To address these limitations, this paper proposes an innovative multi-scale autoregressive-driven framework for MRI-to-PET cross-modal image generation. By explicitly modeling scalewise transformations between MRI and PET via a multi-scale autoregressive mechanism, and incorporating wavelet transform with a linear multi-step connection strategy, our framework effectively enhances structural accuracy and texture detail expression, especially in lesion regions. Experimental results on the ADNI Alzheimer's Disease dataset and a private epilepsy dataset demonstrate that the proposed method consistently outperforms state-of-the-art approaches, generating high-quality PET images efficiently and robustly. Furthermore, it substantially reduces diagnostic costs and radiation exposure, showcasing promising prospects for clinical adoption. Our source code is available at https://github.com/Guanyu-Zhou/WARPNet.
Guanyu Zhou, Yifei Chen 0019, Gaoxiang Ying, Mingxuan Liu 0001, Xuguang Bai, Jialan Zheng, Bixiao Cui, Qiyuan Tian, Jie Lu 0010
BIBM1
2025 Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
abstract
Multimodal Large Language Models (MLLMs) have emerged as a central focus in both industry and academia, but often suffer from biases introduced by visual and language priors, which can lead to multimodal hallucination. These biases arise from the visual encoder and the Large Language Model (LLM) backbone, affecting the attention mechanism responsible for aligning multimodal inputs. Existing decoding-based mitigation methods focus on statistical correlations and overlook the causal relationships between attention mechanisms and model output, limiting their effectiveness in addressing these biases. To tackle this issue, we propose a causal inference framework termed CausalMM that applies structural causal modeling to MLLMs, treating modality priors as a confounder between attention mechanisms and output. Specifically, by employing backdoor adjustment and counterfactual reasoning at both the visual and language attention levels, our method mitigates the negative effects of modality priors and enhances the alignment of MLLM's inputs and outputs, with a maximum score improvement of 65.3% on 6 VLind-Bench indicators and 164 points on MME Benchmark compared to conventional methods. Extensive experiments validate the effectiveness of our approach while being a plug-and-play solution. Our code is available at: https://github.com/The-Martyr/CausalMM.
Guanyu Zhou, Xin Zou 0001, Kun Wang 0056, Aiwei Liu, Xuming Hu
ICLR1
2025 LPUWF-LDM: Enhanced latent diffusion model for precise late-phase UWF-FA generation on limited dataset
Zhaojie Fang, Guanyu Zhou, Ke Zhuang, Yifei Chen 0019, Ruiquan Ge, Changmiao Wang, Gangyong Jia, Qing Wu 0008, Juan Ye, Maimaiti Nuliqiman, Peifang Xu, Ahmed El-Azab
Expert Syst. Appl.3
2013 Learning and Selecting Features Jointly with Point-wise Gated Boltzmann Machines
abstract
Unsupervised feature learning has emerged as a promising tool in learning representations from unlabeled data. However, it is still challenging to learn useful high-level features when the data contains a significant amount of irrelevant patterns. Although feature selection can be used for such complex data, it may fail when we have to build a learning system from scratch (i.e., starting from the lack of useful raw features). To address this problem, we propose a point-wise gated Boltzmann machine, a unified generative model that combines feature learning and feature selection. Our model performs not only feature selection on learned high-level features (i.e., hidden units), but also dynamic feature selection on raw features (i.e., visible units) through a gating mechanism. For each example, the model can adaptively focus on a variable subset of visible nodes corresponding to the task-relevant patterns, while ignoring the visible units corresponding to the task-irrelevant patterns. In experiments, our method achieves improved performance over state-of-the-art in several visual recognition benchmarks.
Kihyuk Sohn, Guanyu Zhou, Chansoo Lee, Honglak Lee
ICML (2)2