Cunhan Guo

dblp:358/7368 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-4546-3928ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Coherence-aware and snap-triggered: A novel mechanism for audio-visual cooperative tasks
Cunhan Guo, Heyan Huang, Ruiqi Hu, Danjie Han
Expert Syst. Appl.1
2026 Transferring cross-Dimensional knowledge via proxy task for medical image segmentation
Cunhan Guo, Heyan Huang, Yang-Hao Zhou, Danjie Han, Changsen Yuan
Expert Syst. Appl.1
2026 Leveraging mamba for reference audio-visual segmentation with vote and cache mechanism
Cunhan Guo, Heyan Huang, Yang-Hao Zhou, Changsen Yuan, Danjie Han
Expert Syst. Appl.1
2026 Selective and contrastive mechanism for distantly supervised relation extraction
Danjie Han, Heyan Huang, Shumin Shi, Cunhan Guo, Yanghao Zhou, Changsen Yuan
Neurocomputing4
2026 Towards unified scene understanding in audio-visual semantic segmentation
Danjie Han, Changsen Yuan, Xuan Zhao 0026, Cunhan Guo, Yanghao Zhou
Knowl. Based Syst.5
2026 Seeing With Words: Interpretable Language-Guided Drone Geo-Localization via LLM-Enriched Semantic Attribute Alignment
abstract
Natural language-guided drone geo-localization (DGL) provides an intuitive and scalable mode of human-drone interaction for tasks such as search, rescue, and surveillance. Recent Vision-Language Models (VLMs) can learn semantic correspondences between text and images during fine-tuning. However, their performance in DGL tasks remains constrained, as complex instructions and cluttered scenes often cause semantic dilution and granularity mismatch, leading to weak cross-modal alignment. Consequently, the models struggle with ambiguous targets and suffer from reduced localization accuracy. To address these challenges, we propose SAA-DGL, a framework for interpretable language-guided Drone Geo-Localization that enriches Semantic Attribute Alignment (SAA) with large language models (LLMs). It introduces two parameter-free cross-modal fusion modules: (1) the LLM-driven Cross-modal Semantic Attribute Enrichment (LCSAE) module, which extracts fine-grained attributes (e.g., color, shape, position) from text and embeds them into visual features as explicit semantic anchors, producing semantically enriched cross-modal representations; and (2) the Bidirectional Feature Alignment (BFA) module, which builds fusion relationships between visual and textual features via similarity-driven mechanisms, enabling effective integration of enriched visual and textual information. This design improves cross-modal consistency and interpretability while preserving pretrained alignment priors and enhancing training stability. Experiments on the GeoText-1652 benchmark show that SAA-DGL achieves state-of-the-art performance and strong robustness under complex visual and linguistic disturbances, validating its effectiveness for challenging geo-localization scenarios. We will release the code.
Changsen Yuan, Yang-Hao Zhou, Cunhan Guo, Danjie Han, Ge Shi 0002, Wenwu Wang 0001
IEEE Trans. Multim.3
2025 CARE: Contextual Augmentation with Retrieval Enhancement for Relation Extraction in Large Language Models
Danjie Han, Heyan Huang, Shumin Shi, Cunhan Guo, Yanghao Zhou, Changsen Yuan
NLPCC (1)4
2025 Enhancing camouflaged object detection through contrastive learning and data augmentation techniques
Cunhan Guo, Heyan Huang
Eng. Appl. Artif. Intell.1
2025 CoFiNet: Unveiling camouflaged objects with multi-scale finesse
Cunhan Guo, Heyan Huang
Neurocomputing1
2025 Distantly Supervised relation extraction with multi-level contextual information integration
Danjie Han, Heyan Huang, Shumin Shi, Changsen Yuan, Cunhan Guo
Neurocomputing5
2025 ALOHA: Adapting Local Spatio-Temporal Context to Enhance the Audio-Visual Semantic Segmentation
abstract
Audio-Visual Semantic Segmentation (AVSS) plays a crucial role in pixel-level multi-modal perception for real-world applications such as robotic navigation and autonomous driving. Existing methods typically rely on global spatio-temporal modules to fuse audio and visual representations, which aids in generating pixel-level semantic masks. However, these approaches often overlook the importance of local spatio-temporal context in understanding semantics, leading to suboptimal performance. This limitation makes it difficult for models to accurately distinguish sound-emitting objects from irrelevant background noise, resulting in erroneous segmentation across the spatio-temporal dimension. To address this issue, we propose the ALOHA framework, which A dapts LO cal spatio-temporal context to en HA nce AVSS. The framework introduces two key components designed to leverage and enhance local spatio-temporal context information: the LOHA adapter and the Selective Context Enhancement (SCE) module. Specifically, the LOHA adapter adaptively captures essential modality information across spatio-temporal dimensions, while implicitly learning fine-grained local context through the local attention mechanism. Furthermore, the SCE module selectively enhances the local context related to the semantics, thereby facilitating the distinction between the sounding object and irrelevant background and improving segmentation accuracy. Moreover, to better adapt to embodied AI systems, our framework utilizes a parameter-shared encoder and applies the adapters in a staged manner. This design significantly reduces the number of trainable parameters, making it more parameter-efficient. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance on the AVSBench-Semantic benchmark dataset and shows competitive results on the AVSBench-Object benchmark, while exhibiting broad adaptability across different visual backbone networks.
Yang-Hao Zhou, Heyan Huang, Cunhan Guo, Rongcheng Tu, Zeyu Xiao 0002, Bo Wang 0134, Xianling Mao
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Leveraging Vote and Cooperate Mechanism for Brain Tumor Segmentation
abstract
The task of brain tumor segmentation necessitates the processing of long Magnetic Resonance Imaging (MRI) sequences across multiple imaging modalities. Traditional convolutional neural networks often exhibit suboptimal ability of long-term memory, while Transformers demand substantial computational resources. To mitigate computational requirements and optimally utilize the information from various imaging modalities, we propose Mamba-based Vote and Cooperate Segmentation (VCSeg), for brain tumor segmentation. Features from the imaging sequences of each modality are extracted from multiple spatial orientations, and assigned different weights based on both modality and orientation considerations, enabling the model to adaptively learn modality-specific feature distribution for adult and child patient situation. The model employs a multi-modal cooperate module to further enhance the feature encoding capabilities. Additionally, deep supervision is applied to achieve progressive mask generation, thereby improving decoding quality. Experiments conducted on the BraTS2019-MEN and BraTS2023-PED datasets have demonstrated that VCSeg achieves state-of-the-art results, demonstrating the effectiveness of our advanced method.
Cunhan Guo, Heyan Huang, Changsen Yuan, Yanghao Zhou
BIBM1
2024 Enhance audio-visual segmentation with hierarchical encoder and audio guidance
Cunhan Guo, Heyan Huang, Yanghao Zhou
Neurocomputing1