Yanguo Sun

dblp:349/8837 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0003-1426-3402ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Light-UNet: A simple 3D brain tumor segmentation network
Zhenping Lan, Yanguo Sun, Yuheng Sun, Yuepeng Guo, Yuru Wang
Pattern Recognit.3
2025 Research on occlusion pedestrian re-identification based on ViT model
Yuepeng Guo, Zhenping Lan, Yanguo Sun, Yuheng Sun, Yuru Wang
J. Supercomput.3
2025 Visible-infrared pedestrian re-identification based on local feature enhancement
Yuepeng Guo, Zhenping Lan, Yanguo Sun, Yuheng Sun, Yuru Wang, Yuwei Meng
J. Supercomput.3
2025 DPF-Unet: a CNN-swin transformer fusion network for 3D brain tumor segmentation in MRI images
Zhenping Lan, Yanguo Sun, Yuheng Sun, Yuepeng Guo, Yuru Wang, Aixia Yuan
J. Supercomput.3
2025 Ldstd: low-altitude drone aerial small target detector
Yuheng Sun, Zhenping Lan, Yanguo Sun, Yuepeng Guo, Yuru Wang
J. Supercomput.3
2025 Htfd-yolo: Small target detection in drone aerial photography based on YOLOv8s
Yuheng Sun, Zhenping Lan, Yanguo Sun, Yuepeng Guo, Yuru Wang, Yuwei Meng
J. Supercomput.3
2025 Few-Shot Learning Based on Multimodal Information Processing
abstract
Few-shot learning aims to develop models with strong generalization capabilities using a small number of training samples. However, most learning methods rely solely on the visual features of a few samples to represent entire categories, leading to poor category representativeness. In contrast, humans can utilize multimodal information to learn category features, thereby making them more representative. Hence, this article emulates the human multimodal learning mechanism by integrating visual features with textual information, thereby facilitating the model's acquisition of more representative and robust category features. Specifically, this article introduces a novel multimodal fusion mechanism-the visual-semantic fusion selection mechanism (VSFSM)-which comprises a fusion selection module (FS-Module) and a category enhancement module (CE-Module). These two modules collaboratively enhance the model's classification performance. The FS-Module aligns and fuses semantic information with visual features across both channel and spatial dimensions, performing feature selection and reconstruction. This process not only generates representative category features but also mitigates the impact of noise. The CE-Module guides the model to emphasize category-specific features in the query images, ultimately yielding representative visual-semantic category features while reducing the interference of noise in the query images. Additionally, to better facilitate few-shot learning, this article introduces a novel objective loss function for optimized training. Extensive comparative and ablation experiments conducted on multiple datasets further validate the effectiveness of the proposed method.
Zhenping Lan, Yanguo Sun, Jiansong Li, Xincheng Yang
IEEE Trans. Neural Networks Learn. Syst.3