Nannan Hu

dblp:86/7405 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-4080-7891ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Regularization-based semi-supervised generative adversarial learning for text classification with limited supervision
Nannan Hu, Yuefeng Zhao, Zongpeng Li, Qibin Li, Nianmin Yao, Nai Zhou
Eng. Appl. Artif. Intell.1
2026 HPRNet : A holistic position-aware residual network for image captioning
Yuefeng Zhao, Qianqian Tao, Jinrun Huang, Wenkai Song, Nannan Hu
Pattern Recognit.6
2025 CrossMed-SAM: Cross-Modal Medical Image Segmentation via Frequency-Seale-Semantic Awareness
abstract
Although vision foundation models such as SAM excel at natural image segmentation, their transfer to cross-modal medical segmentation remains difficult due to frequency mismatches, extreme scale variability, and modality-specific se-mantics. Existing adaptations typically require heavy fine-tuning and bespoke modules, which undermine generalization across diverse anatomies and imaging protocols. To address these chal-lenges, we propose CrossMed-SAM, which integrates three syner-gistic modules. The Cross-Frequency Attention Module (CFAM) leverages discrete wavelet transform to handle frequency domain variations across imaging modalities, while the Multi-Scale At-tention Module (MSAM) employs parallel dilated convolutions to capture extreme scale variations from microscopic lesions to large anatomical structures. The Adaptive Semantic Fusion Mod-ule (ASFM) integrates CLIP-based medical knowledge through dynamic gating mechanisms to provide semantic guidance when visual features are ambiguous. Extensive experiments on twelve datasets across five diverse medical imaging modalities demon-strate that CrossMed-SAM significantly outperforms existing state-of-the-art methods, achieving Dice coefficient improvements of 1.7% to 5.3% over the strong baseline MedSAM, with superior generalization capabilities across unseen datasets.
Qifei Wang, Yuefeng Zhao, Nai Zhou, Nannan Hu
BIBM4
2025 Adaptive Mixture of Experts for Cross-Domain Medical Image Segmentation with Vision Foundation Models
Qifei Wang, Yuefeng Zhao, Nai Zhou, Qianqian Tao, Nannan Hu
PRCV (18)6
2025 DSSLNet: dual-stream self-supervised learning network for end-to-end speech recognition
Boyang Lyu, Chunxiao Fan 0001, Yue Ming 0001, Nannan Hu
Appl. Intell.4
2024 F2D-SIFPNet: a frequency 2D Slow-I-Fast-P network for faster compressed video action recognition
Yue Ming 0001, Jiangwan Zhou, Xia Jia, Nannan Hu
Appl. Intell.7
2024 Action recognition in compressed domains: A survey
Yue Ming 0001, Jiangwan Zhou, Nannan Hu, Panzi Zhao, Boyang Lyu, Hui Yu 0001
Neurocomputing3
2024 CDGAN-BERT: Adversarial constraint and diversity discriminator for semi-supervised text classification
Nai Zhou, Nianmin Yao, Nannan Hu, Jian Zhao 0029
Knowl. Based Syst.3
2024 CSS-Net: A Consistent Segment Selection Network for Audio-Visual Event Localization
abstract
Audio-visual event (AVE) localization aims to localize the temporal boundaries of events that contains visual and audio contents, to identify event categories in unconstrained videos. Existing work usually utilizes successive video segments for temporal modeling. However, ambient sounds or irrelevant visual targets in some segments often cause the problem of audio-visual semantics inconsistency, resulting in inaccurate global event modeling. To tackle this issue, we present a consistent segment selection network (CSS-Net) in this paper. First, we propose a novel bidirectional guided co-attention (BGCA) block, containing two distinct attention paths from audio to vision and from vision to audio, to focus on sound-related visual regions and event-related sound segments. Then, we propose a novel context-aware similarity measure (CASM) module to select semantic consistent visual and audio segments. A cross-correlation matrix is constructed using the correlation coefficients between the visual and audio feature pairs in all time steps. By extracting highly correlated segments and discarding low correlated segments, visual and audio features can learn global event semantics in videos. Finally, we propose a novel audio-visual contrastive loss to learn the similar semantics representation for visual and audio global features under the constraints of cosine and L2 similarities. Extensive experiments on public AVE dataset demonstrates the effectiveness of our proposed CSS-Net. The localization accuracies achieve the best performance of 80.5% and 76.8% in both fully- and weakly-supervised settings compared with other state-of-the-art methods.
Yue Ming 0001, Nannan Hu, Hui Yu 0001
IEEE Trans. Multim.3
2023 Frequency Enhancement Network for Efficient Compressed Video Action Recognition
abstract
The existing frequency-based action recognition methods achieve impressive performance in improving efficiency. However, they ignore the low-frequency texture and edge clues, leading to accuracy degradation. To address this problem, we propose a novel frequency enhancement (FE) block for efficient compressed video action recognition, including a temporal-channel two-heads attention (TCTHA) module and a frequency overlapping group convolution (FOGC) module. First, the TCTHA module emphasizes the inter-frame temporal context and the inner-frame informative frequency semantics by attention. Then, the FOGC module groups channels in different frequency bands with overlap, to extract low-frequency texture and edge clues, while maintaining the interaction of groups. We integrate the FE block into 2D-CNNs with frequency I-frame input, termed FENet, focusing on the pivotal low-frequency spatio-temporal semantics for action recognition. Experiments on HMDB-51, UCF-101, Kinetics-400, and Kinetics-700 verify that our FENet achieves comparable accuracy compared with RGB-based methods with high efficiency.
Yue Ming 0001, Xia Jia, Jiangwan Zhou, Nannan Hu
ICIP7
2023 See, move and hear: a local-to-global multi-modal interaction network for video action recognition
Yue Ming 0001, Nannan Hu, Jiangwan Zhou
Appl. Intell.3
2023 MAENet: A novel multi-head association attention enhancement network for completing intra-modal interaction in image captioning
Nannan Hu, Chunxiao Fan 0001, Yue Ming 0001
Neurocomputing1
2023 En-HACN: Enhancing Hybrid Architecture With Fast Attention and Capsule Network for End-to-end Speech Recognition
abstract
Automatic speech recognition (ASR) is a fundamental technology in the field of artificial intelligence. End-to-end (E2E) ASR is favored for its state-of-the-art performance. However, E2E speech recognition still faces speech spatial information loss and text local information loss, which results in the increase of deletion and substitution errors during inference. To overcome this challenge, we propose a novel Enhancing Hybrid Architecture with Fast Attention and Capsule Network (termed En-HACN), which can model the position relationships between different acoustic unit features to improve the discriminability of speech features while providing the text local information during inference. Firstly, a new CNN-Capsule Network (CNN-Caps) module is proposed to capture the spatial information in the spectrogram through the capsule output and dynamic routing mechanism. Then, we design a novel hybrid structure of LocalGRU Augmented Decoder (LA-decoder) that generates text hidden representations to obtain text local information of the target sequences. Finally, we introduce fast attention instead of self-attention in En-HACN, which improves the generalization ability and efficiency of the model in long utterances. Experiments on corpora Aishell-1 and Librispeech demonstrate that our En-HACN has achieved the state-of-the-art compared with existing works. Besides, experiments on the long utterances dataset based on Aishell-1-long show that our model has a high generalization ability and efficiency.
Boyang Lyu, Chunxiao Fan 0001, Yue Ming 0001, Panzi Zhao, Nannan Hu
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 TSFNet: Triple-Steam Image Captioning
abstract
Image captioning is a challenging task that generates a natural language description based on the visual understanding of the given image. Significant region representation is a milestone in image captioning. Despite the great success of existing region-based works, they only focus on the salient objects and encode these objects independently, still plagued by the lack of global contextual information and visual relationships. In fact, the global contextual information and structured visual relationships are exactly the merits of traditional grid features and emerging scene graph features. In this paper, we present a Triple-Steam Feature Fusion Network (TSFNet) to leverage the complementary advantages of the grid, region, and scene graph triple-steam visual representations in image captioning. Concretely, in our TSFNet, a novel Dual-level Attention (DA) mechanism is proposed to simultaneously explore visual intrinsic properties and word-related attributes uniformly of different features. Then attention enhanced features of different modalities are mapped into a joint representation to guide the caption generation. Moreover, we design a new global-aware decoder, which leverages the concatenated representation of triple-steam features and the joint attention representation to obtain global visual guidance information, further refine the complex multimodal reasoning. To verify the effectiveness of our feature fusion model, we perform extensive experiments on the highly competitive MSCOCO dataset to evaluate the model quantitatively and qualitatively. The results illustrate that the proposed framework outperforms many state-of-the-art image captioning approaches in various evaluation metrics, and generates more accurate and abundant captions.
Nannan Hu, Yue Ming 0001, Chunxiao Fan 0001, Boyang Lyu
IEEE Trans. Multim.1
2022 SSLNet: A network for cross-modal sound source localization in visual scenes
Yue Ming 0001, Nannan Hu
Neurocomputing3
2021 Faster-FCoViAR: Faster Frequency-Domain Compressed Video Action Recognition
Xia Jia, Yue Ming 0001, Jiangwan Zhou, Nannan Hu
BMVC6