Fudong Nian

dblp:168/0848 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
15since 2021 · last 2025
0000-0001-9604-7564ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DCMCF: Dynamic Cross-Modal Context Fusion for Image-Based Person Search
abstract
ABSTRACT Image‐based person search, which locates and identifies a query person from uncropped gallery images, remains challenging due to occlusions, viewpoint changes, and ambiguous appearances. While existing methods increasingly leverage visual context, they are fundamentally limited by static fusion strategies for multi‐view features and their neglect of semantic context from other modalities. To overcome these limitations, we propose the dynamic cross‐modal context fusion (DCMCF) framework, which establishes a core multi‐modal fusion architecture by explicitly leveraging self‐generated textual descriptions to provide semantics‐driven matching guidance. DCMCF consists of three core components: (1) multi‐modal feature extraction, which employs an enhanced multi‐view visual encoder with dynamic gating and a zero‐shot pipeline using pre‐trained VLMs to generate multi‐level textual descriptions from images; (2) a dynamic cross‐modal fusion (DCMF) module, which hierarchically and adaptively fuses the visual and textual features through cross‐modal attention and input‐dependent gating; and (3) a text‐guided inter‐image group context ranking (T‐IGCR) algorithm, which refines retrieval results by measuring holistic image consistency in both visual and textual spaces. Experiments on CUHK‐SYSU and PRW datasets demonstrate state‐of‐the‐art performance, achieving 96.9%/97.3% (mAP/top‐1) and 56.1%/91.2% , respectively. This work demonstrates the significant potential of dynamic cross‐modal context fusion for advancing image‐based person search.
Fudong Nian, Yingfang Wang, Aoyu Liu, Yun Fu 0009, Yanhong Gu
IET Image Process.1
2025 Dual-Branch Enhancement and Multi-Modal Fusion for Low-Light Visible Polarization Image Object Detection in Dense Smog Environments
abstract
ABSTRACT In scenarios with heavy smog, the accuracy of object detection in low‐light visible polarization images significantly decreases. To address this issue, we propose a dual‐branch enhancement and multi‐modal fusion network for object detection in low‐light visible polarization images in dense smog environments. Specifically, the network consists of an image enhancement stage and an object detection stage. In the image enhancement stage, a dual‐branch enhancement structure comprising greyscale feature map prediction and atmospheric light transmission network is proposed to remove noise from the images and enhance texture information, jointly generating enhanced visible polarization images. In the object detection stage, feature maps of the enhanced visible polarization images and the degree of visible polarization images are fused, and their fused texture‐enhanced feature maps are fed into the detection module for object detection. Additionally, we have collected a dataset of low‐light visible polarization images under real smog conditions. Extensive experiments demonstrate that our method can generate visually improved enhanced images and significantly increase detection accuracy and the number of detected objects in low‐light and dense smog environments.
Fudong Nian, Jianguo Huang, Teng Li 0001
IET Image Process.3
2025 Rwkv-vg: visual grounding with RWKV-driven encoder-decoder framework
Fudong Nian, Yanhong Gu, Aoyu Liu, Fanding Li
Multim. Syst.1
2025 Visual-language collaborative multimodal transformer network for group activity detection in surveillance videos
Fudong Nian, Weijie Lu, Chengqian Li, Yun Fu 0009, Zhize Wu
Multim. Syst.1
2025 Local and global self-attention enhanced graph convolutional network for skeleton-based action recognition
Zhize Wu, Long Wan, Teng Li 0001, Fudong Nian
Pattern Recognit.5
2024 Fine-Grained Cross-Modal Contrast Learning for Video-Text Retrieval
Yanhong Gu, Fudong Nian
ICIC (5)4
2024 Swelling-ViT: Rethink Data-Efficient Vision Transformer from Locality
Chuanrui Hu, Fudong Nian, Teng Li 0001
PRCV (4)4
2024 Video-text retrieval via multi-modal masked transformer and adaptive attribute-aware graph convolutional network
Yining Sun, Fudong Nian
Multim. Syst.3
2023 VQA-CLPR: Turning a Visual Question Answering Model into a Chinese License Plate Recognizer
Xuhao Jiang, Yining Sun, Weiya Ni, Fudong Nian
ICIG (2)5
2023 COME: Clip-OCR and Master ObjEct for text image captioning
Yining Sun, Fudong Nian, Maofei Zhu, Wenliang Tang
Image Vis. Comput.3
2023 Hierarchical cross-modal contextual attention network for visual grounding
Yining Sun, Yuxia Hu, Fudong Nian
Multim. Syst.5
2022 Stitching High Resolution Notebook Keyboard Surface Based on Halcon Calibration
Zuchang Ma, Yining Sun, Fudong Nian
ICIC (1)5
2022 Visible light polarization image desmogging via Cycle Convolutional Neural Network
Teng Li 0001, Yuzhou Zeng, Fudong Nian
Multim. Syst.6
2021 Geometric Context Sensitive Loss and Its Application for Nonrigid Structure from Motion
Fudong Nian, Shimeng Yang, Xia Chen 0008
ICIG (3)1
2021 Distributed Policy Evaluation with Fractional Order Dynamics in Multiagent Reinforcement Learning
abstract
The main objective of multiagent reinforcement learning is to achieve a global optimal policy. It is difficult to evaluate the value function with high-dimensional state space. Therefore, we transfer the problem of multiagent reinforcement learning into a distributed optimization problem with constraint terms. In this problem, all agents share the space of states and actions, but each agent only obtains its own local reward. Then, we propose a distributed optimization with fractional order dynamics to solve this problem. Moreover, we prove the convergence of the proposed algorithm and illustrate its effectiveness with a numerical example.
Wei Wang 0176, Zhongtian Mao, Ruwen Jiang, Fudong Nian, Teng Li 0001
Secur. Commun. Networks5
2020 Relative coordinates constraint for face alignment
Fudong Nian, Teng Li 0001, Bing-Kun Bao, Changsheng Xu
Neurocomputing1
2017 Multi-Modal Knowledge Representation Learning via Webly-Supervised Relationships Mining
abstract
Knowledge representation learning (KRL) encodes enormous structured information with entities and relations into a continuous low-dimensional semantic space. Most conventional methods solely focus on learning knowledge representation from single modality, yet neglect the complementary information from others. The more and more rich available multi-modal data on Internet also drive us to explore a novel approach for KRL in multi-modal way, and overcome the limitations of previous single-modal based methods. This paper proposes a novel multi-modal knowledge representation learning (MM-KRL) framework which attempts to handle knowledge from both textual and visual modal web data. It consists of two stages, i.e., webly-supervised multi-modal relationship mining, and bi-enhanced cross-modal knowledge representation learning. Compared with existing knowledge representation methods, our framework has several advantages: (1) It can effectively mine multi-modal knowledge with structured textual and visual relationships from web automatically. (2) It is able to learn a common knowledge space which is independent to both task and modality by the proposed Bi-enhanced Cross-modal Deep Neural Network (BC-DNN). (3) It has the ability to represent unseen multi-modal relationships by transferring the learned knowledge with isolated seen entities and relations into unseen relationships. We build a large-scale multi-modal relationship dataset (MMR-D) and the experimental results show that our framework achieves excellent performance in zero-shot multi-modal retrieval and visual relationship recognition.
Fudong Nian, Bing-Kun Bao, Teng Li 0001, Changsheng Xu
ACM Multimedia1
2017 Learning explicit video attributes from mid-level representation for video captioning
Fudong Nian, Teng Li 0001, Yan Wang 0059, Xinyu Wu 0001, Bingbing Ni, Changsheng Xu
Comput. Vis. Image Underst.1
2017 Robust face anti-spoofing with depth information
Yan Wang 0059, Fudong Nian, Teng Li 0001, Kongqiao Wang
J. Vis. Commun. Image Represent.2
2016 Pornographic image detection utilizing deep convolutional neural networks
Fudong Nian, Teng Li 0001, Yan Wang 0059, Mingliang Xu 0001
Neurocomputing1
2016 Dense crowd counting from still images with convolutional neural networks
Yaocong Hu, Fudong Nian, Yan Wang 0059, Teng Li 0001
J. Vis. Commun. Image Represent.3
2016 Efficient video copy detection using multi-modality and dynamic path search
Teng Li 0001, Fudong Nian, Xinyu Wu 0001, Qingwei Gao, Yixiang Lu
Multim. Syst.2
2016 Efficient near-duplicate image detection with a local-based binary representation
Fudong Nian, Teng Li 0001, Xinyu Wu 0001, Qingwei Gao, Feifeng Li
Multim. Tools Appl.1