EDBT 2026 Demo / reviewers in the wild / expert
Fudong Nian
dblp:168/0848
· DBLP profile ↗
23ranked-venue papers
9as first author
15since 2021 · last 2025
0000-0001-9604-7564ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DCMCF: Dynamic Cross-Modal Context Fusion for Image-Based Person SearchabstractABSTRACT Image‐based person search, which locates and identifies a query person from uncropped gallery images, remains challenging due to occlusions, viewpoint changes, and ambiguous appearances. While existing methods increasingly leverage visual context, they are fundamentally limited by static fusion strategies for multi‐view features and their neglect of semantic context from other modalities. To overcome these limitations, we propose the dynamic cross‐modal context fusion (DCMCF) framework, which establishes a core multi‐modal fusion architecture by explicitly leveraging self‐generated textual descriptions to provide semantics‐driven matching guidance. DCMCF consists of three core components: (1) multi‐modal feature extraction, which employs an enhanced multi‐view visual encoder with dynamic gating and a zero‐shot pipeline using pre‐trained VLMs to generate multi‐level textual descriptions from images; (2) a dynamic cross‐modal fusion (DCMF) module, which hierarchically and adaptively fuses the visual and textual features through cross‐modal attention and input‐dependent gating; and (3) a text‐guided inter‐image group context ranking (T‐IGCR) algorithm, which refines retrieval results by measuring holistic image consistency in both visual and textual spaces. Experiments on CUHK‐SYSU and PRW datasets demonstrate state‐of‐the‐art performance, achieving 96.9%/97.3% (mAP/top‐1) and 56.1%/91.2% , respectively. This work demonstrates the significant potential of dynamic cross‐modal context fusion for advancing image‐based person search. Fudong Nian, Yingfang Wang, Aoyu Liu, Yun Fu 0009, Yanhong Gu |
IET Image Process. | 1 |
| 2025 | Dual-Branch Enhancement and Multi-Modal Fusion for Low-Light Visible Polarization Image Object Detection in Dense Smog EnvironmentsabstractABSTRACT In scenarios with heavy smog, the accuracy of object detection in low‐light visible polarization images significantly decreases. To address this issue, we propose a dual‐branch enhancement and multi‐modal fusion network for object detection in low‐light visible polarization images in dense smog environments. Specifically, the network consists of an image enhancement stage and an object detection stage. In the image enhancement stage, a dual‐branch enhancement structure comprising greyscale feature map prediction and atmospheric light transmission network is proposed to remove noise from the images and enhance texture information, jointly generating enhanced visible polarization images. In the object detection stage, feature maps of the enhanced visible polarization images and the degree of visible polarization images are fused, and their fused texture‐enhanced feature maps are fed into the detection module for object detection. Additionally, we have collected a dataset of low‐light visible polarization images under real smog conditions. Extensive experiments demonstrate that our method can generate visually improved enhanced images and significantly increase detection accuracy and the number of detected objects in low‐light and dense smog environments. Fudong Nian, Jianguo Huang, Teng Li 0001 |
IET Image Process. | 3 |
| 2025 | Rwkv-vg: visual grounding with RWKV-driven encoder-decoder framework
Fudong Nian, Yanhong Gu, Aoyu Liu, Fanding Li |
Multim. Syst. | 1 |
| 2025 | Visual-language collaborative multimodal transformer network for group activity detection in surveillance videos
Fudong Nian, Weijie Lu, Chengqian Li, Yun Fu 0009, Zhize Wu |
Multim. Syst. | 1 |
| 2025 | Local and global self-attention enhanced graph convolutional network for skeleton-based action recognition
Zhize Wu, Long Wan, Teng Li 0001, Fudong Nian |
Pattern Recognit. | 5 |
| 2024 | Fine-Grained Cross-Modal Contrast Learning for Video-Text Retrieval
Yanhong Gu, Fudong Nian |
ICIC (5) | 4 |
| 2024 | Swelling-ViT: Rethink Data-Efficient Vision Transformer from Locality
Chuanrui Hu, Fudong Nian, Teng Li 0001 |
PRCV (4) | 4 |
| 2024 | Video-text retrieval via multi-modal masked transformer and adaptive attribute-aware graph convolutional network
Yining Sun, Fudong Nian |
Multim. Syst. | 3 |
| 2023 | VQA-CLPR: Turning a Visual Question Answering Model into a Chinese License Plate Recognizer
Xuhao Jiang, Yining Sun, Weiya Ni, Fudong Nian |
ICIG (2) | 5 |
| 2023 | COME: Clip-OCR and Master ObjEct for text image captioning
Yining Sun, Fudong Nian, Maofei Zhu, Wenliang Tang |
Image Vis. Comput. | 3 |
| 2023 | Hierarchical cross-modal contextual attention network for visual grounding
Yining Sun, Yuxia Hu, Fudong Nian |
Multim. Syst. | 5 |
| 2022 | Stitching High Resolution Notebook Keyboard Surface Based on Halcon Calibration
Zuchang Ma, Yining Sun, Fudong Nian |
ICIC (1) | 5 |
| 2022 | Visible light polarization image desmogging via Cycle Convolutional Neural Network
Teng Li 0001, Yuzhou Zeng, Fudong Nian |
Multim. Syst. | 6 |
| 2021 | Geometric Context Sensitive Loss and Its Application for Nonrigid Structure from Motion
Fudong Nian, Shimeng Yang, Xia Chen 0008 |
ICIG (3) | 1 |
| 2021 | Distributed Policy Evaluation with Fractional Order Dynamics in Multiagent Reinforcement LearningabstractThe main objective of multiagent reinforcement learning is to achieve a global optimal policy. It is difficult to evaluate the value function with high-dimensional state space. Therefore, we transfer the problem of multiagent reinforcement learning into a distributed optimization problem with constraint terms. In this problem, all agents share the space of states and actions, but each agent only obtains its own local reward. Then, we propose a distributed optimization with fractional order dynamics to solve this problem. Moreover, we prove the convergence of the proposed algorithm and illustrate its effectiveness with a numerical example. Wei Wang 0176, Zhongtian Mao, Ruwen Jiang, Fudong Nian, Teng Li 0001 |
Secur. Commun. Networks | 5 |
| 2020 | Relative coordinates constraint for face alignment
Fudong Nian, Teng Li 0001, Bing-Kun Bao, Changsheng Xu |
Neurocomputing | 1 |
| 2017 | Multi-Modal Knowledge Representation Learning via Webly-Supervised Relationships MiningabstractKnowledge representation learning (KRL) encodes enormous structured information with entities and relations into a continuous low-dimensional semantic space. Most conventional methods solely focus on learning knowledge representation from single modality, yet neglect the complementary information from others. The more and more rich available multi-modal data on Internet also drive us to explore a novel approach for KRL in multi-modal way, and overcome the limitations of previous single-modal based methods. This paper proposes a novel multi-modal knowledge representation learning (MM-KRL) framework which attempts to handle knowledge from both textual and visual modal web data. It consists of two stages, i.e., webly-supervised multi-modal relationship mining, and bi-enhanced cross-modal knowledge representation learning. Compared with existing knowledge representation methods, our framework has several advantages: (1) It can effectively mine multi-modal knowledge with structured textual and visual relationships from web automatically. (2) It is able to learn a common knowledge space which is independent to both task and modality by the proposed Bi-enhanced Cross-modal Deep Neural Network (BC-DNN). (3) It has the ability to represent unseen multi-modal relationships by transferring the learned knowledge with isolated seen entities and relations into unseen relationships. We build a large-scale multi-modal relationship dataset (MMR-D) and the experimental results show that our framework achieves excellent performance in zero-shot multi-modal retrieval and visual relationship recognition. Fudong Nian, Bing-Kun Bao, Teng Li 0001, Changsheng Xu |
ACM Multimedia | 1 |
| 2017 | Learning explicit video attributes from mid-level representation for video captioning
Fudong Nian, Teng Li 0001, Yan Wang 0059, Xinyu Wu 0001, Bingbing Ni, Changsheng Xu |
Comput. Vis. Image Underst. | 1 |
| 2017 | Robust face anti-spoofing with depth information
Yan Wang 0059, Fudong Nian, Teng Li 0001, Kongqiao Wang |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Pornographic image detection utilizing deep convolutional neural networks
Fudong Nian, Teng Li 0001, Yan Wang 0059, Mingliang Xu 0001 |
Neurocomputing | 1 |
| 2016 | Dense crowd counting from still images with convolutional neural networks
Yaocong Hu, Fudong Nian, Yan Wang 0059, Teng Li 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Efficient video copy detection using multi-modality and dynamic path search
Teng Li 0001, Fudong Nian, Xinyu Wu 0001, Qingwei Gao, Yixiang Lu |
Multim. Syst. | 2 |
| 2016 | Efficient near-duplicate image detection with a local-based binary representation
Fudong Nian, Teng Li 0001, Xinyu Wu 0001, Qingwei Gao, Feifeng Li |
Multim. Tools Appl. | 1 |