VLDB 2026 Research / reviewers in the wild / expert
Chunxiao Fan 0001
dblp:83/8378-1
· DBLP profile ↗
24ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0002-3607-4904ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MIND: A Multi-agent Framework for Zero-shot Harmful Meme DetectionabstractThe rapid expansion of memes on social media has highlighted the urgent need for effective approaches to detect harmful content. However, traditional data-driven approaches struggle to detect new memes due to their evolving nature and the lack of up-to-date annotated data. To address this issue, we propose MIND, a multi-agent framework for zero-shot harmful meme detection that does not rely on annotated data. MIND implements three key strategies: 1) We retrieve similar memes from an unannotated reference set to provide contextual information. 2) We propose a bi-directional insight derivation mechanism to extract a comprehensive understanding of similar memes. 3) We then employ a multi-agent debate mechanism to ensure robust decision-making through reasoned arbitration. Extensive experiments on three meme datasets demonstrate that our proposed framework not only outperforms existing zero-shot approaches but also shows strong generalization across different model architectures and parameter scales, providing a scalable solution for harmful meme detection. Chunxiao Fan 0001, Haoran Lou, Yuexin Wu, Kaiwei Deng |
ACL (1) | 2 |
| 2025 | LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
Haoran Lou, Chunxiao Fan 0001, Yuexin Wu |
ICCV | 2 |
| 2025 | DSSLNet: dual-stream self-supervised learning network for end-to-end speech recognition
Boyang Lyu, Chunxiao Fan 0001, Yue Ming 0001, Nannan Hu |
Appl. Intell. | 2 |
| 2023 | Diffusion-Adapter: Text Guided Image Manipulation with Frozen Diffusion Models
Rongting Wei, Chunxiao Fan 0001, Yuexin Wu |
ICANN (2) | 2 |
| 2023 | MAENet: A novel multi-head association attention enhancement network for completing intra-modal interaction in image captioning
Nannan Hu, Chunxiao Fan 0001, Yue Ming 0001 |
Neurocomputing | 2 |
| 2023 | En-HACN: Enhancing Hybrid Architecture With Fast Attention and Capsule Network for End-to-end Speech RecognitionabstractAutomatic speech recognition (ASR) is a fundamental technology in the field of artificial intelligence. End-to-end (E2E) ASR is favored for its state-of-the-art performance. However, E2E speech recognition still faces speech spatial information loss and text local information loss, which results in the increase of deletion and substitution errors during inference. To overcome this challenge, we propose a novel Enhancing Hybrid Architecture with Fast Attention and Capsule Network (termed En-HACN), which can model the position relationships between different acoustic unit features to improve the discriminability of speech features while providing the text local information during inference. Firstly, a new CNN-Capsule Network (CNN-Caps) module is proposed to capture the spatial information in the spectrogram through the capsule output and dynamic routing mechanism. Then, we design a novel hybrid structure of LocalGRU Augmented Decoder (LA-decoder) that generates text hidden representations to obtain text local information of the target sequences. Finally, we introduce fast attention instead of self-attention in En-HACN, which improves the generalization ability and efficiency of the model in long utterances. Experiments on corpora Aishell-1 and Librispeech demonstrate that our En-HACN has achieved the state-of-the-art compared with existing works. Besides, experiments on the long utterances dataset based on Aishell-1-long show that our model has a high generalization ability and efficiency. Boyang Lyu, Chunxiao Fan 0001, Yue Ming 0001, Panzi Zhao, Nannan Hu |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | TSFNet: Triple-Steam Image CaptioningabstractImage captioning is a challenging task that generates a natural language description based on the visual understanding of the given image. Significant region representation is a milestone in image captioning. Despite the great success of existing region-based works, they only focus on the salient objects and encode these objects independently, still plagued by the lack of global contextual information and visual relationships. In fact, the global contextual information and structured visual relationships are exactly the merits of traditional grid features and emerging scene graph features. In this paper, we present a Triple-Steam Feature Fusion Network (TSFNet) to leverage the complementary advantages of the grid, region, and scene graph triple-steam visual representations in image captioning. Concretely, in our TSFNet, a novel Dual-level Attention (DA) mechanism is proposed to simultaneously explore visual intrinsic properties and word-related attributes uniformly of different features. Then attention enhanced features of different modalities are mapped into a joint representation to guide the caption generation. Moreover, we design a new global-aware decoder, which leverages the concatenated representation of triple-steam features and the joint attention representation to obtain global visual guidance information, further refine the complex multimodal reasoning. To verify the effectiveness of our feature fusion model, we perform extensive experiments on the highly competitive MSCOCO dataset to evaluate the model quantitatively and qualitatively. The results illustrate that the proposed framework outperforms many state-of-the-art image captioning approaches in various evaluation metrics, and generates more accurate and abundant captions. Nannan Hu, Yue Ming 0001, Chunxiao Fan 0001, Boyang Lyu |
IEEE Trans. Multim. | 3 |
| 2022 | M-CoTransT: Adaptive spatial continuity in visual trackingabstractAbstract Visual tracking is an important area in computer vision. Based on the Siamese network, current tracking methods employ the self‐attention block in convolutional networks to extract semantic features containing the image structure information of an object. However, spatial continuity is a point of contradiction between two seemingly unrelated challenges, that is, occlusion and similar distractor, in tracking methods. At the same time, it is a spatially discontinuous task to locate a target reappearing after occlusion accurately. The prediction of bounding boxes should be constrained by spatial continuity to prevent them from jumping into similar distractors. This study proposes a novel tracking method for introducing spatial continuity in visual tracking called M‐CoTransT; the novel tracking method is developed through the confidence‐based adaptive Markov motion model (M‐model) and a novel correlation‐based feature fusion network (CoTransT). In particular, the M‐model provides confidence for the nodes of the Markov motion model to estimate the motion state continuity. It also predicts a more accurate search region for CoTransT, which then adds a cross‐correlation branch into the self‐attention tracking network to enhance the continuity of target appearance in the feature space. Extensive experiments on five challenging datasets (LaSOT, GOT‐10k, TrackingNet, OTB‐2015 and UAV123) demonstrated the effectiveness of the proposed M‐CoTransT in visual tracking. Chunxiao Fan 0001, Runqing Zhang, Yue Ming 0001 |
IET Comput. Vis. | 1 |
| 2022 | Knowledge base question answering via path matching
Chunxiao Fan 0001, Wentong Chen, Yuexin Wu |
Knowl. Based Syst. | 1 |
| 2022 | CORNet: Context-Based Ordinal Regression Network for Monocular Depth EstimationabstractMonocular depth estimation, as one of the fundamental tasks of computer vision, plays a crucial role in three-dimensional (3D) scene understanding and perception. Usually, deep learning methods recover monocular depth maps using continuous regression manners by minimizing the errors between the ground-truth depth and the predicted depth. However, fine depth features may not be fully captured through layer-by-layer coding, which is prone to low spatial resolution depth maps and insufficient details. Furthermore, it usually converges slowly and suffers from unsatisfactory results. To tackle these issues, we propose a novel model, named context-based ordinal regression network (CORNet), to reconstruct monocular depth maps in the ordinal regression manner with context information in this paper. Firstly, we put forward a novel context-based encoder with a feature transformation (FT) module to learn context information and details from inputs, and output multi-scale feature maps. Then, we design a boundary enhancement module (BEM) with a spatial attention mechanism following each operation of feature fusion, which captures boundary features in the scene to enhance the border depth. Finally, a feature optimization module (FOM) is designed to fuse and optimize the multi-scale features and boundary features to strengthen depth learning. What’s more, we introduce an ordinal weighted inference to predict depth maps from probabilities and discretization values. Experiments and results on two challenging datasets, KITTI and NYU Depth V2, demonstrate that our proposed CORNet can estimate monocular depth maps effectively and obtain superior performance in capturing geometric features over existing methods. Xuyang Meng, Chunxiao Fan 0001, Yue Ming 0001, Hui Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | MP-LN: motion state prediction and localization network for visual object tracking
Chunxiao Fan 0001, Runqing Zhang, Yue Ming 0001 |
Vis. Comput. | 1 |
| 2021 | Re-Identify Deformable Targets for Visual Tracking
Runqing Zhang, Chunxiao Fan 0001, Yue Ming 0001 |
PRCV (1) | 2 |
| 2021 | Deep learning for monocular depth estimation: A review
Yue Ming 0001, Xuyang Meng, Chunxiao Fan 0001, Hui Yu 0001 |
Neurocomputing | 3 |
| 2020 | An Effective Hierarchical Resolution Learning Method for Low-Resolution Targets TrackingabstractSuffering from the low-resolution target's visual quality, the precisions of visual object trackers are reduced. This paper proposes an effective hierarchical resolution learning method for low-resolution targets tracking, abbreviated as HRT. We adopt a hierarchical structure to exploit information from different resolution levels. (1) At the high level: the super-resolution (SR) images, determining the target's shape, contains richer image textures and clearer target contours, and transmits the search region to the low level. (2) At the low level: low-resolution (LR) images maintain the spatial structure information of the original target, providing the precise center coordinates of the target. Experimental results demonstrate the effectiveness of the proposed tracker, which HRT achieves 90.3% precision on OTB100 LR sequences and 78.5% precision on LR sequences from UAV123 datasets, gaining 2.0%, 2.4% improvement over state-of-the-art trackers respectively. Runqing Zhang, Chunxiao Fan 0001, Yue Ming 0001, Hao Fu 0013, Xuyang Meng |
ICIP | 2 |
| 2020 | MD-ST: Monocular Depth Estimation Based on Spatio-Temporal Correlation Features
Xuyang Meng, Chunxiao Fan 0001, Yue Ming 0001, Runqing Zhang, Panzi Zhao |
PRCV (1) | 2 |
| 2020 | A novel perceptual loss function for single image super-resolution
Chunxiao Fan 0001, Yong Li 0025, Yang Li 0032 |
Multim. Tools Appl. | 2 |
| 2018 | Sparse Tikhonov-Regularized Hashing for Multi-Modal LearningabstractThis paper mainly focuses on the role of regularization in Multi-Modal Learning (MML). Existing MML studies devote most of the efforts in maximizing the consensus of models from cues of different modalities. However, regularization methods are still far from fully explored. To fill in this gap, we propose a compact and efficient coding solution, termed by sparse Tikhonov-Regularized Hashing (STRH). The STRH enforces both the ℓ0-norm induced sparsity constraints and the Tikhonov regularization on the binary solution vectors which maximize cross-modal correlation. In addition, we raise the concerns on the challenging testing scenario of `Multi-modal Learning and Single-modal Prediction' (MLSP). Finally, we demonstrate that the STRH is an efficient hashing solutions by showing its superiority under the MLSP scenario. Lei Tian 0002, Xiaopeng Hong, Chunxiao Fan 0001, Yue Ming 0001, Matti Pietikäinen, Guoying Zhao 0001 |
ICIP | 3 |
| 2018 | Sparse projections matrix binary descriptors for face recognition
Chunxiao Fan 0001, Lei Tian 0002, Yue Ming 0001, Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen |
Neurocomputing | 1 |
| 2018 | Improving deep neural network with Multiple Parametric Exponential Linear Units
Yang Li 0032, Chunxiao Fan 0001, Yong Li 0025, Yue Ming 0001 |
Neurocomputing | 2 |
| 2017 | Learning spherical hashing based binary codes for face recognition
Lei Tian 0002, Chunxiao Fan 0001, Yue Ming 0001 |
Multim. Tools Appl. | 2 |
| 2016 | Learning iterative quantization binary codes for face recognition
Lei Tian 0002, Chunxiao Fan 0001, Yue Ming 0001 |
Neurocomputing | 2 |
| 2015 | 3D human behavior recognition based on spatiotemporal texture featuresabstractNowadays, more and more activity recognition algorithms begin to improve recognition performance by combining the RGB and depth information. Although, the space-time volumes (STV) algorithm and the space-time local features algorithm can combine the RGB and depth information effectively, they also have their own defects. Such as they need expensive computational cost and they are not suitable for modeling nonperiodic activity. In this paper, we propose a novel algorithm for three dimensional human activity recognition that combines spatial-domain local texture features and spatio-temporal local texture features. On the one hand, in order to extract spatial local texture features, we mix the RGB and depth image sequence which have been applied with ViBe (Visual Background extractor) and binarization operator. Then we obtain the RGB-MOHBBI and depth-MOBHBI respectively and perform intersect operation on them. Afterwards, we extract LBP feature from the mixed MOHBBI to describe spatial domain feature. On the other hand, we follow the same background subtraction and binarization method to process the RGB and depth image sequences and get the spatial-temporal local texture features. And then, we project the three dimensional image volume on plane X-T and plane Y-T to get the spatio-temporal behavior volume change image to which we apply LBP operator to extract features that can represent human activity feature in spatio-temporal domain. At last, we combine the two local features that are extracted by LBP algorithm as one integrated feature of our model final output. Extensive experiments are conducted on the BUPT Arm Activity Dataset and the BUPT Arm And Finger Activity Dataset. The experimental results demonstrate the algorithm we proposed in this paper can make up for the deficiency of traditional activity recognition algorithms effectively and provide excellent experiment results on different databases of various complexities. Chunxiao Fan 0001, Lei Tian 0002, Guangchao Wang, Yue Ming 0001, Jiakun Shi, Yi Jin 0001 |
HSI | 1 |
| 2015 | Reliable and Fast Mapping of Keypoints on Large-Size Remote Sensing Images by Use of Multiresolution and Global InformationabstractThis letter proposes a multiresolution technique to address the high computational cost in remote sensing image registration. The scale-invariant feature transform is applied to detect keypoints and descriptors, and then, global information combined with descriptors is utilized to establish keypoint mappings. Keypoints are first classified according to their octaves. Then, in the lowest resolution, the keypoints of the largest octave are mapped with descriptors and the global information, giving an initial affine transformation$T_0$. In the next octave, the keypoints of the second largest octave are mapped by employing$T_0$to narrow the space of matching keypoints. By this means, the process of establishing keypoint correspondences is conducted from one resolution (octave) to the next as the obtained transformation gets finer until we get to the highest resolution. Due to the high computational expense of computing global information, the proposed technique is important for aligning large-size remote sensing imagery. Experimental results show that the proposed method can achieve a comparable registration accuracy but with a less computational cost than directly building keypoint mappings on images of large size. Yong Li 0025, Wei Qiao 0002, Hongbin Jin, Jing Jing 0003, Chunxiao Fan 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2013 | Coordinated transceiver in MIMO heterogeneous network with physical-layer network codingabstractIn this contribution, a multiple-input multiple-output heterogeneous network with physical-layer network coding is considered. The article proposes a coordinated joint transmitter and receiver design to tackle the interference problem between two-way relaying channel in small cell and uplink channel in macro cell. We design the transceiver on the basis that the data streams from two-way communication nodes are aligned to the same direction to perform multi-stream decode-and-forward physical-layer network coding, while other streams are adjusted to orthogonal direction for spatial multiplexing. Coordinated beam-forming and joint multi-cell signal processing are both considered. Based on the criterion that minimizes system-wide mean square error, we formulate the transceiver optimization problem with transmit power constraints on each user and provide an approximate optimal solution for the non-convex problem through iterative algorithm. During the iteration, each of the transmitters and receivers is solved by Lagrange multiplier method and Karush-Kuhn-Tucker condition. The scheme reduces interferences when all the nodes transmit multi-stream symbols simultaneously and significantly improves the system performance. Numerical results such as BER and MSE performances are provided to support the proposed scheme and show that a below 10-3BER could be reached when SNR≥8dB which is close to the lower bound. Zhigang Wen, Zibo Meng, Dongjian Chen, Chunxiao Fan 0001 |
PIMRC | 6 |